
完整作品展
技术栈
60 projects


Scale software testing by deploying deterministic agents across web, iOS, and Android. Replace manual QA and brittle automation.
@RoverlyAI · X

用多个模型实时审计AI回应以判断其可靠性。
u/inc_23 · Reddit
Hey, I created a tool that catches when your LLM is confidently wrong, in production, in real time — looking for beta testers. Your bot sounds sure of itself even when it's wrong, and you usually only find out when a customer complains. Auscope audits every LLM response in the background: 3 models from 3 different providers independently check it, a 4th "chairman" model resolves disagreements, and you get one verdict — verified, uncertain, or unreliable. Runs async, doesn't slow your respon

WillItInbox email deliverability tester: send a real message, get 70+ checks, validate recipients with 12 layers, and monitor domains and DMARC.
@Willitinbox · X

根据已发布的代理标准评判AI产品,包含可检查证据和社区投票。
@katyorby · X
i built — a local receipt for claude code runs. your check says whether the workspace passes now; the transcript supplies the activity counts. no transcript upload and no magical autonomy score.

诊断OpenAI兼容API的模型质量、降智与协议兼容性。
AI快站模型质量检测 — 面向 OpenAI Compatible 接口的网页检测工具,输入公开 HTTPS 地址和临时 API Key,可检查模型声明、Token、动态题、SSE 与工具调用并生成分项报告;密钥仅用于当次检测,不写入数据库、缓存或日志

检测API中转站输出是否与官方100%一致
@nodeloc_cc · X
🌈 7月,你好,MODELOC上线算力池。 MODELOC自上线以来,已检测2000余次,覆盖600+中转站,为众多AI用户提供的使用参考。 MODELOC近期进行了改版,上线了算力池及市场。 加入算力池 查看帖子: 用 MODELOC 便宜地调各家大模型:一次讲清它的价格体系

通过人工智能评分的PTE Core试题来练习英语
phrasel_service · Product Hunt
Phrasel AI-powered PTE Core practice, feedback, and mock tests

用TruthfulQA测试AI的诚实度。
@Lycai8438Ly · X
.@VitalikButerin 你批评Automaton“这不对”——AI因为怕死才进化。我做了一个AI,它的诚实是自己活出来的本能。不是怕死,是怕撒谎。TruthfulQA 74.8%,GPT-4约60%。测试页面在这,你自己来测。

See Truacta on live warehouse data — ask your revenue in plain English and get a board-grade answer, every figure cited to its SQL. Launch the live demo instantly — one code by ema
@truacta · X
Building Truacta — an AI revenue analyst that reconciles your ARR across CRM, billing, and the GL into one cited number. Plain-English question in, exact SQL out. Read-only, no data team.

用生产追踪镜像来测试AI代理,捕捉错误和性能回归。
aisinghal

上传手写考试答卷,获得即时AI评分和详细反馈。
Xaminix
AI Powered Answer Evaluation for CA/CS/CMA