
完整作品展
技术栈
61 projects


根据已发布的代理标准评判AI产品,包含可检查证据和社区投票。
@katyorby · X
i built — a local receipt for claude code runs. your check says whether the workspace passes now; the transcript supplies the activity counts. no transcript upload and no magical autonomy score.

提交目标和标准,让AI模型面板评估,由仲裁员基于证据作出判决。
nadermx · HN
Referee.Chat - Set the goal. An AI panel works. Referee clears it done

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

检查AI推理轨迹,评估模型的可信度。
malik_dixon1 · Product Hunt
TraceLogicAI: AI Architecture Evaluation Compare AI architectures with evidence, not guesswork

评估AI系统的负责任和伦理实践。
@TheWhiz351 · X
I built two versions of the same responsible AI app using different vibe-coding platforms. Perplexity Computer: Base44: Try both. Which has the better design and user experience? #VibeCoding #ResponsibleAI

指挥专门的AI工作人员执行创意任务,每步都需人工审核。
@VisionAIWS · X

查看和对比主流AI模型的公众意见和基准评分。
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

AI平台,提问、生成图像、语音交互,辅助更好的决策。
@Akinzoooo · X

免费AI评估,找出孩子的学习优势和不足,推荐个性化Khan Academy课程。
@edsull · X
My vibe coding team of agents set up a series of AI assessments for each subject and each grade level. Check it out!

上传手写考试答卷,获得即时AI评分和详细反馈。
Xaminix
AI Powered Answer Evaluation for CA/CS/CMA

Submit your AI drafts for human review and editing. Verified reviewers refine tone, fix formulaic AI patterns, and return polished content with a change report in ~2 hours.
@WeCatchAI · X