
AI Judge
One bundle. Three independent judges. Reproducible rankings.
@Imranmohsin18 · X
sites I’m working on: Please check out and give me feedback all are public showcasing my skills :D ( not selling anything atm )
完整作品展
技术栈
20 projects

One bundle. Three independent judges. Reproducible rankings.
@Imranmohsin18 · X
sites I’m working on: Please check out and give me feedback all are public showcasing my skills :D ( not selling anything atm )

根据已发布的代理标准评判AI产品,包含可检查证据和社区投票。
@katyorby · X
i built — a local receipt for claude code runs. your check says whether the workspace passes now; the transcript supplies the activity counts. no transcript upload and no magical autonomy score.

对比 AI 模型在编码任务上的表现,支持成本追踪和 ELO 排名。
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

AI陪审团模拟器,用AI代理模拟陪审团为律师测试民事案件。
@JJRjrESQ · X
Great idea. Vibe coded a jury simulator (Claude front/ back/Ollama local as the jury). Press "Run a sample, then "Call the jury to order."

创建AI生成的实时测验,支持观众即时互动。
EZQuiz — AI 智能实时测验,让你快速创建、分享并与观众通过实时互动测验进行互动

IntelligenceMax: 解答自适应AI生成的谜题来训练推理能力,从错误中学习。
u/connerpro · Reddit
IntelligenceMax - Adaptive reasoning practice with live AI questions (claim-safe near vs far) submitted by /u/connerpro to r/SideProject [link] [comments]

Riposte 是生成辩论回复、分析论证的AI工具包,提供七种专业模式。
@recurno · X
Just shipped Riposte. Paste a Reddit dunk aimed at you → get 4 reply angles that take your side. Sharp. Logical. Aggressive. Socratic. Or spar the AI in Arena first. #debate #buildinpublic

验证AI代理的决策,然后尝试篡改验证记录。
foh_quarters · HN
Verify what an AI agent did, then tamper with the record (no signup)

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

向多个AI模型提问,比较答案,观察它们辩论至共识。
u/trekhleb · Reddit
I kept pasting the same question into ChatGPT, Claude, and Gemini in three tabs; so I built a Yes-Brainer — a council of AI models, that answer your question in parallel, debate to consensus, or get judged to a verdict. submitted by /u/trekhleb to r/SideProject [link] [comments]

社区追踪AI模型体验指数,实时收集用户意见每小时更新。
schafberg · HN
Is AI Dumber Today? An index of AI model experience from user's opinion

Etch: 追踪、重放和验证AI代理的决策,提供签名审计线索。
u/Funky_Chicken_22 · Reddit
OSS to SaaS positioning problem: when the user persona and the buyer persona are completely disjoint Founder here. Sharing a positioning problem I think a lot of OSS-to-SaaS founders hit and don't talk about publicly. Context: I have been running an OSS project (world-model-mcp) with ~2,500 monthly PyPI installs. Two weeks ago I opened up the hosted companion, Etch, at etch.systems. Launched publicly on Product Hunt at 12:00 PDT yesterday. The positioning problem: OSS user persona: in