
AI Judge
One bundle. Three independent judges. Reproducible rankings.
@Imranmohsin18 · X
sites I’m working on: Please check out and give me feedback all are public showcasing my skills :D ( not selling anything atm )
完整作品展
技术栈
17 projects

One bundle. Three independent judges. Reproducible rankings.
@Imranmohsin18 · X
sites I’m working on: Please check out and give me feedback all are public showcasing my skills :D ( not selling anything atm )

根据已发布的代理标准评判AI产品,包含可检查证据和社区投票。
@katyorby · X
i built — a local receipt for claude code runs. your check says whether the workspace passes now; the transcript supplies the activity counts. no transcript upload and no magical autonomy score.

社区追踪AI模型体验指数,实时收集用户意见每小时更新。
schafberg · HN
Is AI Dumber Today? An index of AI model experience from user's opinion

与朋友一起画秘密提示,让AI评判和吐槽你的绘画。
u/Mental_Training4132 · Reddit
Doodle Judge — built a game where an AI roasts your drawings, solo project Everyone in the room draws the same secret prompt under a timer. An AI (Claude Haiku) judges every drawing on a different criteria each round. The AI might roast you if you're not the best doodler. Built it solo. React/Tailwind frontend, FastAPI backend, deployed on Railway. Play with friends via a room code, or try it solo first to see how it works. No sign-up beyond picking a name. Would love any feedback or im

Riposte 是生成辩论回复、分析论证的AI工具包,提供七种专业模式。
@recurno · X
Just shipped Riposte. Paste a Reddit dunk aimed at you → get 4 reply angles that take your side. Sharp. Logical. Aggressive. Socratic. Or spar the AI in Arena first. #debate #buildinpublic

用 AI 分析面部吸引力,获得多维度美貌评分。
FaceRatingAI — AI 颜值打分网站,包括不同维度评分、整体排名等

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

IntelligenceMax: 解答自适应AI生成的谜题来训练推理能力,从错误中学习。
u/connerpro · Reddit
IntelligenceMax - Adaptive reasoning practice with live AI questions (claim-safe near vs far) submitted by /u/connerpro to r/SideProject [link] [comments]

分析网站是 AI 生成还是手工编码,评分 0-100。
@mukparekh · X
I built tool for fun. roast website Free. No signup

向多个AI模型提问,比较答案,观察它们辩论至共识。
u/trekhleb · Reddit
I kept pasting the same question into ChatGPT, Claude, and Gemini in three tabs; so I built a Yes-Brainer — a council of AI models, that answer your question in parallel, debate to consensus, or get judged to a verdict. submitted by /u/trekhleb to r/SideProject [link] [comments]

对比 AI 模型在编码任务上的表现,支持成本追踪和 ELO 排名。
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

Aurora - 本地AI代理验证系统,提供透明的推理过程和MCP集成。
brandon_grutkowski · Product Hunt
Aurora Glass-box Quantitative AI for Humans and Agents