
完整作品展
技术栈
4 projects


对比 AI 模型在编码任务上的表现,支持成本追踪和 ELO 排名。
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

为 AI agents 提供请求人工批准的结构化方式,支持安全重试和决策验证。
@GetAgentHail · X
AgentHail — a control layer for AI agents to request approval, execute work, and return verifiable results. Would you sign up or leave?

检查 AI 推理踪迹,评估模型真实性。
malik_dixon1 · Product Hunt
TraceLogicAI: AI Architecture Evaluation Compare AI architectures with evidence, not guesswork