
AI Benchmark Leaderboards & Model Evals | BenchmarkList
对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
完整作品展
技术栈
60 projects

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Choicease帮助你用AI逐步分析复杂决策,理清权衡并获得建议。
@arunkumarkundra · X
I vibe coded this app: My experience has been mixed. It is an iterative process, and you need to push back. But over time, I have seen Claude get better. If not for AI, I couldn't have developed the app in the first place, so I am thankful 🙂

用AI验证创业想法,提供市场研究、竞争分析和可行性评分。
@xpertvex · X
Building A platform to validate your startup idea with AI before you build it.

追踪目标和失败的LLM智能体,跨会话保存状态并展示每步推理。
u/OGMYT · Reddit
I built LOLM, a lower-cost LLM agent that shows what it actually did — looking for blunt feedback I’m one of the founders/builders behind LOLM. Most AI products show an answer but hide whether the system retrieved anything useful, verified the result, switched models, hit a limit, or simply stopped. LOLM exposes those parts through controller events and run receipts. It includes: - Live agent - CLI - Coding and small app-building workflows - Memory and self-hosting options - Control decis

Independent model readings, explicit disagreement, human decision gates, and verifiable research receipts.
@chatmouthai · X

测量您的品牌在ChatGPT、Claude、Gemini和Perplexity中的AI可见度。
@RHMDigitalSeo · X

Deciding is the part you can’t delegate. Capture from anywhere, including inside your assistant, and keep one list on web, Windows, Mac, and iPhone.
u/Mean-Papaya3532 · Reddit
Done Bear - a local-first task manager with MCP, an API and a CLI Done Bear is a task manager with the five GTD lists, Inbox through Someday, and not much else. https://donebear.com Your AI assistant can use it. A hosted MCP server lets Claude or ChatGPT read your tasks, add them and tick them off, with your permission and nothing running locally. There's a GraphQL API and a CLI too, if you'd rather script it. It runs on the web, Mac, Windows, Linux and iPhone. Tasks live on your device

Etch: 追踪、重放和验证AI代理的决策,提供签名审计线索。
u/Funky_Chicken_22 · Reddit
OSS to SaaS positioning problem: when the user persona and the buyer persona are completely disjoint Founder here. Sharing a positioning problem I think a lot of OSS-to-SaaS founders hit and don't talk about publicly. Context: I have been running an OSS project (world-model-mcp) with ~2,500 monthly PyPI installs. Two weeks ago I opened up the hosted companion, Etch, at etch.systems. Launched publicly on Product Hunt at 12:00 PDT yesterday. The positioning problem: OSS user persona: in

通过5层代码安全测验证明你对编程的理解。
u/rontop151 · Reddit
I vibe-coded an anti-vibe-coding app. Yes. I used AI to build a whole academy that trains you to stop blindly shipping AI code. Five levels. Security quizzes. Badges. A streak. The whole “read the error before you panic” curriculum. It’s like building a gym while eating pizza on the treadmill. Play it here → https://videcodingacademy.web.app This post was also vibe-coded. Happy Vibe Coding 😂 submitted by /u/rontop151 to r/SideProject [link] [comments]

验证AI代理的决策,然后尝试篡改验证记录。
foh_quarters · HN
Verify what an AI agent did, then tamper with the record (no signup)

向多个AI模型提问,观看它们辩论获得答案。
@ZUKOWEB3166 · X
Vibe coding is more interesting when you use Truth AI

Riposte 是生成辩论回复、分析论证的AI工具包,提供七种专业模式。
@recurno · X
Just shipped Riposte. Paste a Reddit dunk aimed at you → get 4 reply angles that take your side. Sharp. Logical. Aggressive. Socratic. Or spar the AI in Arena first. #debate #buildinpublic