
AI Benchmark Leaderboards & Model Evals | BenchmarkList
Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
The full gallery
Tech stack
60 projects

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Run your prompts through Claude, GPT, and Gemini to get AI-audited consensus answers.
@StevenJdotCom · X
One AI makes mistakes. Three catch each other's. AI Consensus runs your prompt through Claude, GPT and Gemini. They work it independently, then critique each other until they reach consensus — handing you an AI audited, combined answer.

Compare and run AI models for text, image, and video generation in one workspace.
AI Onekit — AI 聚合模型创作平台 - [更多介绍](https://aionekit.com/models)

Pick your favorite page and guess whether Claude or Codex created it.
rubenflamshep · HN
I built a blind taste test for Claude and Codex designs

Chat with multiple AI models to run deepresearch and coding tasks.
@Shekar77hima · X
run deepresearch, coding agents for free.

Find AI models optimized for your hardware with performance and pricing estimates.
cdnsteve · HN
Tokenstead, find AI models for your hardware

Bring decisions to a private panel of 5 AI advisors who debate live and reach a synthesized recommendation.
@yaseenvalji · X
built this in a couple hours with Claude Code on Fable 5 ultracode. any hard decision goes to a board of 5 AI advisors: they debate live, a Chair calls it, and it remembers the outcome. open source, on the Claude API. @AnthropicAI

Access 207+ AI models from different providers through a single unified API.
TaylorM492 · HN
InferAll – One API for OpenAI, Anthropic, Google, Nvidia Nim

Axon grades your AI agents’ real conversations with an LLM judge: A–F scorecards, cited evidence, fixes and FinOps.
@tech_maju · X

Take a 16-question personality test to discover which AI model matches your thinking style.
@RedeatIX · X
GPT、Claude、Gemini、Grok……如果把你的思维方式映射成一款 AI,你会是谁?🤖我做了一个娱乐向 AI 人格测试: 16 道题 · 约 3 分钟 27 个公开型号 + 4 个隐藏款 它不考你认不认识 AI,而是看你如何判断、表达、协作和行动。测完把结果发到评论区📷 #AI人格测试 #AI测试 #你是哪个AI #人格测试 #AI #人工智能 #大模型 #ChatGPT #Claude #Gemini #Grok #趣味测试 #在线测试 #3分钟测试

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

Build and refine AI prompts to improve model responses.
markquis91 · HN
Prompting Refinement Tool [requesting testing]