
AI Benchmark Leaderboards & Model Evals | BenchmarkList
Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
The full gallery
Tech stack
32 projects

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Ad-free Vedic astrology engine on the Swiss Ephemeris. Compute your sidereal chart, dasha timing and planetary strengths — then query it with AI.
@saketposwal · X

Test large model API relay stations for consistency and authenticity with official versions.
@nodeloc_cc · X
🌈 7月,你好,MODELOC上线算力池。 MODELOC自上线以来,已检测2000余次,覆盖600+中转站,为众多AI用户提供的使用参考。 MODELOC近期进行了改版,上线了算力池及市场。 加入算力池 查看帖子: 用 MODELOC 便宜地调各家大模型:一次讲清它的价格体系

Benchmark AI models by having them animate a 3D banana plant's full lifecycle.
fran-mora · HN
I gave 5 AI coding agents one prompt: grow a banana plant through its whole life in three.js: sprout, leaves, flower, fruit, rot, then pups that restart the loop. It's deceptively simple and yet very hard to get right from procedural code: you have to write working three.js and understand how the plant is actually built; how it hangs, ages and decays. Get the biology wrong and the code renders something weird. These are agents, not bare models (Claude Code and Codex for now). They can use tools, including playwright to check their work and improve it.

Generate 3D models, characters, and animations from images or text.
zaczuo · Product Hunt
V2Fun Generate 3D character with 8K textures and AI motion capture

Find AI models optimized for your hardware with performance and pricing estimates.
cdnsteve · HN
Tokenstead, find AI models for your hardware

Evaluate AI products against published agent criteria with inspectable evidence and community votes.
@katyorby · X
i built — a local receipt for claude code runs. your check says whether the workspace passes now; the transcript supplies the activity counts. no transcript upload and no magical autonomy score.

View and compare public opinions and benchmark ratings for leading AI models.
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

Run real models against benchmarks in your browser to detect performance regressions before production.
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser

Access leading AI models through one API with transparent token pricing.
DustinPham12 · HN
1endpoint – Cheaper access to AI models

Compare today's leading AI models by price, intelligence and more.
@spectragai · X

Compare how different AI models generate frontend code and view accessibility scores.
u/12qwww · Reddit
I built a live benchmark to see which AI actually writes the best frontend code Hey everyone! I built OpenVibeEval because I was tired of "vibe-checking" AI-generated frontend code. I wanted to know which model actually produces the most accessible and clean React/Tailwind output. What I built: •A leaderboard of 24 models (Claude, GPT, DeepSeek, etc.) ranked by axe-core accessibility scores. •A Harness Comparator to show how different system prompts change the same model's output. •