
AI Benchmark Leaderboards & Model Evals | BenchmarkList
Compare AI model benchmarks across coding, reasoning, agents, and multiple evaluation domains.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
The full gallery
Tech stack
23 projects

Compare AI model benchmarks across coding, reasoning, agents, and multiple evaluation domains.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Compare AI language models by performance across official benchmarks.
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

See your Claude Code typing ranked on a live leaderboard with daily, weekly, and weekend scores.
@justdohank · X
— a live leaderboard of Claude Code usage. I've always wondered how much the builders actually making money prompt every day. If you use CC, come claim your spot 👀

Measure code review performance across GitHub teams with leaderboards and reviewer analytics.
u/SnooStrawberries827 · Reddit
my team had 47 open PRs and nobody was reviewing them, so I gamified it our team hit 47 open PRs at one point last month and nobody was reviewing them. tried slack reminders, deadlines, rotating reviewers, none of it really stuck. might be related to the fact that everyone's hyped about how fast AI can write code now, copilot cranking out entire features in hours, but none of that matters if the PR just sits there for a week. feels like writing code stopped being the bottleneck a while back

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

Browse community-ranked products, services, and companies rated by real people.
@durinthegreat · X
Building and launched - let’s connect!

Participate in community predictions; let time verify your guesses and track your rank.
geniushui · V2EX
我的第 4 个 Vibe coding 项目——猜一猜 网址: https://cai1cai.com/ 猜一猜,让时间验证你的判断 花了 2 天时间整出来的,请各位大佬给点建议!

A thought-provoking sewing machine game with a competitive leaderboard.
dkhcyx · V2EX
欢迎大家体验我做的发人深省,劝人向善的精品游戏 https://cassiangroup.uk/sewing/ https://cassiangroup.uk/salvate/ 最近 codex 重置太多了,蹬不完的 token 拿来蹬游戏了 期待大家荣登排行榜

Organize tasks with kanban boards, notes, and calendar on web or Android.
@ak14053 · X

Cross-post to 9 platforms from one dashboard with live Viral Score and AI Hook Generator.
@D_Frank_88 · X
SyncPost — cross-post to 9 platforms from one dashboard, with a live Viral Score climbing as you type and a built-in Hook Generator — Neo (optional multi model AI sidekick if you want one.)

Create polls, surveys, forms, and live stages with AI verification and reputation-based scoring.
@Eli_Greenfeld · X

Discover Reddit trends, research keywords, and grade posts before publishing.
u/hachishikiga · Reddit
You don't need LLMs for everything "Just ask (insert any LLM)" is crazy to me. I'm sure you've also seen it, if you go around the saas/entrepreneur subreddit space the most obvious tell is how EVERY homepage looks. Recently the trend has been terminal green with slight glowing elements, and slowly revealing features as you scroll down. This is a bigger issue than the scope of this post, but I will still plant the thought in your head: if everything comes from the same source (even more so if