
AI Benchmark Leaderboards & Model Evals | BenchmarkList
Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
The full gallery
Tech stack
26 projects

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Compare today's leading AI models by price, intelligence and more.
@spectragai · X

Browse e-commerce SaaS tools and identify which can be rebuilt with AI in one prompt.
@andreasacca · X
Inspired by @robj3d3’s that went viral last week. Over the past year working with ecommerce brands I kept hearing: “I only use 10% of this SaaS and I’m paying hundreds for it.” So I built the Ecommerce focused version: 200+ tools analysed. Honest verdicts on what you can realistically vibecode. Vibecoded entirely with @Lovable + some fixes, starting from the prompt taken from First detailed analyses coming soon. Reply with the tool you pay for every month. Thanks @robj3d3 for the inspiration! Loving X

Finally. Everything in one place
@LumetaAI · X
Best all-in-one suite of AI Image/Video/Audio tools in a unified easy-to-use platform and app

Compare AI language models by performance across official benchmarks.
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

Find AI models optimized for your hardware with performance and pricing estimates.
cdnsteve · HN
Tokenstead, find AI models for your hardware

Blogr.ai is an AI Blogger Tool that researches, plans, writes, publishes, and continuously optimizes SEO content using live SERP analysis, topical authority mapping, and Google Sea
@karakhanyanS · X
- AI Blogger Tool for SaaS, Ecom & Local SEO

Analyzes AI prompts to identify failures and return corrected versions.
u/pulptaken · Reddit
I built a website to audit AI prompts I'm giving away a limited number of early access invites for anyone who wants to try Marzel. You'll also be able to try the product once directly from the landing page, without joining the early access. This initial phase is focused on validating the core features and collecting feedback. After that, Marzel will also include a browser extension and a desktop app for auditing Claude Code and Codex CLI sessions, which is the project's main goal. If yo

Design, validate, and compare AI agent deployments with governance controls.
@paulrodturner · X

AI-powered retrospectives for IT teams that connect to your tools and identify improvement patterns.
@akadhanu · X
an ai sprint retrospectives for it teams (

Compare AI model coverage, pricing, uptime, and latency across different AI relays.
zizheruan · HN
XTokenChecker – Verifies model identities of your AI gateway