
AI Benchmark Leaderboards & Model Evals | BenchmarkList
Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
The full gallery
Tech stack
19 projects

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Create AI-powered product roadmaps, manage tasks, and collect feature feedback.
@jimmy_harika · X
TLDR: Notion shipped my exact app that I have using in my daily workflow from last 2 years. Try here: It got a mcp that wires to your claude code and codex. Git integration is almost complete and will ship in coming days

Ingest, explore, and analyze datasets with autonomous data processing in an interactive workspace.
@kashyap_ai · X

Privacy-first analytics platform with AI copilot for SEO content research and publishing.
@VertCodeEU · X

Analyze sensitive data with AI while keeping it encrypted end-to-end with client-side keys.
@JackiePeters · X

Query spreadsheets and datasets with plain-English questions to generate instant answers and reports.
u/maybeImakemoney · Reddit
I built the thing. Now I am not sure the base use case is one people will pay for. Founder here. This started as a side learning project to see whether an LLM could answer questions about Excel data, back when they could not do it well. I built the first version on n8n, with workflows that ingested files, generated metadata with an LLM, and answered questions against the converted data plus that metadata. Then I started using it for my own analysis and report generation, saw that the time sav

Track and analyze product user events with an interactive analytics dashboard.
@kitbasedev · X
We just launched, check us out at

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

Analyze your product to diagnose market fit and uncover growth bottlenecks.
@teslaptimus · X

Create and launch AI services with integrated payments and analytics.
EasyLaunch — 能力产品化平台,把专家方法论和 AI Skill 一键上线成带登录、收款与数据统计的可售服务

AI-powered virtual data room that transforms documents into guided experiences for deal teams.
@Puneeeeeeet · X
We help brands go viral on X with organic video campaigns that people actually want to watch. wanna try for

Upload a CAS PDF to get AI portfolio analysis with benchmark comparisons and allocation insights.
@iASHeeesh · X