
AI Benchmark Leaderboards & Model Evals | BenchmarkList
Compare AI model benchmarks across coding, reasoning, agents, and multiple evaluation domains.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
The full gallery
Tech stack
60 projects

Compare AI model benchmarks across coding, reasoning, agents, and multiple evaluation domains.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Run a free SEO audit and track your visibility in Google and AI search engines.
@Sathibuilds

AI agent that audits landing pages, tests CTAs, and recommends conversion optimizations.
@raihankhan_rk · X

Analyze how your brand appears in AI search engines like SGE, ChatGPT, and Perplexity.
u/No_Pangolin3678 · Reddit
We built AiVisis to measure brand visibility in AI search — looking for honest feedback Hi r/SideProject , I’m part of the team building AiVisis, a GEO and AI visibility analysis platform for businesses, marketers, and agencies. Most companies already track their search rankings, website traffic, and social media performance. However, many brands still cannot clearly answer a newer question: How visible is our brand across AI-powered search and answer platforms? AiVisis analyzes a we

Interactive AI audit funnel for analyzing systems with multiple visual approaches.
@jozef_x1 · X
🚨 I Vibe Coded F0ur (4) different versions of @coreyganim's "Ai Audit" Funnel, which is your favourite? VERSION #3:

Analyze why your product is invisible to AI search and get code fixes to improve visibility.
@dreamingfounder · X
Tell me your thoughts on , we let you learn how to appear on AI recommendations

AI scans competitors and market data to validate if a startup idea is worth building in 90 seconds.
@Ebrahim_Rio · X
Most founders skip validation and pray. I automated the "worth building?" check. AI scans competitors, Reddit, and market data → Cook or Kill in 90 seconds. Killed? It surfaces the pivot the data actually backs.

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

Add evaluation reports to your AI agent with a shareable URL that scores performance.
adeeonline · HN
AgentsProof – a small project for testing AI agents

AI platform to ask questions, generate images, and speak hands-free for better decision-making.
@Akinzoooo · X

Play strategic games against AI models and see how different LLMs rank on an objective leaderboard.
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

Score your landing page copy for AI-generated tone and get improvement suggestions.
parweb · HN
Detecting AI-written copy without an LLM (deterministic, client-side)