
100 Questions — AI Visibility Audit & Benchmark Tool
Audit where your brand appears in major AI search engines and identify visibility gaps.
@kylekane · X
The full gallery
Tech stack
61 projects

Audit where your brand appears in major AI search engines and identify visibility gaps.
@kylekane · X

Benchmark local coding models on consumer hardware to measure accuracy, latency, and throughput across 27 tasks.
u/Unfair_Association89 · Reddit
I built a reproducible benchmark for local coding models (Ollama, 27 tasks, live leaderboard) ran it on my 8GB card, here's what I found I kept eyeballing "vibes" to decide whether one quant of a coding model was actually better than another on my machine, so I built Sakura to get real numbers instead. What it does: - Points at any Ollama model and runs it through 27 hand-curated tasks: codegen, bugfix, SQL, refactor, systems design, protocol implementation, and terminal-agent episode

Measure your GPU's real memory-bandwidth ceiling for local AI in 30 seconds.
Ar5en1c · HN
Headroom – measure your GPU's true bandwidth ceiling for local AI

A transparent, research-backed benchmark for startup ideas and websites.
@xmangonic · X

Discover which local AI models your machine can run with verified benchmarks.
@Carl0sFelipe · X
Just shipped — a tool that helps you discover which local AI models actually run on your hardware, with community benchmarks, quantization support, and estimated speed. Building in public from here. #BuildingPublic #AIDevelopment #rust #benchmaks #aimodel

Submit kernel patches and engine optimizations for LLM inference speed, benchmarked on dedicated hardware.
carsenk · HN
Frontier.fast – Help push the frontier of LLM speed forward

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

View and compare public opinions and benchmark ratings for leading AI models.
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

Compare AI language models by performance across official benchmarks.
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

Run real models against benchmarks in your browser to detect performance regressions before production.
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser

Analyze Python code across 14 quality dimensions to detect violations and measure capabilities.
@KSFirasa · X
Hello! I built a tool that profiles code (python only atm) across 14 dimensions detecting violations and capabilities outputting a full report. A bit more nuanced than "AI-powered insights". Free while in beta. Thank you!

Compare latency and throughput performance across LLM API providers.
@QAInsights · X