
Trunchbull
Run real models against benchmarks in your browser to detect performance regressions before production.
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser
The full gallery
Tech stack
18 projects

Run real models against benchmarks in your browser to detect performance regressions before production.
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

Measure code review performance across GitHub teams with leaderboards and reviewer analytics.
u/SnooStrawberries827 · Reddit
my team had 47 open PRs and nobody was reviewing them, so I gamified it our team hit 47 open PRs at one point last month and nobody was reviewing them. tried slack reminders, deadlines, rotating reviewers, none of it really stuck. might be related to the fact that everyone's hyped about how fast AI can write code now, copilot cranking out entire features in hours, but none of that matters if the PR just sits there for a week. feels like writing code stopped being the bottleneck a while back

Benchmark AI models by having them animate a 3D banana plant's full lifecycle.
fran-mora · HN
I gave 5 AI coding agents one prompt: grow a banana plant through its whole life in three.js: sprout, leaves, flower, fruit, rot, then pups that restart the loop. It's deceptively simple and yet very hard to get right from procedural code: you have to write working three.js and understand how the plant is actually built; how it hangs, ages and decays. Get the biology wrong and the code renders something weird. These are agents, not bare models (Claude Code and Codex for now). They can use tools, including playwright to check their work and improve it.

Speech to text dictation and multi-engine speed benchmarking. Compare OpenAI GPT-Transcribe, Deepgram Nova-3, NVIDIA Parakeet, and Fish Audio with local IndexedDB privacy.
@alvaisy · X
finished voice to text small web app for my own itch. it's opensource. use openrotuer key. and use it with 4 models.

Compare how different AI models generate frontend code and view accessibility scores.
u/12qwww · Reddit
I built a live benchmark to see which AI actually writes the best frontend code Hey everyone! I built OpenVibeEval because I was tired of "vibe-checking" AI-generated frontend code. I wanted to know which model actually produces the most accessible and clean React/Tailwind output. What I built: •A leaderboard of 24 models (Claude, GPT, DeepSeek, etc.) ranked by axe-core accessibility scores. •A Harness Comparator to show how different system prompts change the same model's output. •

Compare AI language models by performance across official benchmarks.
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Visualize hardware performance metrics while running LLM inference on your system.
dev_dan_2 · HN
WatchMachineGo – A visualizer to show hardware performing LLM inference

Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · HN
Echo – Fable-level results at 1/3 the cost using open-weight models

Compare candidates on evidence-cited dimensions and generate hiring documents.
facundobon · HN
Verdict – AI hiring verdicts where every score cites the CV verbatim

Collect and analyze feedback from multiple platforms in a single dashboard.
@GeorgiG26929167 · X
Remarkd lets you create a feedback page for anything—ideas, landing pages, designs, pricing, features or prototypes. Share one link anywhere and collect structured, anonymous feedback in one place, with AI-powered insights to help you spot patterns faster