
Trunchbull
在浏览器中运行AI模型基准测试以检测性能回归。
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser
完整作品展
技术栈
18 projects

在浏览器中运行AI模型基准测试以检测性能回归。
pepperpoppins · HN
Trunchbull, run real models against any benchmark in your browser

对比 AI 模型在编码任务上的表现,支持成本追踪和 ELO 排名。
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

PRcade 通过团队排行榜和分析可视化GitHub代码审查性能
u/SnooStrawberries827 · Reddit
my team had 47 open PRs and nobody was reviewing them, so I gamified it our team hit 47 open PRs at one point last month and nobody was reviewing them. tried slack reminders, deadlines, rotating reviewers, none of it really stuck. might be related to the fact that everyone's hyped about how fast AI can write code now, copilot cranking out entire features in hours, but none of that matters if the PR just sits there for a week. feels like writing code stopped being the bottleneck a while back

让AI模型通过3D动画展现香蕉植物的完整生命周期来比较性能。
fran-mora · HN
I gave 5 AI coding agents one prompt: grow a banana plant through its whole life in three.js: sprout, leaves, flower, fruit, rot, then pups that restart the loop. It's deceptively simple and yet very hard to get right from procedural code: you have to write working three.js and understand how the plant is actually built; how it hangs, ages and decays. Get the biology wrong and the code renders something weird. These are agents, not bare models (Claude Code and Codex for now). They can use tools, including playwright to check their work and improve it.

Speech to text dictation and multi-engine speed benchmarking. Compare OpenAI GPT-Transcribe, Deepgram Nova-3, NVIDIA Parakeet, and Fish Audio with local IndexedDB privacy.
@alvaisy · X
finished voice to text small web app for my own itch. it's opensource. use openrotuer key. and use it with 4 models.

在OpenVibeEval中对比不同AI模型生成前端代码和可访问性评分。
u/12qwww · Reddit
I built a live benchmark to see which AI actually writes the best frontend code Hey everyone! I built OpenVibeEval because I was tired of "vibe-checking" AI-generated frontend code. I wanted to know which model actually produces the most accessible and clean React/Tailwind output. What I built: •A leaderboard of 24 models (Claude, GPT, DeepSeek, etc.) ranked by axe-core accessibility scores. •A Harness Comparator to show how different system prompts change the same model's output. •

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

实时可视化硬件在运行LLM推理时的性能指标
dev_dan_2 · HN
WatchMachineGo – A visualizer to show hardware performing LLM inference

Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · HN
Echo – Fable-level results at 1/3 the cost using open-weight models

基于证据分析比较候选人,生成招聘决策文件。
facundobon · HN
Verdict – AI hiring verdicts where every score cites the CV verbatim

在一个仪表板中收集和分析来自多个平台的反馈。
@GeorgiG26929167 · X
Remarkd lets you create a feedback page for anything—ideas, landing pages, designs, pricing, features or prototypes. Share one link anywhere and collect structured, anonymous feedback in one place, with AI-powered insights to help you spot patterns faster