
知行录 · leaderboard.cn
在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~
完整作品展
技术栈
27 projects

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

对比 AI 模型在编码任务上的表现,支持成本追踪和 ELO 排名。
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

用多个AI模型同时翻译并对比结果。
orion1 · V2EX
vibe conding 撸了两个工具。 都是自己经常再用的 1 ,翻译站,模仿的 deepl ,主要是翻译网站只给一个翻译,有的时候未必是自己想要的那个,这个能调用多个模型都翻译一遍。 https://translate.v8ce.com/ 2 ,音视频处理工具,不上传数据,只调用浏览器能力。有的时候想简单裁切一下视频,压缩或者加速,打开大型处理软件有点太重了。 https://video-tools.v8ce.com/

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

发现你的电脑能运行的本地 AI 模型,包含已验证的基准数据。
@Carl0sFelipe · X
Just shipped — a tool that helps you discover which local AI models actually run on your hardware, with community benchmarks, quantization support, and estimated speed. Building in public from here. #BuildingPublic #AIDevelopment #rust #benchmaks #aimodel

查看和对比主流AI模型的公众意见和基准评分。
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

对比语言模型在 Redactle 谜题上的表现排名。
pampas · HN
Redactle LLM Leaderboard

在OpenVibeEval中对比不同AI模型生成前端代码和可访问性评分。
u/12qwww · Reddit
I built a live benchmark to see which AI actually writes the best frontend code Hey everyone! I built OpenVibeEval because I was tired of "vibe-checking" AI-generated frontend code. I wanted to know which model actually produces the most accessible and clean React/Tailwind output. What I built: •A leaderboard of 24 models (Claude, GPT, DeepSeek, etc.) ranked by axe-core accessibility scores. •A Harness Comparator to show how different system prompts change the same model's output. •

用你的数据微调定制语言模型,支持密码学删除证明。
@MBrew26730 · X
Dataset cleaning + fine tuning + continual learning at

比较多个语音转文字引擎的速度和准确度,支持本地隐私保护。
@alvaisy · X
finished voice to text small web app for my own itch. it's opensource. use openrotuer key. and use it with 4 models.

构建或自动生成具有逻辑、支付、签名和集成等功能的多语言表单。
@hkbonur · X
that digitalise paper forms and make everything multi language, so businesses can get better conversion. One form any language.

可视化语言模型在各层回答前的思考内容。
ada1981 · HN
I built a web tool to see and edit what an AI thinks before it answers