
AI Benchmark Leaderboards & Model Evals | BenchmarkList
对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
完整作品展
技术栈
26 projects

对比和评估 AI 模型在编码、推理、代理和其他基准测试中的表现。
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

Compare today's leading AI models by price, intelligence and more.
@spectragai · X

浏览220多个电商工具,识别哪些可以用AI重构。
@andreasacca · X
Inspired by @robj3d3’s that went viral last week. Over the past year working with ecommerce brands I kept hearing: “I only use 10% of this SaaS and I’m paying hundreds for it.” So I built the Ecommerce focused version: 200+ tools analysed. Honest verdicts on what you can realistically vibecode. Vibecoded entirely with @Lovable + some fixes, starting from the prompt taken from First detailed analyses coming soon. Reply with the tool you pay for every month. Thanks @robj3d3 for the inspiration! Loving X

Finally. Everything in one place
@LumetaAI · X
Best all-in-one suite of AI Image/Video/Audio tools in a unified easy-to-use platform and app

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

对比 AI 模型在编码任务上的表现,支持成本追踪和 ELO 排名。
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

查找与您硬件兼容的AI模型并查看性能和价格估计。
cdnsteve · HN
Tokenstead, find AI models for your hardware

Blogr.ai is an AI Blogger Tool that researches, plans, writes, publishes, and continuously optimizes SEO content using live SERP analysis, topical authority mapping, and Google Sea
@karakhanyanS · X
- AI Blogger Tool for SaaS, Ecom & Local SEO

分析 AI 提示词,识别问题并返回改进版本。
u/pulptaken · Reddit
I built a website to audit AI prompts I'm giving away a limited number of early access invites for anyone who wants to try Marzel. You'll also be able to try the product once directly from the landing page, without joining the early access. This initial phase is focused on validating the core features and collecting feedback. After that, Marzel will also include a browser extension and a desktop app for auditing Claude Code and Codex CLI sessions, which is the project's main goal. If yo

ClawMaven — AI 代理部署治理层,支持设计和验证
@paulrodturner · X

AI驱动的IT团队回顾工具,连接各类工具并识别改进模式。
@akadhanu · X
an ai sprint retrospectives for it teams (

XTokenChecker是一个AI模型目录,用于对比不同AI网关上的覆盖率、定价和延迟。
zizheruan · HN
XTokenChecker – Verifies model identities of your AI gateway