
知行录 · leaderboard.cn
在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~
完整作品展
技术栈
16 projects

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

CoBro 用 AI 扫描竞争对手和市场数据,90 秒内判断初创企业创意是否值得构建。
@Ebrahim_Rio · X
Most founders skip validation and pray. I automated the "worth building?" check. AI scans competitors, Reddit, and market data → Cook or Kill in 90 seconds. Killed? It surfaces the pivot the data actually backs.

使用已验证的数据对比和基准测试 SaaS API,决定是否自建或购买。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

每日阅读AI精选新闻摘要,关注全球发展。
@ChenglongW98225 · X
做了一个AI新闻日报,感兴趣的可以点下方链接看一下 有什么需要改进的也可以直接评论我,我都会认真回复

聊天对比 AI 模型,投票参与排行榜排名。
u/Rabus · Reddit
I got TestingModels too overcomplicated over the month it is running: looking for some feedback how to make it more useful and simpler I run a benchmark like arena.ai , but with pre-generated prompts. So far, nearly 6k people came in and like 30k comparisons has been made - which means the thing is genuinely useful for people to compare the models. The problem is the more features i started adding the more overblown and complicated UI became - like old internet explorer tab bars Old: ht

Run a $9 AI visibility audit across OpenAI, Claude, Gemini, and Grok. See where AI overlooks your brand, who appears instead, and what to fix next.
@kylekane · X

通过InferAll统一API访问207+个AI模型
TaylorM492 · HN
InferAll – One API for OpenAI, Anthropic, Google, Nvidia Nim

用 AI 创建产品路线图、管理任务并收集功能反馈。
@jimmy_harika · X
TLDR: Notion shipped my exact app that I have using in my daily workflow from last 2 years. Try here: It got a mcp that wires to your claude code and codex. Git integration is almost complete and will ship in coming days

Contexi 用 AI 简报追踪 AI 更新、产品新闻和人物提及。
@Mileson07 · X
今天Codex、Claude Code 重置了吗? 我做了一个追踪的网站,每小时追踪 Tibo、Boris Cherny 两位主理人,以及对应的官方推特账号, 快速了解到,有没有可能重置,最近是不是已经重置了 而且还能详细看到历史的重置情况,分析未来的重置可能性,过去两周真是疯狂的重置~

为人类和AI代理记录和检索决策,追踪完整来源。
@burn2delete · X

对版本控制系统和编码代理进行性能基准测试。
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler