
知行录 · leaderboard.cn
在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~
完整作品展
技术栈
15 projects

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

对版本控制系统和编码代理进行性能基准测试。
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler

使用已验证数据对比和基准测试 SaaS API,决定是否构建或购买。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

跨14个维度分析Python代码,检测违规并提供详细报告。
@KSFirasa · X
Hello! I built a tool that profiles code (python only atm) across 14 dimensions detecting violations and capabilities outputting a full report. A bit more nuanced than "AI-powered insights". Free while in beta. Thank you!

提交网站获得每日速度排名和性能监控。
@thefastestweb · X
daily speed monitoring for indie sites. Submit your URL, get ranked on a public leaderboard, and know the moment your performance drops.

检查并移除 AI 生成文本中的隐藏 Unicode 工件,无需修改可见内容。
u/nategdd · Reddit
I published reproducible fixtures for a lossless AI text artifact scanner I built AI Text Watermark Remover to inspect copied AI text without rewriting visible words. It reports exact hidden Unicode code points, removes only supported literal artifacts locally, and does not claim that hidden characters prove AI authorship. I just published the browser compatibility fixtures, artifact coverage benchmark, self-hosted API, Docker image, and open-source scanner so the claims can be tested inste

识别阻碍你的B2B AI或SaaS初创企业发展的最大瓶颈。
@FounderUnstuck · X
If you’re building a startup and want to identify your biggest bottleneck, try the free assessment: 🔗 We’d love to hear if the results match your experience.

与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

发布任务来评估不同的AI代理和工具,用排行榜找出最佳方案。
u/Ruqii-ruqii · Reddit
I built an open Eval to compare different AI agents/tools/pipelines and find which solution works the best (not very pretty╥﹏╥, but practical) The original reason I built it was because I wanted to find a good PDF parser. Every PDF parser claims to be the best, but none of them can get my PDF 100% correct. They would either miss numbers or hallucinate some. Or they get PDF A and B correct but failed at C. Or get C correct but failed at A and B. Very frustrating. So I create

CoBro 用 AI 扫描竞争对手和市场数据,90 秒内判断初创企业创意是否值得构建。
@Ebrahim_Rio · X
Most founders skip validation and pray. I automated the "worth building?" check. AI scans competitors, Reddit, and market data → Cook or Kill in 90 seconds. Killed? It surfaces the pivot the data actually backs.

对比700+ Azure虚拟机在各地区的实时定价。
@honitec · X
做了一个每天自动刷新的模型定价 API 替代 LiteLLM,每天拉最新的推理定价 为什么需要「每天」? 因为 AI 模型价格一周一变 3月 $8/M → 6 月 $3/M,半年跌了一半多 半年前到处喊 GPU 短缺 现在全员降价抢客户 估值故事讲的是稀缺 现实是价格比成本跌得快 这中间的落差,就是泡沫

每周自动追踪竞争对手的定价、功能和消息变化。
@Caden1Fenn · X
Intel Brief — tracks what your competitors change (pricing, features, positioning) and delivers it to your inbox weekly. Built for indie SaaS founders.