
Openbenchmarks for Agents
使用已验证数据对比和基准测试 SaaS API,决定是否构建或购买。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs
完整作品展
技术栈
24 projects

使用已验证数据对比和基准测试 SaaS API,决定是否构建或购买。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

提交网站获得每日速度排名和性能监控。
@thefastestweb · X
daily speed monitoring for indie sites. Submit your URL, get ranked on a public leaderboard, and know the moment your performance drops.

在 leaderboard 上按官方基准对比 AI 大模型的性能排名
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

发布任务来评估不同的AI代理和工具,用排行榜找出最佳方案。
u/Ruqii-ruqii · Reddit
I built an open Eval to compare different AI agents/tools/pipelines and find which solution works the best (not very pretty╥﹏╥, but practical) The original reason I built it was because I wanted to find a good PDF parser. Every PDF parser claims to be the best, but none of them can get my PDF 100% correct. They would either miss numbers or hallucinate some. Or they get PDF A and B correct but failed at C. Or get C correct but failed at A and B. Very frustrating. So I create

检查并移除 AI 生成文本中的隐藏 Unicode 工件,无需修改可见内容。
u/nategdd · Reddit
I published reproducible fixtures for a lossless AI text artifact scanner I built AI Text Watermark Remover to inspect copied AI text without rewriting visible words. It reports exact hidden Unicode code points, removes only supported literal artifacts locally, and does not claim that hidden characters prove AI authorship. I just published the browser compatibility fixtures, artifact coverage benchmark, self-hosted API, Docker image, and open-source scanner so the claims can be tested inste

与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

查看LLM模型在10个基准问题上的评分和排名。
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol

对版本控制系统和编码代理进行性能基准测试。
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler

分析TikTok视频在你的基线和细分市场中的表现,发现爆款内容格式。
@AnthonyCasauria · X
no human touched these — ViralVault's blog just shipped 2 more posts on autopilot. imaged, quality-checked, and published by the pipeline itself. 100-article backlog, chipping away. check it out → @viralvaultapp #buildinpublic #solofounder

审计您的品牌在各大AI搜索引擎中的可见性并识别差距。
@kylekane · X

跨14个维度分析Python代码,检测违规并提供详细报告。
@KSFirasa · X
Hello! I built a tool that profiles code (python only atm) across 14 dimensions detecting violations and capabilities outputting a full report. A bit more nuanced than "AI-powered insights". Free while in beta. Thank you!

A public leaderboard for websites and X profiles. No algorithm, no votes. Your rank is what you paid.
@alok8feb · X