
Openbenchmarks for Agents
使用已验证数据对比和基准测试 SaaS API,决定是否构建或购买。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs
完整作品展
技术栈
29 projects

使用已验证数据对比和基准测试 SaaS API,决定是否构建或购买。
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

提交网站获得每日速度排名和性能监控。
@thefastestweb · X
daily speed monitoring for indie sites. Submit your URL, get ranked on a public leaderboard, and know the moment your performance drops.

审计您的品牌在各大AI搜索引擎中的可见性并识别差距。
@kylekane · X

用0-100 AGI分数对标前沿AI模型的基准性能。
baraklaniado · HN
I audited my AI leaderboard scale – every score dropped 6-15 points

Vee Group diagnoses the real constraint stalling a founder-led business — benchmarked against its closest structural peers, not generic industry averages — then builds the fix and
@VeeGrp · X
AI-native business diagnostic platform. Business intelligence built on human behaviour.

发布任务来评估不同的AI代理和工具,用排行榜找出最佳方案。
u/Ruqii-ruqii · Reddit
I built an open Eval to compare different AI agents/tools/pipelines and find which solution works the best (not very pretty╥﹏╥, but practical) The original reason I built it was because I wanted to find a good PDF parser. Every PDF parser claims to be the best, but none of them can get my PDF 100% correct. They would either miss numbers or hallucinate some. Or they get PDF A and B correct but failed at C. Or get C correct but failed at A and B. Very frustrating. So I create

AI代理24小时监控您的投资组合,提供漂移、收益和基准追踪警报。
@SigVestAI · X

对比多个LLM API提供商的延迟和吞吐量性能。
@QAInsights · X

检查并移除 AI 生成文本中的隐藏 Unicode 工件,无需修改可见内容。
u/nategdd · Reddit
I published reproducible fixtures for a lossless AI text artifact scanner I built AI Text Watermark Remover to inspect copied AI text without rewriting visible words. It reports exact hidden Unicode code points, removes only supported literal artifacts locally, and does not claim that hidden characters prove AI authorship. I just published the browser compatibility fixtures, artifact coverage benchmark, self-hosted API, Docker image, and open-source scanner so the claims can be tested inste

与AI模型进行策略游戏,查看大语言模型在排行榜上的排名。
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

查看LLM模型在10个基准问题上的评分和排名。
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol

对版本控制系统和编码代理进行性能基准测试。
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler