
Openbenchmarks for Agents
Compare and benchmark SaaS APIs with verified data to decide whether to build or buy.
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs
The full gallery
Tech stack
24 projects

Compare and benchmark SaaS APIs with verified data to decide whether to build or buy.
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

Submit your website for daily speed benchmarking and ranking on a public leaderboard.
@thefastestweb · X
daily speed monitoring for indie sites. Submit your URL, get ranked on a public leaderboard, and know the moment your performance drops.

Audit where your brand appears in major AI search engines and identify visibility gaps.
@kylekane · X

Compare AI language models by performance across official benchmarks.
fcten · V2EX
做了一个大模型 leaderboard 网站 最近一个月 CodeX 疯狂送重置,token 根本用不完,顺手做点东西。 地址:[知行录]( https://leaderboard.cn/) 排行依据主要为模型官方基准测试成绩。非主观排名。 数据会持续更新。如果有点用,欢迎各位 v 友收藏~

Publish tasks to evaluate and benchmark different AI agents and tools on a leaderboard.
u/Ruqii-ruqii · Reddit
I built an open Eval to compare different AI agents/tools/pipelines and find which solution works the best (not very pretty╥﹏╥, but practical) The original reason I built it was because I wanted to find a good PDF parser. Every PDF parser claims to be the best, but none of them can get my PDF 100% correct. They would either miss numbers or hallucinate some. Or they get PDF A and B correct but failed at C. Or get C correct but failed at A and B. Very frustrating. So I create

Analyze your TikTok videos against baseline and niche benchmarks to identify winning formats.
@AnthonyCasauria · X
no human touched these — ViralVault's blog just shipped 2 more posts on autopilot. imaged, quality-checked, and published by the pipeline itself. 100-article backlog, chipping away. check it out → @viralvaultapp #buildinpublic #solofounder

Inspect and remove hidden Unicode artifacts in AI-generated text without altering visible content.
u/nategdd · Reddit
I published reproducible fixtures for a lossless AI text artifact scanner I built AI Text Watermark Remover to inspect copied AI text without rewriting visible words. It reports exact hidden Unicode code points, removes only supported literal artifacts locally, and does not claim that hidden characters prove AI authorship. I just published the browser compatibility fixtures, artifact coverage benchmark, self-hosted API, Docker image, and open-source scanner so the claims can be tested inste

Play strategic games against AI models and see how different LLMs rank on an objective leaderboard.
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

View LLM model rankings across 10 benchmark questions.
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol

Benchmark version-control systems and coding agents on realistic development tasks.
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler

Analyze Python code across 14 quality dimensions to detect violations and measure capabilities.
@KSFirasa · X
Hello! I built a tool that profiles code (python only atm) across 14 dimensions detecting violations and capabilities outputting a full report. A bit more nuanced than "AI-powered insights". Free while in beta. Thank you!

A public leaderboard for websites and X profiles. No algorithm, no votes. Your rank is what you paid.
@alok8feb · X