
Openbenchmarks for Agents
Compare and benchmark SaaS APIs with verified data to decide whether to build or buy.
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs
The full gallery
Tech stack
18 projects

Compare and benchmark SaaS APIs with verified data to decide whether to build or buy.
fenilsuchak · HN
OpenBenchmarks – Helping agents discover and pick the right SaaS APIs

View LLM model rankings across 10 benchmark questions.
fristovic · HN
She watched me look at model rankings and asked what do the numbers mean... I literally had no good way of explaining it to her so I just came up with something that is approximately in the same ballpark as some of the benchmarks out there lol

Benchmark version-control systems and coding agents on realistic development tasks.
videlov · HN
I was interested in answering this question so I built a benchmark comparing git, jj and gitbutler in agentic context https://vcbench.dev/ Disclaimer - I am a co-founder of GitButler

Compare latency and throughput performance across LLM API providers.
@QAInsights · X

Play strategic games against AI models and see how different LLMs rank on an objective leaderboard.
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

Analyze Python code across 14 quality dimensions to detect violations and measure capabilities.
@KSFirasa · X
Hello! I built a tool that profiles code (python only atm) across 14 dimensions detecting violations and capabilities outputting a full report. A bit more nuanced than "AI-powered insights". Free while in beta. Thank you!

AI scans competitors and market data to validate if a startup idea is worth building in 90 seconds.
@Ebrahim_Rio · X
Most founders skip validation and pray. I automated the "worth building?" check. AI scans competitors, Reddit, and market data → Cook or Kill in 90 seconds. Killed? It surfaces the pivot the data actually backs.

Centralize and analyze Google Business Profile reviews with AI-powered weekly performance briefings.
@smooseo · X
For the moment only in French but soon in english : - analyse GBP’s reviews

Analytics dashboard for LLM API spending by model and environment with optimization suggestions.
ATsimbalistov · HN
Show HN: Tracking GenAI cost and endpoint fragility so app teams don't have to

Compare and use multiple large language models through a unified API interface.
u/DanTahirCode · Reddit
I built an open source coding agent with a personality - meet Klenny Code 🐾 Hey r/SideProject, my name is Dan Tahir, and I'm here to show off something I'm really proud of: Klenny Code, the open source coding agent with personality. A fully capable coding agent with memory and cross-project referencing, plus an assistant who can read your email, run scheduled tasks, pilot your browser, and be your corgi pal. Here's the pitch: bring your own OpenRouter API key, and Klenny wil

Get AI-powered investment research tools like DCF models and stock screening for individual investors.
@StephanInvests · X
Building a financial saas at the lowest price for people to put more towards their financial goals

Model and evaluate your project's quality standards with specifications, CLI, and agent skills.
craigsmitham