
AI Benchmark Leaderboards & Model Evals | BenchmarkList
Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities
The full gallery
Tech stack
32 projects

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

AI agents that scan web apps and APIs for vulnerabilities through a chat interface.
@SableOffensive · X
Building Sable. An AaaS where specialized AI security agents perform penetration testing for web apps and APIs, helping teams find vulnerabilities before they reach production.

Build and deploy AI agents with persistent threads, webhooks, and scheduled tasks.
@computer_agents · X

An AI agent that analyzes code repositories and writes code autonomously in your browser.
@vlipadev · X

Write objectives and let AI agents (Claude, Cursor, Codex) decompose and execute coding tasks.
dudemanAtl · HN
PlanWright – A control plane for AI coding agents

Assign tasks to an AI team lead that dispatches AI specialists to complete work.
@StevenCen75554 · X
我们最新推出的 Product Hunt 上线了! 专为电商团队而生,让你只需一句话,就能拥有一整支 AI 团队——写SEO Blog、做用户调研、优化Listing、生成AI短视频,全都不在话下。不用招人,不用在一堆工具间来回切换,只需要说出你想要什么。 专为想要「一个人打出一支团队的仗」的电商运营团队打造。 如果这个理念打动了你,今天的一个 upvote 对我们来说意义重大 🙏 #ProductHunt #BuildInPublic #AIAgents #IndieHackers

Assign and coordinate work for AI agents using a prioritized task board with dependencies.
Olscore · HN
Pullboard – a work queue for agents, built to run a quant desk

Search and discover AI Skills by professional scenario, save collections, and install into Claude Code with one click.
SkillForge — Claude Skill 发现与分发平台,按职业场景组织 5700+ skill 覆盖 30 个垂直领域,一行命令装到 Claude Code / Cursor,登录后可留存自己的工具集