
System 2 Arena - Objective AI Strategy Benchmarks
Play strategic games against AI models and see how different LLMs rank on an objective leaderboard.
masterchef2209 · HN
I created a platform to check which AI models is the best gamer
The full gallery
Tech stack
25 projects

Play strategic games against AI models and see how different LLMs rank on an objective leaderboard.
masterchef2209 · HN
I created a platform to check which AI models is the best gamer

Chat with AI models, compare them, and vote to shape a community leaderboard.
u/Rabus · Reddit
I got TestingModels too overcomplicated over the month it is running: looking for some feedback how to make it more useful and simpler I run a benchmark like arena.ai , but with pre-generated prompts. So far, nearly 6k people came in and like 30k comparisons has been made - which means the thing is genuinely useful for people to compare the models. The problem is the more features i started adding the more overblown and complicated UI became - like old internet explorer tab bars Old: ht

Scores AI-generated ad creative and returns verdicts (run/fix/kill) via MCP and REST API.
ds246 · HN
Spendict – a performance marketer's verdict for AI agents, over MCP

Write a bot and watch it compete against others in real-time battles.
@hoofader · X
Get ready for robots age:

Play a multiplayer strategy game where an AI serves as the game master.
jklewis · HN
Machinations – a multiplayer strategy game where LLM is the game master

Log and retrieve decisions with complete provenance for humans and AI agents.
@burn2delete · X

Practice interviews, negotiations and high-stakes conversations with an AI roleplay partner.
@auditormusic19 · X
Building iGrow, an AI roleplay app to practice high stakes conversations before they happen. Check it here:

Chat with multiple AI models to run deepresearch and coding tasks.
@Shekar77hima · X
run deepresearch, coding agents for free.

Give your AI product a sense of user taste through preference learning.
@marcellafjacob · X

PvP trading battles where AI agents compete for rewards.
@CreatorBid · X
You can make money vibe-coding also on

Bring decisions to a private panel of 5 AI advisors who debate live and reach a synthesized recommendation.
@yaseenvalji · X
built this in a couple hours with Claude Code on Fable 5 ultracode. any hard decision goes to a board of 5 AI advisors: they debate live, a Chair calls it, and it remembers the outcome. open source, on the Claude API. @AnthropicAI

Scan websites for AI usage and test chatbots for bias and safety with this open-source EU AI Act compliance toolkit.
@sebuzdugan · X
building @getfrai , an open source toolkit that helps ML engineers navigate EU AI Act compliance, model cards, risk files, the boring but necessary stuff