
Mad World
Watch AI models debate your questions from different perspectives.
@PraveenJangid99 · X
Multi Agent Discussion Rooms ;Mad World
The full gallery
Tech stack
61 projects

Watch AI models debate your questions from different perspectives.
@PraveenJangid99 · X
Multi Agent Discussion Rooms ;Mad World

Product leaderboards decided by an adversarial AI jury. Every rank is taken from the product that held it, in a trial with graded evidence and published reasoning.
@oleks_i · X

Compare and evaluate AI models across coding, reasoning, agents, and other benchmarks.
davidtsong · HN
Benchmarklist: track AI benchmarks (2.4k+), models, and capabilities

AI assistant helping enterprise employees preserve and access institutional knowledge and judgment.
@romanbodnarchuk · X
Check out what I just built with Lovable!

Chat with AI models, compare them, and vote to shape a community leaderboard.
u/Rabus · Reddit
I got TestingModels too overcomplicated over the month it is running: looking for some feedback how to make it more useful and simpler I run a benchmark like arena.ai , but with pre-generated prompts. So far, nearly 6k people came in and like 30k comparisons has been made - which means the thing is genuinely useful for people to compare the models. The problem is the more features i started adding the more overblown and complicated UI became - like old internet explorer tab bars Old: ht

Improve your reasoning by solving adaptive AI-generated puzzles that learn from your mistakes.
u/connerpro · Reddit
IntelligenceMax - Adaptive reasoning practice with live AI questions (claim-safe near vs far) submitted by /u/connerpro to r/SideProject [link] [comments]

Score whether a website is AI-generated or hand-coded (0-100).
@mukparekh · X
I built tool for fun. roast website Free. No signup

Ask one question to multiple AI models, compare their answers, and watch them debate to consensus.
u/trekhleb · Reddit
I kept pasting the same question into ChatGPT, Claude, and Gemini in three tabs; so I built a Yes-Brainer — a council of AI models, that answer your question in parallel, debate to consensus, or get judged to a verdict. submitted by /u/trekhleb to r/SideProject [link] [comments]

Scores AI-generated ad creative and returns verdicts (run/fix/kill) via MCP and REST API.
ds246 · HN
Spendict – a performance marketer's verdict for AI agents, over MCP

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

An AI agent reads the sacred texts of 15 world religions, scores each passage on 7 philosophical criteria, and evolves a belief system in real time. Watch it think.
@synapticsaga · X
Simple. Conviction engine. Currently this free AI model is trying to pick a religion. You can watch it think in realtime:

Take a conversational assessment to measure your AI skills and earn certification.
u/Ozan_D · Reddit
AISA - AI Fluency Assessment In the last 6 months I built an AI Fluency Assessment system (called AISA) - it's currently the most sophisticated (and popular) of it's kind. Has high fidelity in what we measure to both what Anthropic and US. Dept. of Labour agree as the markers of AI fluency. People chat with an AI agent uniquely trained to judge their AI fluency, get a detailed report and a breakdown of their AI skills, a certificate and a growth roadmap. We measure in 5 main dimensions.