
Curio — Where humans review the documents AI creates
Pin comments on AI-generated documents for the AI to read and revise.
ashaney · HN
Curio, a place for HTML files
The full gallery
Tech stack
25 projects

Pin comments on AI-generated documents for the AI to read and revise.
ashaney · HN
Curio, a place for HTML files

Compare AI coding models on real tasks with live previews, cost tracking, and ELO rankings.
@intheworldofai · X
On the World of AI Bench (vibe-coding composite): Claude Fable 5 → 85.2 GPT-5.6-sol → 82.4 kimi-k3 → 81.5 Moonshot’s K3 just walked in and claimed bronze on one of the toughest coding-focused leaderboards out there.

Score whether a website is AI-generated or hand-coded (0-100).
@mukparekh · X
I built tool for fun. roast website Free. No signup

Desktop app for running multiple AI coding agents in parallel across git worktrees
vuphanse

View and compare public opinions and benchmark ratings for leading AI models.
u/TasteMysterious5285 · Reddit
I built AI Census, a live field bulletin for how people are actually talking about AI models I’ve been building AI Census, a public “field bulletin” for how people are talking about current AI models. I kept running into the same problem: benchmark tables tell me how a model performs on a test, but not whether people are actually finding it useful, frustrating, reliable, etc. So I built a rolling view from public technical conversations across Reddit, Hacker News, Bluesky, GitHub, and Huggi

Multiple AI agents develop code in parallel with automatic error detection and fixing
@ForgeLab_Brain · X
Open beta launched. See more here: 🌐 #VibeCoding #BuildInPublic #IndieHacker #AIdev

Analyzes AI prompts to identify failures and return corrected versions.
u/pulptaken · Reddit
I built a website to audit AI prompts I'm giving away a limited number of early access invites for anyone who wants to try Marzel. You'll also be able to try the product once directly from the landing page, without joining the early access. This initial phase is focused on validating the core features and collecting feedback. After that, Marzel will also include a browser extension and a desktop app for auditing Claude Code and Codex CLI sessions, which is the project's main goal. If yo

Measure code review performance across GitHub teams with leaderboards and reviewer analytics.
u/SnooStrawberries827 · Reddit
my team had 47 open PRs and nobody was reviewing them, so I gamified it our team hit 47 open PRs at one point last month and nobody was reviewing them. tried slack reminders, deadlines, rotating reviewers, none of it really stuck. might be related to the fact that everyone's hyped about how fast AI can write code now, copilot cranking out entire features in hours, but none of that matters if the PR just sits there for a week. feels like writing code stopped being the bottleneck a while back

Upload code to automatically scan for security vulnerabilities.
@POONAMSING9999 · X
Forget this fight. The problem in AI is vibe coding security risk so i made vibe Guard ai that analysis user code and found out security risks. For the sake of humanity, to solve a painful problem I made this It had free version try now

Collect and manage user feedback and feature requests with an open-source AI alternative to Canny.
@kngkng182542 · X
If you're paying for Canny, give FeedLog a look. Same workflow for collecting feedback and managing feature requests, but free during launch.

Web-based terminal multiplexer for running AI coding agents in parallel.
04mg · Product Hunt
Caw Open source web terminal multiplexer for AI agents

Compare how different AI models generate frontend code and view accessibility scores.
u/12qwww · Reddit
I built a live benchmark to see which AI actually writes the best frontend code Hey everyone! I built OpenVibeEval because I was tired of "vibe-checking" AI-generated frontend code. I wanted to know which model actually produces the most accessible and clean React/Tailwind output. What I built: •A leaderboard of 24 models (Claude, GPT, DeepSeek, etc.) ranked by axe-core accessibility scores. •A Harness Comparator to show how different system prompts change the same model's output. •