
LLM Leaderboard | Redactle
Compare language model performance at solving Redactle puzzles.
pampas · HN
Redactle LLM Leaderboard
The full gallery
Tech stack
60 projects

Compare language model performance at solving Redactle puzzles.
pampas · HN
Redactle LLM Leaderboard

Compare how different AI models generate frontend code and view accessibility scores.
u/12qwww · Reddit
I built a live benchmark to see which AI actually writes the best frontend code Hey everyone! I built OpenVibeEval because I was tired of "vibe-checking" AI-generated frontend code. I wanted to know which model actually produces the most accessible and clean React/Tailwind output. What I built: •A leaderboard of 24 models (Claude, GPT, DeepSeek, etc.) ranked by axe-core accessibility scores. •A Harness Comparator to show how different system prompts change the same model's output. •

The end-to-end platform for small language models. Tuned to your task, a small open model matches frontier accuracy at a fraction of the cost. Own your intelligence: private, compa
@dayoffdev · X

Upload your data to fine-tune custom language models with cryptographic deletion proofs.
@MBrew26730 · X
Dataset cleaning + fine tuning + continual learning at

Compare speech-to-text engines (OpenAI, Deepgram, NVIDIA, Fish Audio) with real-time benchmarking and local privacy.
@alvaisy · X
finished voice to text small web app for my own itch. it's opensource. use openrotuer key. and use it with 4 models.

Build or auto-generate multilingual forms with logic, payments, signatures, and integrations.
@hkbonur · X
that digitalise paper forms and make everything multi language, so businesses can get better conversion. One form any language.

Master vocabulary with AI-powered translations, visual mnemonics, and pronunciation guides.
LandoLorinse · HN
Vocab Top – AI-powered vocabulary builder that helps you retain words

Chat with an open-source language model trained on European Portuguese.
wbemaker · HN
I Am Hosting Amalia – The First Portuguese LLM

See the concepts a language model holds at each layer before it answers.
ada1981 · HN
I built a web tool to see and edit what an AI thinks before it answers

Benchmark AI models by having them animate a 3D banana plant's full lifecycle.
fran-mora · HN
I gave 5 AI coding agents one prompt: grow a banana plant through its whole life in three.js: sprout, leaves, flower, fruit, rot, then pups that restart the loop. It's deceptively simple and yet very hard to get right from procedural code: you have to write working three.js and understand how the plant is actually built; how it hangs, ages and decays. Get the biology wrong and the code renders something weird. These are agents, not bare models (Claude Code and Codex for now). They can use tools, including playwright to check their work and improve it.

Submit kernel patches and engine optimizations for LLM inference speed, benchmarked on dedicated hardware.
carsenk · HN
Frontier.fast – Help push the frontier of LLM speed forward

Try semantic adapters that modify how frozen language models perceive marked text spans.
joshua_s_penman · HN
Semantic Overlays – an NX bit for LLM prompt injection (live demo)