Project detail · Developer Tools
Sakura — Benchmark for Local Coding Models
Benchmark local coding models on consumer hardware to measure accuracy, latency, and throughput across 27 tasks.
sakura.vaansh.dev

Try live demo Live checked14 hours ago
Try this first: View benchmark results for local coding models on the live leaderboard.
- Open source
- Yes
AnalyticsDeveloper ToolBenchmarkLlm“reproducible benchmark”“local coding models”“live leaderboard”“accuracy, latency, and throughput”
Tech Profile
- Built by
- Solo maker
- Time
- A weekend
- Open source
- Yes
Site health
🛠 3 site-health suggestions await the maker — claim this project (sign in with X) to view.
Source
I built a reproducible benchmark for local coding models (Ollama, 27 tasks, live leaderboard) ran it on my 8GB card, here's what I found I kept eyeballing "vibes" to decide whether one quant of a coding model was actually better than another on my machine, so I built Sakura to get real numbers instead. What it does: - Points at any Ollama model and runs it through 27 hand-curated tasks: codegen, bugfix, SQL, refactor, systems design, protocol implementation, and terminal-agent episode
For the maker: hang the plaque, claim the project
This project is unclaimed — sign in, hang the plaque on your site, and it's yours.
For the maker: hang the plaque, claim the project
This project is unclaimed — sign in, hang the plaque on your site, and it's yours.
Keep exploring
Similar projects
Sign in to report a problem with this project.














Comments
Sign in to comment.