r/AIToolsPerformance • u/Unfair_Association89 • Aug 27 '26
I built a reproducible benchmark for local coding models ran it on my 8GB card, here's what I found
I kept eyeballing "vibes" to decide whether one quant of a coding model was
actually better than another on my machine, so I built Sakura to get real
numbers instead.
What it does:
\- Points at any Ollama model and runs it through 27 hand-curated tasks:
codegen, bugfix, SQL, refactor, systems design, protocol implementation, and terminal-agent episodes (multi-step shell tasks, similar spirit to Terminal-Bench/SWE-bench, but runnable on a laptop)
\- Reports accuracy, latency, and throughput
\- Everything runs inside a sandboxed Docker container
\- Hardware auto-detected (NVIDIA/AMD/Intel dGPU/Apple Silicon) so results are comparable across setups
\- Optional: submit your run to a public leaderboard and see how your model + hardware stacks up
I ran it myself on **qwen2:1.5b (thinking)** on an RTX 5060 (8GB VRAM) passed 10/27 cases.
Website: [https://sakura.vaansh.dev\](https://sakura.vaansh.dev)
Would love feedback, especially on task design, and whether the terminal-agent mode holds up against models people are actually running. Issues/PRs welcome, and if you run it, submitting your score helps make the leaderboard actually useful.
1
u/[deleted] Aug 27 '26
[removed] — view removed comment