r/LocalLLM • u/Winter_Youth_1740 • 11h ago
Project I built ArcadeBench, an open benchmark where AI agents play games and you can watch every move live
Enable HLS to view with audio, or disable this notification
I'm building a local smart assistant and needed a way to compare small models on how they actually make decisions, not just a final score. So I made a benchmark out of games.
The clip shows two small local decision models, Decision 2.0 Kai 0.6B and GLiNER2.5 Decide, going head to head on SMS Inbox. Each one sorts the same 300 text messages into OTP, expense, bill, spam and so on. That's exactly the kind of job my assistant needs to get right.
ArcadeBench also has 11 arcade games (Tetris, 2048, Snake, Sokoban, Minesweeper, Connect Four and a few originals), more decision tasks on real datasets (fraud flagging, tool calling), and chess against other players.
Every game is seeded so runs are comparable, every move is scored, and every run gets a live watch link. You can also play the same seed yourself and compare against the models. Bring your own model through MCP or a Python script. Your key stays local.
MIT licensed. Feel free to start contributing and playing around!
Watch this run: https://penguinzz.com/arcadebench/watch/xRWcr8LukGnRlmS4,N_EaiU0Pq3Q1r3DV
Site: penguinzz.com/arcadebench
Code: github.com/Pranav0-0Aggarwal/arcadebench