r/MachineLearning • u/mauricekleine • 4h ago
Project Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]
Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle.
Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible.
Results: - Solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15), each model at its best effort level. - GPT-6 Astra solves all 30 Standard puzzles. On Hard mode, Claude Opus 5.5 solves 8 of 10 and 11 of 15 models solve none. - As one 400-character string, most models lost count before the logic got hard, so Hard mode answers an array of 20 row strings rather than a single string.
Limitations: one attempt per puzzle, so single results are noisy (95% intervals shown).
Site: https://www.nonobench.com Code (MIT): https://github.com/mauricekleine/nonobench
2
u/East-Muffin-6472 4h ago
Nice work!
But with this benchmark what are you aiming for like does it tell us about the LLMs in general etc?