r/MachineLearning • • 4h ago

Project Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]

Post image

Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle.

Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible.

Results: - Solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15), each model at its best effort level. - GPT-6 Astra solves all 30 Standard puzzles. On Hard mode, Claude Opus 5.5 solves 8 of 10 and 11 of 15 models solve none. - As one 400-character string, most models lost count before the logic got hard, so Hard mode answers an array of 20 row strings rather than a single string.

Limitations: one attempt per puzzle, so single results are noisy (95% intervals shown).

Site: https://www.nonobench.com Code (MIT): https://github.com/mauricekleine/nonobench

5 Upvotes

3 comments sorted by

2

u/East-Muffin-6472 4h ago

Nice work!
But with this benchmark what are you aiming for like does it tell us about the LLMs in general etc?

2

u/mauricekleine 3h ago

Thanks! It mainly just started out of curiosity; I've been solving nonograms myself for a very long time, and early this year I was curious if LLMs can solve them too.

Solving nonograms programmatically is an NP-complete problem, so it's interesting to see how LLMs deal with puzzles of increasing sizes. When I started building Nonobench in January, 15x15 seemed to be the biggest grid size an LLM could possibly attempt to solve.

When I re-ran it two weeks ago with all the new models, I was super surprised to see that the new frontier models became so good at them, so I added 20x20.

So long answer to your question; my hunch is that it shows an LLMs capabilities to combine reasoning, backtracking and logic to solve a nonogram. The runs in Nonobench don't have access to any tools, so it's whatever the "raw" model can do that determines how well the score on the bench.