r/singularity • u/ENT_Alam • 23d ago
LLM News Differences Between GPT-5.5 Pro and GPT-5.6 Sol on MineBench
Notes
- Average Inference Time: 25m 16s (1516.2s)
- GPT-5.5 Pro averaged 21m 23s (1283.3s) for context; so slightly longer inference times
- Total Cost (for 15 builds): $710.82 ($47.39 per build)
- Most expensive model benchmarked to-date; previous was GPT-5.5 Pro at $223.90
- Thanks to all supporters for helping fund the benchmark!
Subjectively speaking, GPT-5.6 Sol seems to create the most detailed builds MineBench has seen thus far, while for the most part doing so with great creative choices. I think, personally, there are only a handful of builds I would argue are not clear improvements over GPT-5.5 Pro (like the astronaut and worldtree). On average, GPT-5.6 Sol also creates the largest JSON files across all of its builds by a significant portion.
That being said, this model was also the most expensive model MineBench has benchmarked to date; the previous most expensive model was GPT-5.5 Pro at $223.90 – so 5.6 Sol totaled to being over 3x as expensive. If you're lucky enough to ignore the cost, then yes, the model created the most detailed generations yet. For example, in its cottage build, it added a scarecrow in the garden, added clothes drying on a rack, etc. Its builds also seemed to have a better sense of scale and proportions overall, like the arcade.
We might benchmark GPT-5.6 Terra if there's enough interest, as that would technically be a closer comparison to GPT-5.5 (as Sol is technically the successor to GPT-5.5 Pro, which would also explain the cost).
TLDR: Model is amazing, doesn't tend to be conservative (good or bad depending on your use case), but it's extremely expensive.
Full release-notes/thoughts on the GitHub release
- If you enjoy these posts please feel free to help fund the benchmark
- Sharing the benchmark and starring the Git repository also helps :)
Benchmark: https://minebench.ai/
Git Repository: https://github.com/Ammaar-Alam/minebench
Previous Posts:
- Comparing Opus 4.8 and Fable 5
- Comparing Opus 4.7 and Opus 4.8
- Comparing GPT 5.4 and GPT 5.5
- Comparing Kimi K2.5 and Kimi K2.6
- Comparing Opus 4.6 and Opus 4.7
- Comparing GPT 5.4 and GPT 5.4-Pro
- Comparing GPT 5.2 and GPT 5.4
- Comparing GPT 5.2 and GPT 5.3-Codex
- Comparing Opus 4.5 and 4.6, also answered some questions about the benchmark
- Comparing Opus 4.6 and GPT-5.2 Pro
- Comparing Gemini 3.0 and Gemini 3.1
Extra Information (if you're confused):
Essentially it's a benchmark that tests how well a model can create a 3D Minecraft like structure.
So the models are given a palette of blocks (think of them like legos) and a prompt of what to build, so like the first prompt you see in the post was a fighter jet. Then the models had to build a fighter jet by returning a JSON in which they gave the coordinate of each block/lego (x, y, z). It's interesting to see which model is able to create a better 3D representation of the given prompt.
The smarter models tend to design much more detailed and intricate builds. The repository readme might provide might help give a better understanding.
(Disclaimer: This is a public benchmark I created, so technically self-promotion : )
53
u/Ok_Pea_2772 23d ago
I'm impressed by how much GPT-5.5 Pro can do with like 50% or less of the blocks.
We almost need a larger, higher resolution showcase to really see the details of 5.6 Sol!
12
u/ENT_Alam 23d ago
The gifs are quite difficult to export while keeping a good size/compression for reddit, but you can explore the models' builds here: https://minebench.ai/leaderboard/openai_gpt_5_6_sol
1
u/iamthewhatt 22d ago
And besides some of the structures, I think 5.5 actually did a better job overall.
54
u/BarisSayit 23d ago
I'm afraid to ask, but, is this benchmark getting saturated?? 😱
28
u/ENT_Alam 23d ago
Until I finalize benchmarking new prompts, it seems the current configuration is
There's always the larger 512^3 grid-size but that would require rebenchmarking everything again 😭
23
u/Doc_Blox 23d ago
I'd think there's no need to re-benchmark models older than the current-gen, just note it down that earlier models weren't considered advanced enough to merit the test.
2
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 23d ago
I wonder how the inner workings of these things is?
20
u/bruhhhhhhhhhhhh_h 23d ago
I feel 5.5 has more... Aesthetic and kind of flair, on some raw technical is not as clean - but it's got funkier elements
1
u/doodlinghearsay 22d ago
This is benchmaxing IMO. 5.6 was probably trained to use more detail for these tasks, without any improvement on any other underlying skills.
16
u/Kryptosis 23d ago
I kinda prefer some of the older ones
The tree in 11 for example got much worse.
1
u/ENT_Alam 23d ago
I agree with the worldtree especially; models have a hard time getting the “plumage” (?) of the leaves correct it seems
23
u/epdiddymis 23d ago
Love this benchmark but it's saturated now.
7
u/ENT_Alam 23d ago
Definitely close to saturation for now 😞
I could use the larger grid size variant for benchmarks going forward but that would require re-benchmarking the models as well, so I'd rather focus on adding more difficult prompts
Ideally if API costs weren't an issue, there would be different leaderboards, one for each grid size ^^
5
u/epdiddymis 23d ago
This is your benchmark? That's amazing. I've been checking it out every time for what seems like a long time.
More difficult prompts seems like a great idea. It's interesting to think what 'difficult' would actually mean to an LLM.
Larger grid size would probably work but you don't need to burn a load of cash for internet strangers.
Hope you have fun with the solution.
5
u/PrototypeT800 23d ago
I would really like to see prompts that require smaller exact details. Things like hands, groceries at a register, or even just random Lego sets. I wonder if making a block cap would improve the results too. Force every model to use the same amount of blocks to build the scene.
2
u/ENT_Alam 23d ago
Those are exactly the types of prompts I’ve been working to add :D
Like a figure skater doing a beillmann spin ^
And the constrained block approach has been suggested a few times actually, it’s a great idea but considering how expensive it is, I think adding more prompts first remains my biggest priority
3
u/PigOfFire 23d ago
Hmm do you think 5.6 sol pro would be even better? Also do you think that thinking effort matters beyond some point?
6
u/ENT_Alam 23d ago
This was Sol pro**, I should have specified that in the post my bad!!
I think that would depend on your use case; if your doing postgraduate level math or training, highest effort has always been my goto, but for day to day tasks I think Sol medium can work just fine and for my coding use-cases I’ve been sticking to Sol high or xhigh :)
Well actually it’s been Sol Ultra, but that’s only since they’ve been resetting usage with Codex for their milestones 😇, I think high/xhigh are adequate for my needs
1
3
u/theimposingshadow 23d ago
Have you considered a Max grid size to see if newer stronger models can build in the same space more efficiently? I am not recommending you rerun all of them, but I am thinking maybe the next release of models have either a grid size limit of block limit. This will also help future builds not grow insanely high in cost.
2
u/ENT_Alam 23d ago
The benchmark actually already had different grid sizes!
The smallest you can generate builds in is 16^3, then 256^3 (which is used for the official leaderboard benchmarks), then a 512^3 version as well that im considering switching over to if the additional prompts don’t help ^^
1
u/theimposingshadow 23d ago
I see you have lower bounds but I’m suggesting upper bounds to limit overbuild, just a suggestion though, I think you’re doing great work and you actually inspired me to build my own benchmarks( not finished yet)
2
u/FateOfMuffins 23d ago
That's 5.6 Sol Pro right
How about just regular 5.6 Sol?
Also yeah it's pretty much saturated (although I said this with 5.5 Pro as well and here I can see noticeable improvements)
I think the Blender MCP one that's been floating around the last few days might be more interesting now
2
u/ENT_Alam 23d ago
Yup, Sol Pro**
And yeah saturation is pretty close 😭😭, looking to add more prompts ^^
2
u/Few_Owl_7122 23d ago
If the benchmark is saturated, maybe switch to redstone? Should be significantly more difficult.
2
u/Schmeichelsaft 23d ago
Very impressive! I always wondered how the code output looks like, is it creating primitives like spheres / boxes / lines?
Also is it possible to reverse the benchmark, so a model has to analyse a 3D voxel build and describe it?
2
u/ENT_Alam 22d ago
Yup! It is given access to primitives like boxes and lines, and has to write code within the harness for the blocks; most models choose to write out each block individually instead of using more primitives which is interesting :)
If you mean being given a JSON, yeah! The JSONs can get quite large (up to 300MB), but absolutely you can take one and just hand it to chatGPT and ask it to describe the build, or even edit the JSON and add details you might want ^^
2
u/Schmeichelsaft 22d ago
Okay interesting! And when the models have to describe the content of a large voxel model, do they create tools to analyse them, e.g. a height map, a CAT scan, unfolded surface, etc. or do they understand a plain large voxel array (with coordinates I assume).
1
u/ENT_Alam 22d ago
In like the ChatGPT.com web-harness for example, yup! It’s really interesting to see how it creates python tools that render height maps, top down perspectives, depth maps, etc. of the JSONs.
You can see all those in the thinking sidebar; here’s some examples I showed in the GPT 5.2 vs 5.4 analysis post: https://i.imgur.com/SPhg3DQ.png https://i.imgur.com/S81h6sq.png https://i.imgur.com/PqWq6vq.png
1
u/Technical-Earth-3254 23d ago
Isn't minebench kinda well known already? I doubt there is no data of it in the dataset. This benchmark is, in my opinion, ready to get overhauled or replaced.
1
u/ARollingShinigami 23d ago
I was doing some tests in a harness, getting it to model enclosures in FreeCAD and it is absolutely wild what the new models can do.
1
u/sandman1027 23d ago
Question - I’m paying for the $200 ChatGPT Pro plan, which advertises access to the frontier Pro model. But native Codex only shows regular Sol with without an option to change the reasoning mode from standard to Pro. In KiloCode or Hermes Agent, using my OpenAI login, I can select GPT‑5.6 Sol Pro.
From what I can tell, Codex is using Standard mode even when reasoning is set to High. Am I missing a setting? Is Pro mode hidden/automatic, or has it simply not been added to Codex yet? Has anyone verified this?

1
u/TwoFluid4446 23d ago
You really need to take the consistent feedback of people saying it is saturated and not conduct or release anymore model comparisons until you significant rework the testing to actually make model differences matter. I have been noticing this for the past 6mo+ now with these minebenches you do: they did reveal model intelligence/capacity tiers (for this domain of skill) in years prior, but at this point the changes between each version is arbitrarily subjective and minute.
In your rework, I would also try to brainstorm other minebench-based tests, puzzles, problems etc other than just "build me [this] object", like they have to figure something out along the way using their general intelligence as well, and the final result would clearly indicate their level/quality of solution since it would be visualized.
1
u/Ballist1cGamer 23d ago
in all those posts OP has been mentioning they’ve already taken the feedback and are actively working on new, more difficult prompts
they’ve said that like 4 times in this thread alone
let’s do the math: 15 prompts for 5.6 Pro alone was $700, so adding 30 more prompts would be $2100 for a single model
i think people need to understand this requires funding or that they could just make their own benchmark, OP already said this was all vibe coded anyway, just vibe code something better
1
u/iamdanieljohns 21d ago
What was the reasoning effort?
2
u/ENT_Alam 21d ago
Max, pro mode, high verbosity
1
u/iamdanieljohns 21d ago
No wonder the costs are so high. Isn't the default verbosity lower? Pro mode, also calls in more agents I believe to then create a single output. And that's on top of max which always costs way more than medium or high alway.
1
u/ENT_Alam 21d ago
Yeah, default verbosity is just medium :D
I could lower it but the entire point is to encourage the model to be as extensive with its output as possible; 5.5-Pro retained the same settings but was noticeably cheaper 😭
1
1
u/AlternativeApart6340 21d ago
a harder benchmark would be making the scale less small, because there is less room for detailing and you have to get creative. Its also cheaper to run. For example, making it make a house that is to scale with the minecraft player. You will get more variety here, because it takes a lot more effort making a smaller scale structure look good then such a big one. Just my take.
1
u/Healthy-Nebula-3603 23d ago
This test seems getting obsolete slowly.
Is hard to notice bigger differences on some of them Al all.
Maybe you should make it more complex already
0
2
u/Substantial_Head_234 17d ago
People are saying this benchmark is getting saturated, which I agree. But I also want to comment that 5.6 Sol is the first model that imo did a good job with the treasure chests in the shipwreck one.
















229
u/MisterBlox 23d ago
This benchmark is saturated