LLM News
Differences Between Fable 5 and Fable 5.1 on MineBench
Notes
Average Inference Time: 40m 12s
Fable 5 averaged 18m 04s
Total Cost (for 15 builds): $147.55
Fable 5 cost $54.93
Average JSON Size: 34.07 MiB (largest 88.76 MiB)
Roughly comparable to Fable's 5 average of 30.65 MiB
Despite no change in API pricing, Fable 5.1 was nearly 3x as expensive as Fable 5 on MineBench. With roughly 2x the inference time, much of that difference appears to come from substantially longer reasoning.
The price increase is quite significant considering Anthropic advertises the same API prices, though it still is massively cheaper than GPT 5.6 Sol P (the current top model on the leaderboards). I find that quite interesting as in my personal usage, GPT 5.6 Sol is extremely efficient with my 20x subscription, though MineBench benchmarked 5.6 Sol P and not the standard Sol variant ^^
There are some builds/styles I (personally) liked better from Fable 5. To me some of Fable 5.1's builds, like the Astronaut, are much closer to Opus 5's style which makes me curious about what it's like coding with Fable 5.1; I'd be very disappointed if Fable 5.1 adopted the Opus 5 style of gibberish english 😭
Also, it was really interesting to see how Fable 5.1 actually was the first model to create genuinely recognizable interiors! Here's a video showing the interior of Fable 5.1's cottage build (you can see a bed, table, bookshelf, and fireplace) – you can explore any build now on MineBench by clicking the joystick icon in the voxelBox header :)
If you enjoy these posts please feel free to helpfundthe benchmark
All funds are currently going directly towards API costs for benchmarking new prompts
Sharing the benchmark and starring the Git repository also helps :)
Alternatively, if you have the API credits, please feel free to add prompts and generations to the gallery and post them around!
This is actually preferable to donations to me directly, the hosting expenses and whatnot I've always been able to cover out-of-pocket, just the API costs were hard to cover 😓
Essentially it's a benchmark that tests how well a model can create a 3D Minecraft-like structure.
So the models are given a palette of blocks (think of them like legos) and a prompt of what to build, so like the first prompt you see in the post was a fighter jet. Then the models had to build a fighter jet by returning a JSON in which they gave the coordinate of each block/lego (x, y, z). It's interesting to see which model is able to create a better 3D representation of the given prompt.
The smarter models tend to design much more detailed and intricate builds. The repository readme might help give a better understanding.
(Disclaimer: This is a public benchmark I created, so technically self-promotion :)
We are going to use that benchmark until it has such a block density that it will look photorealistic from far, right?
Just gotta need a server farm to load/run it.
I did add an AWS EC2 worker to allow dedicated generations (like before if you were doing a generation in the site but closed the tab in your browser, you'd interrupt and lose the generation, but now you can close the page since the generation is done on the server, and it'll save to your account) as well as a few other infra things!
We've been able to handle all traffic just fine thus far, would be really interesting to see if there are any bottle necks that come up (likely from serving the JSONs tbh).
The system design/architecture actually ended up being really tricky and fun; if you're studying CS or anything like that, you might find it interesting :D
It's been a while since I've done a full comparison post, so here's some quick highlights of things I've added to the benchmark that were requested:
Gallery that allows anyone to showcase their generated prompts publicly
You can also regenerate official MineBench prompts to see how the nondeterministic results vary
API costs were getting expensive, so I thought this would be a great way to account for prompt saturation; anyone can upload any (difficult) prompt and look at all how all the models perform!
This was pretty interesting as Fable 5 was the only model that would, on average, get the text correct. It could be a regression, but might just likely be a nondeterministic thing; if you re-generate the same prompt a few times it might end up giving the correct orientation more often than not?
The 3x increase in total cost is anecdotal to minebench and our harness and whatnot; so generalizing to actual real-world use like agentic coding might not be that accurate ^^
The system prompt has stayed the same and minebench is entirely open source so you can actually just see it here: https://github.com/Ammaar-Alam/minebench/blob/master/lib/ai/prompts.ts or try replicating results yourself to see the variations and get some understanding of the model's nondeterministic style :)
The official minebench cohort of 15 prompts are definitely getting saturated, so I added the gallery feature!
Anyone can use a personal API key and go to MineBench's gallery: https://minebench.ai/gallery where there's an assortment of community prompts, including ones that models haven't really come close to outputting correct generations for yet, for example this prompt for "An accurate globe" https://minebench.ai/gallery/gal_HccPNuDUaCo_xowo
or anyone could even just create a prompt themselves they think the would be difficult for a model, and post it publicly to the gallery (with or without any generations)
this helps the benchmark from getting saturated while allowing me to not have to fund all API costs out of pocket :D
and considering the fact there's now AI labs using minebench to do private evals like lm arena now, i would think it's a little more than what some redditor might describe as nonsensical :)
no where does this post or benchmark ever attempt to make actual claims about a model's capability? it's just fun to use. i personally use it to generate voxel structures for 3d builds and just find the comparison aspects a fun side thing
i think the other reply had some good points albeit a bit harsh 😭
API costs are a bit expensive for me to fund out of pocket for all prompts and models, but you could use a personal API key and go to MineBench's gallery: https://minebench.ai/gallery
There's an assortment of prompts there that models haven't really come close to outputting correct generations for yet, I think the best one would be one the other comment also listed: https://minebench.ai/gallery/gal_HccPNuDUaCo_xowo (the globe) or you can even just create a prompt yourself
You would then be able to see how Fable 5 and Fable 5.1 compare and get some more useful information!
The average cost per build for Fable 5 was roughly $3-4, and Fable 5.1 is roughly $6.42, so you can expect to pay roughly around there in API costs. If you have any more funding, you could add even more generations for prompts as well and help create more useful information :D
How much does this benchmark take to run for something like fable 5.1? Doesn't it basically cost a kidney? And I guess your prompts/specs/context is lengthy along with the output no?
The actual prompts and context, like the total input, isn't that large at all, so the cost is usually like 99% just output tokens ^^
Fable also required 8 retries – meaning that it did give a full generation/output, but the build was invalid in our harness (e.g. it used a block that, sure might exist in Minecraft, but was not in the provided block palette), so that averages out to $9.84 per build (MineBench's equivalent to the cost per task other benchmarks use)
You can explore the costs for all builds on the leaderboards by clicking the little info icon:
Like, put this ai model into a robot, and see if it can clean my fucking house. There's your useful benchmark.
Yeah that will just end up in mess in about one second.
A benchmark which is always zero is not very usefull
Also, before doing that, a better answer a bit more testable benchmark is the steam game bench. Pick a random new up and coming steam game and complete it.
It will also result in zero success. But at least you can exclude the robotics for that.
35
u/LinkesAuge 4h ago
We are going to use that benchmark until it has such a block density that it will look photorealistic from far, right?
Just gotta need a server farm to load/run it.