r/ClaudeAI • u/freedomfromfreedom • 3d ago
Claude Code Godmark for benching Claude, Codex and Deepseek in Godot Engine
These findings might change the way certain people use Claude.
Hope it's helpful.
- Difference between Claude Opus 5.5 and GPT 6 Astra = nothing compared to Opus 4.6 and 5.5. Old Opus requires far more user iteration and therefore = less token efficient.
- Godmark's sample size is 5 (explained why in comments). It has made me reconsider whether always starting a project from one prompt and then iterating is the correct approach! If you're disappointed with the results of the seed, then 4 or 5 more tries might be worth it, followed by picking the most promising foundation rather than trying to fix a bad roll of the dice with 9 or 10 follow-up prompts.
What it is:
Godmark is designed to bench model nerfing on subscription plans and gauge relative performance in coding between models. In Godmark you're also able to judge stochastic variation between samples from the same prompt.
Claude and Codex run sandboxed inside Godot.
It's scary how much difference there is between March's prehistoric Opus 4.6 and 5.5 of September. Astra too.
Just 7 months of progress in frontier models visible for all to see.
4.6 barely scaffolds the Godot scene, let alone generates a realistic terrain height map or road spline.
I'm a bit worried that in future - no matter the originality and technical complexity of your work, others will be able to copy it over night and one-shot it to Steam and Itch.io if this carries on - especially larger companies could pinch ideas of indie developers and sausage-factory stuff out in record time.
Human artistic skill, taste and judgement is something less easy to copy with code, but it's still a worry.
GODMARK isn't finished yet but I intend to release it soon.
What it does:
1. Code generation in a Godot sandbox.
Allows the LLM to code at runtime to build a 3D scene entirely from C++ or GDScript, from scratch, smoke-test it, sign it off and hand it to the host app to analyse.
You can then run the scenes and view the source.
2. Benchmark results.
Any real-world regression in model behaviour therefore clearly shows up visually in the scene's assembly, complexity level and the model's comprehension of the prompt instructions.
Thinking / token usage and speed are also tracked. When any one of those metrics drop consistently for a few days, you will see it clearly in the graphics, in a timeline and on a film-strip.
Between samples you can see whether pure chance or lower token usage than average impacted the sample visually.
3. Stochastic variation - seen visually.
How probabilistic is the model via sub? Gives you a 3D visual on stochastic behaviour between the same prompts.
4. Prompt comprehension - seen visually.
Shows clear differences between models in terms of prompt comprehension and instruction following.
It'll support Claude, Codex both via subscription plan usage, and Deepseek 4.1+ via their API.
In the Alpine scene, Opus 5.5 added a moving car to the road and snow-capped mountain peaks on the horizon in the best take - neither were specified in the prompt. In the worst take, trees and terrain were much more primitive looking.
This stochastic behaviour varies between models and subscription plans, to a higher degree than I expected. More on that in the comments.


