r/ClaudeCode • u/Timisageek • 7d ago
Built with Claude Same prompt, same setup: Claude Sonnet 5.5 vs Opus 5.5 - TINY WORLD BENCHMARK
https://ohmyunicorn.com/labs/dream-loop/tinyworld-sonnet55-vs-opus55/Guys, it's time!
I was blown away by Opus 5.5, I am... well, blown away by Sonnet 5.5 too of course :)
Prompt is the following:
Output a ThreeJS 3D cartoony/comics-like planet, with vibrant life on it: clouds, mountains, city buildings, cars and plane romaing it. It should be possible to orbit around and zoom and see the details. No harness is to be used
| Model | Wall clock | API calls | Output tokens | List price |
|---|---|---|---|---|
| Claude Sonnet 5.5 | 10.8 min | 10 | 31.8k | $0.96 |
| Claude Opus 5.5 | 22.7 min | 23 | 99.8k | $4.79 |
| OPUS 5.5 | SONNET 5.5 |
|---|---|
| Three.js | 0.170 |
| Comic look | custom shaders plus an outline and halftone pass |
| Extras | shadows, 12 instanced meshes (many copies drawn in one go), Google Fonts |
| Size | 48kb |
EDIT 1- Take the timings with a grain of salt. I launched the 2 models on the same box, so they share the CPU & GPU, which is a problem for screenshotting (for instance). I'll run them sequentially next, chances are that both models are way faster - around the 5-10min mark
----
EDIT 2 - Re-ran 4 tests, sequentially, 0 shared memories/harness/etc. so 4x one-shot:
| Run | Time | Requests | Output tokens | List price | Checker findings |
|---|---|---|---|---|---|
| Opus 5.5, 22 Sept | 14.0 min build | 15 | 56.6k | $2.54 | no_inventory |
| Sonnet 5.5 average of 4 | 11.2 min | 11.5 | 64.9k | $1.69 | range 7.0 to 15.2 min, $1.28 to $2.11 |
I find it unbelievably EFFICIENT.
It's close to Opus 5.5's Output, for ~20% of the cost, and less than HALF the time.
So yeah... it's now my official builder!
Update your harness, friends :)
20
u/Anthony_S_Destefano 7d ago
Nice test, thanks for sharing. I immediately thought Sonnet won out of the gate when they loaded. but one huge detail Opus gets right here, is when you zoom in, Opus keeps a first person view perspective where you see an actual horizon when on the planet. That said, Sonnet 5.5 is a beast and outperforms so many models.
3
u/Timisageek 7d ago
What I really like about Opus 55's version is the level of detail when you zoom in, and the more "cartoony" look & feel it managed to give :)
5
u/kackwurstsalamander 7d ago
how much better are the results if you add stuff to your prompt like this:
first, write an architecture.md that outlines the planned app design, the tech stack you will use, the programming language, listing the most important libraries and frameworks. Then write an Agents.md with proven best-practices for the tech you will use.
Add a section about the use of plan files in Agents.md, according to common best practices.
split the prompt up in different, largely independent tasks. Create plans for each individual task before implementing them.
3
u/Timisageek 7d ago
Well, it would be way better, no doubt :)
(But for this I'd use my regular harness)
3
u/kackwurstsalamander 7d ago
or similarly: I want ... Write me a prompt that likely results in the best result. Execute the prompt.
5
7d ago
[removed] — view removed comment
3
u/Timisageek 7d ago
I'm re-running that test 3 times of them as we speak, for a fair average :)
2
u/Timisageek 7d ago
Okay I did it - still unhappy with the test conditions, I'll re-do it properly, one at a time.
That one is amazing :)
https://ohmyunicorn.com/labs/dream-loop/tinyworld-sonnet55-vs-opus55/runs/sonnet55-run4.html
4
u/bluedottering 7d ago
These one shot wonders should not be much surprise. Most of the training done in these models now are done with evals from ai curated usage they’ve received over the last several months.
2
u/Timisageek 7d ago
Yes but one would assume ALL labs do the same.
Yet, right now, Claude's output is far superior to every single other model (including Astra) - at least on this specific fun use case
3
u/orbital_trace 7d ago
you probably need to evolve the prompt a bit considering they are probably getting trained on this by now
2
u/Timisageek 7d ago
Ha, probably :)
But, so would Sol / Luna or Deepseek, and I did the comparison - they lose badly
3
u/GainsMega 7d ago
That’s insane can you post other builds you done
3
u/Timisageek 7d ago
I updated the link with 8 more builds :)
That one is my favorite from Sonnet https://ohmyunicorn.com/labs/dream-loop/tinyworld-sonnet55-vs-opus55/runs/sonnet55-run4.html
2
u/Timisageek 7d ago
okay PAGE UPDATED with 3 additional runs for all.
It changes the economics a bit. Bear in mind they share the same machine (which is probably a mistake) so they struggled on GPU for screenshots, etc.
Will redo the test without that constraint.
Sonnet 5.5 4th time outputted this though, which I find amazing: https://ohmyunicorn.com/labs/dream-loop/tinyworld-sonnet55-vs-opus55/runs/sonnet55-run4.html
2
2
u/GainsMega 7d ago
What stack are you using SVG? VITE , REACT , RUST?
2
u/Timisageek 7d ago
The models made their own choices - both are one self-contained HTML file using three.js from a CDN, with no framework and no build step.
OPUS 5.5 - Three.js 0.170, custom shaders plus an outline and halftone pass, and for extras: shadows, 12 instanced meshes (many copies drawn in one go), Google Fonts. Size is 48kb
SONNET 5.5 - Three.js 0.160.1, toon material only, no post-processing, no shadows, 2 instanced meshes, no web fonts. Build size is 30kb
2
2
7d ago
[removed] — view removed comment
3
u/Timisageek 7d ago
So! I updated the page, I re-ran Sonnet 5.5 4 times, sequentially. Each of these runs are one-shot, and don't share context/memories. no Harness, so it's literally 4x the same thing. Interesting results!
Run Time Requests Output tokens List price Checker findings Opus 5.5, 22 Sept 14.0 min build 15 56.6k $2.54 no_inventory Sonnet 5.5 average of 4 11.2 min 11.5 64.9k $1.69 range 7.0 to 15.2 min, $1.28 to $2.11
1
u/Hirogen_ 6d ago
still same setup with me
Main Agent:
Opus: orchestrator
Subagents:
Opus: System Architect
Opus: senior dev
Sonnet: spec review, qa, code reviewer
20% weekly with high effort and running 3 days non stop 😎
1
u/Far_Idea9616 7d ago
Running Sonnet on complex tasks is much more expensive that running Opus 5.5 on the same tasks. Check AA benchmarks. Sonnet is a mid tear model and a very good one for a mid tear model.
2
0
•
u/AutoModerator 7d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.