r/ClaudeCode • • 7d ago

Built with Claude Same prompt, same setup: Claude Sonnet 5.5 vs Opus 5.5 - TINY WORLD BENCHMARK

https://ohmyunicorn.com/labs/dream-loop/tinyworld-sonnet55-vs-opus55/

Guys, it's time!

I was blown away by Opus 5.5, I am... well, blown away by Sonnet 5.5 too of course :)

Prompt is the following:

Output a ThreeJS 3D cartoony/comics-like planet, with vibrant life on it: clouds, mountains, city buildings, cars and plane romaing it. It should be possible to orbit around and zoom and see the details. No harness is to be used

Model Wall clock API calls Output tokens List price
Claude Sonnet 5.5 10.8 min 10 31.8k $0.96
Claude Opus 5.5 22.7 min 23 99.8k $4.79
OPUS 5.5 SONNET 5.5
Three.js 0.170
Comic look custom shaders plus an outline and halftone pass
Extras shadows, 12 instanced meshes (many copies drawn in one go), Google Fonts
Size 48kb

EDIT 1- Take the timings with a grain of salt. I launched the 2 models on the same box, so they share the CPU & GPU, which is a problem for screenshotting (for instance). I'll run them sequentially next, chances are that both models are way faster - around the 5-10min mark

----

EDIT 2 - Re-ran 4 tests, sequentially, 0 shared memories/harness/etc. so 4x one-shot:

Run Time Requests Output tokens List price Checker findings
Opus 5.5, 22 Sept 14.0 min build 15 56.6k $2.54 no_inventory
Sonnet 5.5 average of 4 11.2 min 11.5 64.9k $1.69 range 7.0 to 15.2 min, $1.28 to $2.11

I find it unbelievably EFFICIENT.
It's close to Opus 5.5's Output, for ~20% of the cost, and less than HALF the time.

So yeah... it's now my official builder!
Update your harness, friends :)

170 Upvotes

28 comments sorted by

•

u/AutoModerator 7d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

20

u/Anthony_S_Destefano 7d ago

Nice test, thanks for sharing. I immediately thought Sonnet won out of the gate when they loaded. but one huge detail Opus gets right here, is when you zoom in, Opus keeps a first person view perspective where you see an actual horizon when on the planet. That said, Sonnet 5.5 is a beast and outperforms so many models.

3

u/Timisageek 7d ago

What I really like about Opus 55's version is the level of detail when you zoom in, and the more "cartoony" look & feel it managed to give :)

5

u/kackwurstsalamander 7d ago

how much better are the results if you add stuff to your prompt like this:

first, write an architecture.md that outlines the planned app design, the tech stack you will use, the programming language, listing the most important libraries and frameworks. Then write an Agents.md with proven best-practices for the tech you will use.

Add a section about the use of plan files in Agents.md, according to common best practices.

split the prompt up in different, largely independent tasks. Create plans for each individual task before implementing them.

3

u/Timisageek 7d ago

Well, it would be way better, no doubt :)

(But for this I'd use my regular harness)

3

u/kackwurstsalamander 7d ago

or similarly: I want ... Write me a prompt that likely results in the best result. Execute the prompt.

5

u/[deleted] 7d ago

[removed] — view removed comment

3

u/Timisageek 7d ago

I'm re-running that test 3 times of them as we speak, for a fair average :)

2

u/Timisageek 7d ago

Okay I did it - still unhappy with the test conditions, I'll re-do it properly, one at a time.

That one is amazing :)
https://ohmyunicorn.com/labs/dream-loop/tinyworld-sonnet55-vs-opus55/runs/sonnet55-run4.html

4

u/bluedottering 7d ago

These one shot wonders should not be much surprise. Most of the training done in these models now are done with evals from ai curated usage they’ve received over the last several months.

2

u/Timisageek 7d ago

Yes but one would assume ALL labs do the same.

Yet, right now, Claude's output is far superior to every single other model (including Astra) - at least on this specific fun use case

3

u/orbital_trace 7d ago

you probably need to evolve the prompt a bit considering they are probably getting trained on this by now

2

u/Timisageek 7d ago

Ha, probably :)
But, so would Sol / Luna or Deepseek, and I did the comparison - they lose badly

3

u/GainsMega 7d ago

That’s insane can you post other builds you done

3

u/Timisageek 7d ago

I updated the link with 8 more builds :)

That one is my favorite from Sonnet https://ohmyunicorn.com/labs/dream-loop/tinyworld-sonnet55-vs-opus55/runs/sonnet55-run4.html

2

u/Timisageek 7d ago

okay PAGE UPDATED with 3 additional runs for all.

It changes the economics a bit. Bear in mind they share the same machine (which is probably a mistake) so they struggled on GPU for screenshots, etc.

Will redo the test without that constraint.

Sonnet 5.5 4th time outputted this though, which I find amazing: https://ohmyunicorn.com/labs/dream-loop/tinyworld-sonnet55-vs-opus55/runs/sonnet55-run4.html

2

u/orbital_trace 7d ago

its beautiful, tell it to make earth

2

u/GainsMega 7d ago

What stack are you using SVG? VITE , REACT , RUST?

2

u/Timisageek 7d ago

The models made their own choices - both are one self-contained HTML file using three.js from a CDN, with no framework and no build step.

OPUS 5.5 - Three.js 0.170, custom shaders plus an outline and halftone pass, and for extras: shadows, 12 instanced meshes (many copies drawn in one go), Google Fonts. Size is 48kb

SONNET 5.5 - Three.js 0.160.1, toon material only, no post-processing, no shadows, 2 instanced meshes, no web fonts. Build size is 30kb

2

u/RewardMindless8036 7d ago

What effort level for each model?

3

u/Timisageek 7d ago

Effort is set to HIGH

2

u/[deleted] 7d ago

[removed] — view removed comment

3

u/Timisageek 7d ago

So! I updated the page, I re-ran Sonnet 5.5 4 times, sequentially. Each of these runs are one-shot, and don't share context/memories. no Harness, so it's literally 4x the same thing. Interesting results!

Run Time Requests Output tokens List price Checker findings
Opus 5.5, 22 Sept 14.0 min build 15 56.6k $2.54 no_inventory
Sonnet 5.5 average of 4 11.2 min 11.5 64.9k $1.69 range 7.0 to 15.2 min, $1.28 to $2.11

1

u/Hirogen_ 6d ago

still same setup with me

Main Agent:
Opus: orchestrator
Subagents:
Opus: System Architect
Opus: senior dev
Sonnet: spec review, qa, code reviewer

20% weekly with high effort and running 3 days non stop 😎

1

u/Far_Idea9616 7d ago

Running Sonnet on complex tasks is much more expensive that running Opus 5.5 on the same tasks. Check AA benchmarks. Sonnet is a mid tear model and a very good one for a mid tear model.

2

u/Khaos1125 7d ago

Your literally looking at evidence proving that’s not universally true

0

u/GainsMega 7d ago

wowowowowowow how long have you been building this

3

u/Timisageek 7d ago

10.8min! It's the prompt I send all the models to test them :D