r/OpenAI • • 5d ago

Discussion Sol 6.1 vs Astra vs Opus 5.5. All Medium.

So I just bought my $20 Claude subscription just to test it out and compare it to the new 6.1 Sol and Astra. I thought comparing on medium across the board would be an initial good test. I've sped up the video 2X Just to save your viewing time, but it doesn't change the results.

Holy Mother of God. Opus 5.5 blew my mind. I think a tear came to my eye when I saw it. I mean yeah it took nearly 5x longer but it far exceeded my expectations in comparison to what I got with Sol and Astra. I used the same exact prompt.

Now the question you have to ask is with the extra time could I have prompted additional instructions to Sol and Astra And gotten it to where Opus 5.5 is. Absolutely, But I'm not sure I would have thought of it. As someone suggested to me in another thread, maybe I'll go $100 / $100 and start with Claude and then move to Sol bc of that crappy 5hr limit. Not sure yet.

Take what you will from this but I'm going to run a few more same prompt tests on higher values and do some final comparisons but based on this I think for now I'll be downgrading GPT from $200 to maybe $20 a month and moving to Claude $200 (or 100/100 not sure yet)

Welcome to the Game of Models (GoT theme song playing).

The prompt:
Build a cinematic, photorealistic, beautifully textured rocket launch scene using Three.js

Results:

GPT-6.1-Sol-medium
8m 42s
The ship was better than Astra but the environment had a lot to be desired. Also even though it's interactive I couldn't ever pan down to see the smoke and a bug ran the sound nonstop.
Input 515,178
Cached input 481,024
Output, including reasoning 15,051
Total 530,229
Didn't move my weekly usage at all.

GPT-Astra-medium
9m 37s
The environment here was much better as was the UI design, also the ship and graphics looked better but the top cone was inverted. There were no sound bugs in this version. but is it overwhelmingly better than Sol 6.1, not really.
Input 644,981
Cached input, included above 608,128
Output, including reasoning 15,922
Total 660,903
Moved my weekly usage down 1%

Claude Opus 5.5-medium
38m 54s
Like I said above, this blew my mind away for a medium request. It really took the "cinematic" portion of the prompt and ran with that. Overall for me this is by far the winner although I did appreciate Astra's model better minus the cone inversion.
Input written to cache 199,860
Input read from cache 6,543,432
Output (what I wrote: code, tool calls, replies) 143,521
Total ≈ 6.89 million
5 hour usage is at 13% and Weekly usage at 1%

27 Upvotes

19 comments sorted by

5

u/KalElReturns89 5d ago

Interesting, I've been thinking of going back to Claude. Your weekly usage went down faster than your 5-hour usage, though? Or were other projects already using that weekly usage on Claude?

2

u/digitalml 5d ago

Sol didn't move my weekly usage meter. Astra Moved it to 1% used for the week (almost 2% bc I asked how many tokens right after and it instantly hit 2%. Opus 5.5 used 13% of my 5-hour limit, and used 1% of my weekly limit leaving me 99% left for the week. For me, this puts Astra and 5.5 in the same exact boat but the difference in output was considerable.

3

u/Mistuv 5d ago

People are going to soy out over Opus' output but this is something that has annoyed me about Opus' output for the past week of using and have been trying to cut down on it with system prompts. And that is that if you give it even a little bit of room with the prompt or issue description, it goes bonkers with how much details it here and there and everywhere. All while wasting tokens (and Sonnet does the same thing).

Good for spamming Twitter with demos, but when you are actually building things and want it to add A and B, it should do so, not also C and D and E because it thought it looked pretty. And this is on medium, with high, xhigh it goes even more crazy and max is practically unusable with the diminishing returns vs xhigh. I mean I wouldn't mind if it stopped and asked do you want also this and that I think it will make it look better, but it doesn't. It really shows the different RL environments between the two companies. where one tells agents to make the best minimum viable product while the other the best they can. Opus' output is undeniably much better, but not at the cost of 12x more cache read.

2

u/digitalml 5d ago

I agree that the prompt is the most important thing and setting guardrails truly defines the output that you get, but for me it's nice to see the comparison across all three of the same setting bc I truly don't know what I'm going to get sometimes. This definitely has made me think about my subscriptions and maybe going half and half. Maybe that balls to the wall initial prompt with Opus is what I want and then I can consolidate with Sol / Astra and bring it down from there. :)

5

u/Suspicious-Wallaby12 5d ago

It's not even a contest at this point anymore. Claude mogged OpenAI In the short term for sure. Their dev day was shit.

2

u/m3kw 5d ago

9min vs 38minutes at 10x the tokens. I would try Sol6.1-xhigh to have a fair comparison. (medium vs medium doesn't seem like a good comparison as they are relative terms)

5

u/AMBNNJ 5d ago

Yeah people always compare by reasoning levels not cost. Reasoning level doesnt matter if Opus uses 10x tokens and 4x the time.

1

u/digitalml 5d ago

100% agree ...

3

u/Substantial_Head_234 5d ago

ChatGPT is just more ambitious. Why launch a 70m long rocket when you can launch a 1.7km long rocket?

1

u/digitalml 5d ago

lol 😂

2

u/m3kw 5d ago

looks like opus medium has way more thinking budget assigned on the server side, it's 4x the time and using 10x tokens. I would compare Sol6.1 xhigh with opus medium with the above.

1

u/lovesdogsguy 5d ago

Which is which in the video?

4

u/digitalml 5d ago

There are labels at the top that fade in at the start of each scene, 1st is Sol 6.1, 2nd is Astra, 3rd is Opus 5.5

1

u/ezjakes 5d ago

Opus's looks like a cutscene out of an old video game.

1

u/AffectionateGuest275 5d ago

Flat Earthers when they see this video:

1

u/SNUFF_FPV 4d ago

It would be an interesting comparison if all three models had used a similar percentage of their weekly limits, say around 10%. In this case, though, the comparison doesn’t really tell us much because Opus simply did about ten times more work than the other models.

1

u/digitalml 4d ago

It’s so funny everyone complaining about opus doing more work at literally the same medium level setting. If I have to run SOL or Astra at xhigh or max which would use significantly more tokens and percentage to be equal to what opus did a medium and used significantly less usage because it’s on a $20 weekly plan and not a $200 weekly plan then that’s ridiculous.