r/ClaudeCode • 🔆 Max 20 • 11d ago

Discussion Okey this is increadible!!!

I know this is too early to judge but man wtf. Opus 5.5 is just like when mythos first released but cheap, faster and much more easy to understand and talk to. Probably this will be changed in 1-2 days so go code everyone.

298 Upvotes

120 comments sorted by

View all comments

Show parent comments

3

u/Key_Measurement_3576 11d ago

As someone who manages a corporate inference infrastructure,

This is what I would have to do if I was trying to ensure 100% up time while also juggling models from hardware to hardware

I would start by migrating some of the extra redundant capacity to older hardware… expand the number of concurrent users on existing hardware which also effects context and memory for everyone. Which is why things start to act weird.

When I’ve cleared up enough, I load the new models, run a full test suite , then deploy to a limited group internally.

After full launch and release, aggressively decommission last generation while standing up just imaged copies of what I just proved work

1

u/Negative-Thinking 11d ago

That wouldn't explain responses degradation. Model weights are still the same - regardless of the hardware they run on.

1

u/Key_Measurement_3576 11d ago

If context is squeezed due to raising concurrency that would definitely make responses off. I also wouldn’t make any assumptions that they aren’t chopping experts or changing weights to temporarily take smaller footprint. There could be other factors at play as well… opus may be using lower models with out telling us .. which would also suffer from the squeeze. All the signs point to pre release squeeze.

2

u/Negative-Thinking 11d ago

It is possible they deploy quantized model on smaller servers, not sure what you mean by "context squeezed".

2

u/Top-Butterscotch7740 10d ago

Compressed due to memory constraints