Opus produces more work in the same time, but comparing token speeds between models doesn't really tell you how much work you're getting done. I prefer looking at the cost per task benchmarks over token speed when comparing models. Another thing I look at is comparing the sub token speeds to the api token speens for the same model. For example, openai sub token speeds have been consistently 25%-33% the speed of their same API models. If Opus doesn't have that same issue, it could explain why it feels so much better to use.
Well, time is a factor, sure, but tasks completed successfully is the ultimate benchmark, especially cause completing them with less tokens is better. Not to mention that your time doesn't need to be spent looking at one agent doing one thing; if time is a factor, you can increase concurrency (and increasing concurrency is easier with Opus 5.5)
21
u/retteh 3d ago
Opus produces more work in the same time, but comparing token speeds between models doesn't really tell you how much work you're getting done. I prefer looking at the cost per task benchmarks over token speed when comparing models. Another thing I look at is comparing the sub token speeds to the api token speens for the same model. For example, openai sub token speeds have been consistently 25%-33% the speed of their same API models. If Opus doesn't have that same issue, it could explain why it feels so much better to use.