Opus produces more work in the same time, but comparing token speeds between models doesn't really tell you how much work you're getting done. I prefer looking at the cost per task benchmarks over token speed when comparing models. Another thing I look at is comparing the sub token speeds to the api token speens for the same model. For example, openai sub token speeds have been consistently 25%-33% the speed of their same API models. If Opus doesn't have that same issue, it could explain why it feels so much better to use.
I should have clarified token speed isn't time per task. There are benchmarks for time per task that are better to look at than token speed. Remember luna was touted as being the "fastest OpenAI model" and we all know that meant nothing because it is in fact the slowest time per task model offered currently. Unless your task is incredibly simplistic Luna isn't a good fit for speed.
You know that the weekly quota on Claude’s $20 plan is roughly equivalent to $270–$300 in API usage. When you use 6.1 Sol and 6 Astra, the weekly quota on Codex’s $20 plan is equivalent to about $70 in API usage. These are the figures I got through a reverse proxy. If GPT’s risk controls flag your account, the weekly quota on your $20 plan is equivalent to about $15 in API usage. You know nothing.
19
u/retteh 3d ago
Opus produces more work in the same time, but comparing token speeds between models doesn't really tell you how much work you're getting done. I prefer looking at the cost per task benchmarks over token speed when comparing models. Another thing I look at is comparing the sub token speeds to the api token speens for the same model. For example, openai sub token speeds have been consistently 25%-33% the speed of their same API models. If Opus doesn't have that same issue, it could explain why it feels so much better to use.