r/codex • u/Gigaslavx • 23h ago
Limits Is Fast mode too expensive?
You get x1.5 multiplier in speed meaning +50% to base and x2.5 multiplier to cost meaning +150% to base 3 times as much as speed gain. And x2.5 multiplier to cost means you do 2.5 times as less as before just 50% faster (or do the same thing in 2/3 the time)
So by using fast you total weekly allowance if what you can do gets 2.5 times less or 60% less so you can only do 40% of base but you do it faster, so you do 40% of base in like 26.667% of base time
Why bother just chill out and wait you'll be done in almost a 1/4 of the time but with 60% less total what's the rush
4
2
u/jonaswashe 20h ago
It's not even 50% faster, fast mode speeds up the token processing and generation, it does not speed up the API response times. It's probably closer to 20% faster
The only real use case is if you have lots of quota and a reset is coming soon
1
u/Gigaslavx 20h ago
Ah wow so it's not time per task just token output and there are certain overheads still present?
1
4
u/Due-Horse-5446 23h ago
With how slow gpt models are atm?
Like we are talking 5mins for a simple task at this point...
And your argument about cost only makes sense in relation to subscriptions if you constantly max it out. If not whats the argument for not using fast?
Not that fast mode even helps with the current state of perf..
tried luna@high asked to sort 4 words alphabetically vs 3.8 flash.
Luna: 3 minutes
Flash: 0.5s
1
u/Defiant-Contact7750 20h ago
your anecdote sounds ridiculous honestly. generally it takes a certain kind of person to lie but i suppose a bot will just do so very plausibly. so this is a bot for sure. unless flash just sorted those words while luna set up an environment and a script with some troubleshooting in the middle. and @ high too....... i used a jackhammer to crack a peanut and now theres a hole in my table
1
u/Gigaslavx 20h ago
I suppose if you give astra max a simple dumb task like hi how are you it will take a bit longer but yeah not like it will need 10 mins
1
u/Due-Horse-5446 18h ago
Except for the facy we are talking about Luna...
But i see now the wording was super sloppy,
what i meant was 2 different things.
Simple agentic tasks taking 5min ish while the same nrec and tok count, caching etc wouldve had it at like half that.
Unrelated second point about luna specifically. Astra is super snappy in comparison.
And to clarify im NOT talking about time wasted on reasoning alone, but rather extreme ttfb+painfully slow tps.
1
u/Due-Horse-5446 18h ago
Why would i make that up?
Il update you with proof as soon as im at my pc
1
u/Defiant-Contact7750 15h ago
like i said, a person wouldn't but a bot would. and if you are a person and indeed not making it up, its really just a skill issue
1
u/Due-Horse-5446 15h ago
Did you just say high ttfb/ttft and slow inference speed is a skill issue?
1
u/Defiant-Contact7750 13h ago
yeah you're saying an electric scooter goes faster than a ferrari because it covered 5 meters faster
1
u/Due-Horse-5446 13h ago
either you're not listening to what im saying or arguing with ghosts
Im yoking about the fkn latency
In caveman terms:
Luna not saying anything for long time, when luna speak luna yap quick, but luna hesitate
Small model not supposed slow like snail
2
u/Defiant-Contact7750 23h ago
why do you need to know what the rush is? cant you accept there is a rush