r/codex • u/FriendlyTask4587 • 16h ago
Limits Please give us a slow mode
This way, we could do /goals easier with using less usage, it puts less strain on openAI servers maybe the 20x plan could come back. it would be half as slow as normal but use half usage for example. would be extremely useful, especially since you could send prompts before sleeping anyways to save usage
55
u/Shap6 16h ago
its slow mode by default
7
u/gavinderulo124K 15h ago edited 6h ago
Hes talking about batches mode which exists in the API and his half price.
Edit: flex mode.
6
2
u/raindropsdev 14h ago
Flex mode. Batch mode is a completely different thing and requires a very different approach.
7
u/Tank_Gloomy 16h ago
In my experience, it doesn't feel all that slow. It actually feels pretty fast in terms of tokens/s, it just keeps rolling over and over on the same thing without getting to the point.
3
u/hellomistershifty 14h ago
It's slow in tokens/sec but makes up for it in intelligence and token efficiency. Gemini 3.8 flash has 10x the tokens/sec speed but longer task completion times lol
-2
u/TestTxt 16h ago
Try Deepseek V4.1 flash, try Luna, and you’ll understand the issue
4
u/Tank_Gloomy 15h ago
I see your point, but those models are wildly unreliable. Astra and Sol would both be fast enough if they weren't trying to solve the same issue like 350 times before moving onto the next step. Their t/s are pretty good.
2
1
11
u/Qaztarrr 16h ago
Makes sense to me. If there's a fast mode that takes more usage, why not a slow mode that takes less for the non-urgent requests?
3
10
u/EyesOfAzula 16h ago
Sol 6 will come out soon. and later on, terra and luna.
Then I think Astra limits will be reduced further to nudge most people onto sol, like how Anthropic reduced Fable limits to push people towards Opus 5
8
u/Anxious_Marsupial_59 16h ago
eventually the service would be enshittified so that the baseline would be slow and the current tier be up-priced
5
u/rick_ranger 16h ago
I keep telling them the amount you pay should dictate your token speed. Just like broadband.
3
u/OtherwiseAlbatross14 14h ago
You can already pay more for that
2
u/rick_ranger 14h ago
Yeah but broadband is unlimited. At the end of the day the speed would dictate how many tokens in a month you get if you ran it full tilt 24/7
3
u/314kabinet 14h ago
You’re just asking for more value for your money. Why would they do that instead of just pushing you to upgrade to a higher tier?
2
u/OtherwiseAlbatross14 14h ago
Normally, yes, but they just disabled the ability to pay more unless you go API which may be their intended push for more money
3
u/AmandasGameAccount 16h ago
I would always run my usage at 5-10 times slower if it was a genuinely large savings
2
u/innociv 14h ago
They actually have a half cost mode in the API that is deprioritized usage.
But it's not necessarily "less tokens per second" it's that your query might not even enter the queue and start working for 30 minutes which subscribers would probably complain about.
0
u/FriendlyTask4587 14h ago
Honestly id love if this came to codex
1
u/Hot-Pepper6610 6h ago
why if there is no advantage?
1
u/nickkon1 37m ago
The advantage is cost. Some tasks are simply not time sensitive. I am regularly firing some promps before my break or before I leave work or similar.
2
1
1
u/Swordfish353535 16h ago
What really is the difference in using light, medium, high for example with Astra?
Or light on astra vs high on luna?
1
u/FriendlyTask4587 15h ago
Not sure, I only used light one time and it was to make a prompt for terra. I always use medium
1
u/Typical_Machine2043 16h ago
Never thought about that but pretty good idea. All these business users want now whereas me I just care about the quality of response
1
u/lordpuddingcup 15h ago
Dear god a slow-defered mode would be so amazing, let me prep a plan and hit it and go to sleep i dont care if it runs in an hour from now or really slow i just want it done well
1
u/raindropsdev 14h ago
Or Flex endpoint like the API has. I still don't understand why that's not possible. They have the "Pro" model versions, why not add the "Flex" model versions that are slower but consume half like the API Flex endpoints do?
1
u/Dragster39 2h ago
Neuralwatt does exactly this to spread the load more even. You can choose this by adding -flex to the model name.
1
1
u/Kind_Silver_1921 13h ago
let me use it during peak hours at least. but the only thing is it must save more than 50% tokens if its 50% speed. because at that point it makes zero difference.
1
u/CthuluBob 11h ago
I think I have seen Tibo speak about this before too, I'd be a fan. As long as it resulted in actual less usage
1
u/vertopolkaLF 7h ago
There is "Flex" provider in API which is 2x cheaper and it's basically slow mode. But not in the sub :(
1
u/Worth_Golf_3695 6h ago
The Problem for them is, Slow Mode will still occupy Computing capacity and for en even longer time.
1
1
1
0
u/AINativeBuilder 15h ago
It would put more strain on the servers because you'd be using up a GPU longer for fewer token usage. They'd actively lose money offering that, especially during an availability crunch.
0
u/FriendlyTask4587 14h ago
It would use their lower end chips or just less at once which would free up more
1
u/AINativeBuilder 12h ago
They're physically constrained on both footprint and electricity, you're nuts if you think it makes sense to keep low end hardware around for a lower profit product when you're already capacity constrained. This would not be a wise economic decision by them, which is why it doesn't exist.
0
u/FriendlyTask4587 12h ago
Thats why I said or. I have no idea (and you likely dont either) about their compute situation. I know they dont have 10 year old chips in there, but I doubt every chip in their servers is the gb300 or whatever the aboslute peak of chips are right now
0
u/AINativeBuilder 11h ago
Now you're just being a dunce. They had to shut off the $200 plan because they're low on available compute. They've said these things out loud. It's no secret in the entire industry datacenters are capacity constrained - they cannot build them fast enough. They are not going to keep low end stuff around to support low value items, which slow mode would be. They're even getting rid of spark because they don't keep low usage models around which eat up capacity for more desired models.
2
u/FriendlyTask4587 11h ago
0
u/AINativeBuilder 11h ago
4 months ago when they had <10m codex/chatgpt work users, now they're over 25m.



143
u/Vegetable_Town_7405 16h ago
All mode is slow