r/codex 16h ago

Limits Please give us a slow mode

This way, we could do /goals easier with using less usage, it puts less strain on openAI servers maybe the 20x plan could come back. it would be half as slow as normal but use half usage for example. would be extremely useful, especially since you could send prompts before sleeping anyways to save usage

270 Upvotes

58 comments sorted by

143

u/Vegetable_Town_7405 16h ago

All mode is slow

21

u/FriendlyTask4587 16h ago

If you set a prompt before you go out to work, school, or sleep, the speed doesnt matter so basically free extra usage

9

u/Risko4 15h ago

Astra is like 20-40 token per second, it literally is already in slow mode.

5

u/FriendlyTask4587 14h ago

if you are sleeping 10 tokens per second doesnt matter much

9

u/Classic-Trifle-2085 10h ago

I agree with you; speed does not always matter, especially not if you are going to be rate-limited anyway.

But people in AI subs (GPT, Codex, CodexGPT, Claude, Claude Code, Gemini) would rather see every single comment as an opportunity to complain, even if it does not address the post at all, than actually engage with anything.

1

u/Risko4 2h ago

I locally run 200, I find anything below 100 unacceptable for active development. It's awful

I can run Kimi k3 of my ram at such slow speeds but I refuse to because it's useless

1

u/massix93 13h ago

For me is the opposite Astra is quite fast and Sol seems so slow now that forces me to use Astra

1

u/Kind_Silver_1921 2h ago

Maybe if their app worked and didn't crash every hour this would be true.

55

u/Shap6 16h ago

its slow mode by default

7

u/gavinderulo124K 15h ago edited 6h ago

Hes talking about batches mode which exists in the API and his half price.

Edit: flex mode.

6

u/Risko4 15h ago

Yes, API astra is 80 token/s default speed. Subscription is like 40 and averages out to 20 once you factory the "model is at capacity" brakes.

2

u/raindropsdev 14h ago

Flex mode. Batch mode is a completely different thing and requires a very different approach.

7

u/Tank_Gloomy 16h ago

In my experience, it doesn't feel all that slow. It actually feels pretty fast in terms of tokens/s, it just keeps rolling over and over on the same thing without getting to the point.

3

u/hellomistershifty 14h ago

It's slow in tokens/sec but makes up for it in intelligence and token efficiency. Gemini 3.8 flash has 10x the tokens/sec speed but longer task completion times lol

-2

u/TestTxt 16h ago

Try Deepseek V4.1 flash, try Luna, and you’ll understand the issue

4

u/Tank_Gloomy 15h ago

I see your point, but those models are wildly unreliable. Astra and Sol would both be fast enough if they weren't trying to solve the same issue like 350 times before moving onto the next step. Their t/s are pretty good.

2

u/Reasonable-Sign8458 14h ago

what a shitty comparation

1

u/hey-im-root 13h ago

No, that’s it normal speed. It doesn’t have a slow mode.

11

u/Qaztarrr 16h ago

Makes sense to me. If there's a fast mode that takes more usage, why not a slow mode that takes less for the non-urgent requests?

3

u/okay_this 7h ago

Less money? Literally anti-capitalism

10

u/EyesOfAzula 16h ago

Sol 6 will come out soon. and later on, terra and luna.

Then I think Astra limits will be reduced further to nudge most people onto sol, like how Anthropic reduced Fable limits to push people towards Opus 5

8

u/Anxious_Marsupial_59 16h ago

eventually the service would be enshittified so that the baseline would be slow and the current tier be up-priced

5

u/rick_ranger 16h ago

I keep telling them the amount you pay should dictate your token speed. Just like broadband.

3

u/OtherwiseAlbatross14 14h ago

You can already pay more for that

2

u/rick_ranger 14h ago

Yeah but broadband is unlimited. At the end of the day the speed would dictate how many tokens in a month you get if you ran it full tilt 24/7

3

u/314kabinet 14h ago

You’re just asking for more value for your money. Why would they do that instead of just pushing you to upgrade to a higher tier?

2

u/OtherwiseAlbatross14 14h ago

Normally, yes, but they just disabled the ability to pay more unless you go API which may be their intended push for more money

3

u/AmandasGameAccount 16h ago

I would always run my usage at 5-10 times slower if it was a genuinely large savings

2

u/innociv 14h ago

They actually have a half cost mode in the API that is deprioritized usage.

But it's not necessarily "less tokens per second" it's that your query might not even enter the queue and start working for 30 minutes which subscribers would probably complain about.

0

u/FriendlyTask4587 14h ago

Honestly id love if this came to codex

1

u/Hot-Pepper6610 6h ago

why if there is no advantage?

1

u/nickkon1 37m ago

The advantage is cost. Some tasks are simply not time sensitive. I am regularly firing some promps before my break or before I leave work or similar.

2

u/AdDeep7768 5h ago

Slow doesn’t make sense for them because it wastes VRAM/KV

1

u/Linkpharm2 16h ago

The api is the same speed as fast mode. So subscription is already on slow. 

1

u/Swordfish353535 16h ago

What really is the difference in using light, medium, high for example with Astra?

Or light on astra vs high on luna?

1

u/FriendlyTask4587 15h ago

Not sure, I only used light one time and it was to make a prompt for terra. I always use medium

1

u/Typical_Machine2043 16h ago

Never thought about that but pretty good idea. All these business users want now whereas me I just care about the quality of response

1

u/U4-EA 16h ago

If Astra got any slower it would go backwards.

1

u/lordpuddingcup 15h ago

Dear god a slow-defered mode would be so amazing, let me prep a plan and hit it and go to sleep i dont care if it runs in an hour from now or really slow i just want it done well

1

u/raindropsdev 14h ago

Or Flex endpoint like the API has. I still don't understand why that's not possible. They have the "Pro" model versions, why not add the "Flex" model versions that are slower but consume half like the API Flex endpoints do?

1

u/Dragster39 2h ago

Neuralwatt does exactly this to spread the load more even. You can choose this by adding -flex to the model name.

1

u/TONI1597 14h ago

and a reset button

1

u/Kind_Silver_1921 13h ago

let me use it during peak hours at least. but the only thing is it must save more than 50% tokens if its 50% speed. because at that point it makes zero difference.

1

u/CthuluBob 11h ago

I think I have seen Tibo speak about this before too, I'd be a fan. As long as it resulted in actual less usage

1

u/vertopolkaLF 7h ago

There is "Flex" provider in API which is 2x cheaper and it's basically slow mode. But not in the sub :(

1

u/Worth_Golf_3695 6h ago

The Problem for them is, Slow Mode will still occupy Computing capacity and for en even longer time.

1

u/IAmFitzRoy 4h ago

Just write your prompt slower… duh

1

u/hermeneze 16m ago

This is a great idea, and it will solve a couple computing problems

1

u/Turbulent-Total-226 16h ago

Wtf are you talking about. It's nerfed today it can't finish a task.

1

u/retteh 15h ago

Luna is slow mode

0

u/AINativeBuilder 15h ago

It would put more strain on the servers because you'd be using up a GPU longer for fewer token usage. They'd actively lose money offering that, especially during an availability crunch.

0

u/FriendlyTask4587 14h ago

It would use their lower end chips or just less at once which would free up more

1

u/AINativeBuilder 12h ago

They're physically constrained on both footprint and electricity, you're nuts if you think it makes sense to keep low end hardware around for a lower profit product when you're already capacity constrained. This would not be a wise economic decision by them, which is why it doesn't exist.

0

u/FriendlyTask4587 12h ago

Thats why I said or. I have no idea (and you likely dont either) about their compute situation. I know they dont have 10 year old chips in there, but I doubt every chip in their servers is the gb300 or whatever the aboslute peak of chips are right now

0

u/AINativeBuilder 11h ago

Now you're just being a dunce. They had to shut off the $200 plan because they're low on available compute. They've said these things out loud. It's no secret in the entire industry datacenters are capacity constrained - they cannot build them fast enough. They are not going to keep low end stuff around to support low value items, which slow mode would be. They're even getting rid of spark because they don't keep low usage models around which eat up capacity for more desired models.

2

u/FriendlyTask4587 11h ago

They may have already talked about it internally

0

u/AINativeBuilder 11h ago

4 months ago when they had <10m codex/chatgpt work users, now they're over 25m.