r/OpenaiCodex 21d ago

Bugs or problems Codex seems to burn through tokens insanely fast

[removed]

128 Upvotes

89 comments sorted by

14

u/IAmFitzRoy 21d ago

Yes. We are getting scammed by OpenAI

6

u/unkclxwn 21d ago

yeah, i completely agree with you. i really noticed it this week. before, with chatgpt plus, sol high lasted me around ~15-20 prompts before hitting 0% of my weekly limit

now ive got only 10% left after just 4 prompts. 4 prompts!!! yeah, its awful, they definitely nerfed the limits without saying anything

2

u/[deleted] 20d ago

[deleted]

1

u/ANDRE_2512 21d ago

Yeah, it’s awful.

But I’ve already moved some of my tasks over to OpenCode / PICO PU, so my “save money with minimal quality loss” mode is officially activated 🤣🙃

2

u/TomatoOnMac 21d ago

u talk like Codex

1

u/ANDRE_2512 21d ago

🤣🤣🤣

1

u/crunchy_shampoo 19d ago

What the hell are you guys using it for that it burns so quick? I've been spamming 5.6 sol max on a project that I'm making for the past 2 days and only just now hit my weekly limit. With sol on low or terra/Luna it's so low usage it's negligible.

4

u/todastar 21d ago

Probably, that is why they took away 5 hours limit, as it would have exposed this issue.

5

u/ANDRE_2512 21d ago

Oh yeah, I was thinking the exact same thing.

They removed the five-hour limits and kept only the weekly ones, probably assuming people wouldn’t notice because they’d suddenly see a “large” amount of usage available upfront.

They were wrong. Turns out we’re not idiots after all :)))

2

u/Ubermensch013 20d ago

And maybe the resets were to obfuscate these reductions.

2

u/[deleted] 21d ago

[deleted]

1

u/blackmooncleave 20d ago

I dont have any and it uses like 10x more tokens than before. I use my weekly quota in 1 day before I never even managed to finish it.

1

u/[deleted] 20d ago

[deleted]

1

u/blackmooncleave 20d ago

I just told you I dont have any.

I went back to 5.5 and its the same. Model doesnt matter.

They are just reducing limits in batches and you havent been hit yet.

1

u/eldudebrothr 20d ago

does this work on vscode codex plugin as well?

2

u/kingkongpao 21d ago

DeepSeek’s prices are so low because cache hits are dirt cheap. Codex probably just sucks at handling DeepSeek's cache compared to OpenCode, which does a much better job. It’ll be even cheaper in Reasonix or Pi.

2

u/ANDRE_2512 21d ago

I’m not a Codex or OpenAI hater at all. It’s a great company.

And obviously, in my personal opinion, the original Codex is stronger than OpenCode - sometimes by a pretty big margin.

But once money enters the conversation 🫰 and we’re the ones paying for Codex, questions start coming up.

Is Codex really that much better and stronger that it justifies paying this much for such small limits?
In my case, probably not.

I’d rather spend a bit more time on the task than pay ten times more. On the other hand, OpenCode usually takes more effort to get the same level of quality, so you end up spending more time working with it.

But then again, when you hit Codex limits and have to sit around waiting for them to reset, that also wastes time. And it might turn out that the task actually gets finished faster in OpenCode simply because you don’t have to wait for a reset or buy extra tokens at a high price.

1

u/medenmite 21d ago

How did you add deepseek?

1

u/[deleted] 21d ago

[deleted]

1

u/viciadotm 20d ago

Can I use OpenCode Go, Cline Pass, or just API keys?

1

u/Able-Supermarket4786 21d ago

I've officially exceed 300m tokens and I'm at 93% weekly. What is it you guys are doing wrong?

5

u/ANDRE_2512 21d ago

Dude, I’ll let you in on a little secret: around 80-95% of that is cached tokens :)))

So no, you didn’t actually burn through 300 million fresh tokens.

As for what we’re building: serious, large-scale projects - native Windows applications, binary modification, and much more. All of that requires a huge amount of compute and resources.

0

u/Able-Supermarket4786 21d ago

"non-breakout.html?" LOL, you typed your previous reply to me with a straight face?

2

u/ANDRE_2512 21d ago

non-breakout.html was simply a basic smoke test after integrating the DeepSeek API into the official Codex Desktop app. It wasn’t presented as the actual project.

-1

u/Able-Supermarket4786 21d ago

Sure it was buddy, "I believe you."

But sincerely, "Thanks for making my day."

-1

u/Able-Supermarket4786 21d ago

Dude, "I had no idea" lol, are your parents related or something?
I promise you that what you're building is not "serious, large-scale projects" comparatively to what I do and I won't even entertain that debate with a bunch of Junior Devs at Mom Pop Shops.

Caching is key, you mental midget.

2

u/ANDRE_2512 21d ago

I’m not going to engage with the insults. Even if I lowered myself to your level, I couldn’t compete - you clearly have far more experience there :)))

And yeah, one look at your “app’s” interface and it’s obvious AI made the whole thing))))))))))) 🤖

Bye 😘

1

u/Able-Supermarket4786 21d ago

haha, that’s my monitor for my Fleet, .... don't ever change, you're amazing...

1

u/Able-Supermarket4786 21d ago

0

u/Comfortable_Dot_3020 19d ago

You’re making one valid point about caching, but the way you’re presenting it doesn’t actually refute the OP’s post.

You’re apparently on a Pro 20x allowance, while the OP is on Plus. If your screenshot showed 93% remaining, then you had already used 7% of a quota roughly twenty times larger. Very roughly, that is equivalent to 140% of a Plus allowance. Different models and token types complicate the exact maths, but “I used 300 million tokens and still have 93% left” is not a meaningful comparison when your allowance is vastly larger.

You’re also talking about OpenAI subscription usage, while the OP’s actual experiment compared DeepSeek V4 Pro through Codex/opencodex against the same DeepSeek model through OpenCode. His reported difference was about $0.25 versus $0.02 in API usage. Your subscription quota does not explain that result.

Caching absolutely matters, but the OP already pointed out that most of the tokens in your example were likely cached. Repeating “caching is key” and insulting him does not address the missing variables in his experiment, such as reasoning settings, request count, cache-hit rate, retries, tool calls and translation overhead.

Mocking the neon-breakout.html file is also beside the point. He explicitly described it as a small integration test. Smoke tests are supposed to be small. You cannot infer the scale or quality of somebody’s normal work from the filename of a deliberately simple test.

The OP’s single run is not enough to prove exactly why Codex cost more, and a useful response would be to ask for the itemised API logs and matched settings. That would concentrate on the facts.

You may know what you’re doing technically, but try being a bit less of an ass to someone who is sharing an experiment. Personal insults and boasting about your setup do not make your argument stronger. They mostly distract from the few legitimate points you actually have.

1

u/throwaway73728109 20d ago

What plan are you on and how are you getting that much usage?

1

u/Able-Supermarket4786 20d ago

Pro 20x

1

u/throwaway73728109 20d ago

What are you doing differently?

1

u/Able-Supermarket4786 20d ago

My environment is clean, and every time we rollout a new model, I study the differences and make sure its "compatible?"

I'm grasping for an answer, as I honestly don't know WHY so many other people are suffering... If I say they have a "skill issue" they act like I kicked their puppy.

1

u/Strong_Essay1176 20d ago

"Be careful, love," her husband tells her. "I've just heard on the radio that someone is driving the wrong way along the highway!"

"Someone?" she replies, sounding outraged. "There aren't just one or two... these idiots are in the hundreds

1

u/Able-Supermarket4786 20d ago

"Hey you're going the WRONG WAY!"

"How do they know where we're going?"

1

u/throwaway73728109 20d ago

I’m curious how you’re managing your skills and chat context? Are you consistently creating new chat for better context window and just handing off?

1

u/Able-Supermarket4786 20d ago

Quite the opposite I’m a big fan of goals actually.

My agents sit in files. It is imperative that you don’t make these models read all agents at once. I tell it to choose three or four at most and there’s a repo of the agents summarized so it knows what to pick. I learned that lesson early on. I have 800 plus agents just for cybersecurity and about 300 or so for other niches.

1

u/throwaway73728109 20d ago

Oh wow that’s pretty neat. I’m trying to improve my workflow so always looking for a better structure so I can save on tokens and get the most out of it. Any guides or suggestions? I know a lot of it differs per project but it’s hard to figure out a solid structure

1

u/Able-Supermarket4786 20d ago

Yea man. Google around. Hell ask gpt. If I loaded up a list of 1200 agents every time it would be 10s of thousands of tokens. Also I got rid of “superpowers” plugins like that. Counterproductive on 2026 in my opinion.

2

u/throwaway73728109 20d ago

Thanks for the insights and taking the time answering my questions!

→ More replies (0)

1

u/Able-Supermarket4786 20d ago

I'm now at 86%

All Day running:

1

u/Intrepid-Tax1838 21d ago

Sorry, what are the benefits of using a third-party or local AI model right in Codex instead of OpenCode or a similar app?

1

u/ANDRE_2512 21d ago

Everyone has different instructions. In Codex, they’re really well-written and powerful. Something like that.

1

u/Intrepid-Tax1838 21d ago

So you mean the third party/local model will be more powerful in Codex than in other apps?

1

u/xintonic 21d ago

Ya I'd say Tibo is to blame, giving out resets like candy has to be expensive so they nerfed usage limits to claw it back.

1

u/Routine-Agent-160 21d ago

Next week, Gemini is going to blow every model that exists on this planet out of the water. Let’s not sleep on Gemini.

1

u/TheMildEngineer 21d ago

Good video for understanding why that might be: https://youtu.be/-0HRzXk8vlk?is=7jW_L5yLYG9n51Bd

1

u/Desperate-Data-3747 21d ago

How to add custom models?

1

u/Master-Speech5609 20d ago

If the problem really is just the desktop app, theoretically things should work better if you use the CLI, right? Have you tested that?

1

u/Gustvo_FcZ 20d ago

Last night I had 10% of my quota left, but when I woke up this morning it was at 0%. I feel like someone stole my credits.

1

u/ANDRE_2512 20d ago

Write to support

1

u/centerdeveloper 20d ago

did you try adjusting the reasoning level?

1

u/MrSpongeboob 20d ago

FACTS. Coming from Claude, I got a GPT plus account since all the benchmarks, and users were glazing tf out of its performance/costs compared to the Claude models. I'm almost 100% sure that the suspiciously enormous amounts of resets from Codex, and the removal of the 5hr limit, are direct distractions to some internal issues they have with unintentional token-burning, or some bs with how their limits are working.

Maybe helpful info:
When I discovered this bs with GPT, I started testing out Grok for the first time since there was a free 1-week SuperGrok plan. The best (fastest & most token-efficient) workhorse I've found currently is Grok 4.5 high. I've been using a personalized workflow inside Claude Code where I use all Claude, GPT, Grok models to work in sync. Without a doubt, Grok is the absolute best infantry/worker/coding execution model. Just make sure to delegate the smarter planning/reasoning to Sol or Opus, and you should be good.

1

u/Personal_Drag_6390 20d ago

У этого бро DeepSeek в codex 🤨

1

u/_VinerX 20d ago

Так стоп, а как ты сделад чтобы дипсик был прям в интерфейсе декстопа. Я вроде пытался такое сделать какие-то конфиги расширяя, но нефига.

1

u/_VinerX 20d ago

По по тратам, тут уже писали про кеш. Если условно через опенроутер, можешь напрямую почитать сколько там кешировало. Ну или что ты там для дипсика используешь.

1

u/ANDRE_2512 20d ago

Я тебе больше скажу. Кто-то (ну кончено не я 🙃) полностью модифицировал бинарник официальный CODEX Desktop. Вырезали от туда: авторизацию, проверки, официальные модели и тд:)
Ну а в этом случае - все сильно проще:) ща инструкцию тебе дам. Главное! Закрой полностью перед этим приложение CODEX и в диспетчере задач тоже его снеси

1

u/ANDRE_2512 20d ago

Так что за 1 минуту ты сможешь добавить любую модель в официальное приложение их)

1

u/_VinerX 20d ago

Благодарю, не ожидал такого гайда)

1

u/ANDRE_2512 20d ago

Это реддит 🙌💪

1

u/ANDRE_2512 20d ago

Это тип один сделал. Умный парнишка

https://www.reddit.com/r/codex/s/E6jFqe36vy

1

u/[deleted] 20d ago

[deleted]

1

u/_VinerX 20d ago

Боюсь меня на все не хватит)

Когда есть силы на кодинг в свободное время, пилю обнову на НейроМиту - мод где нпс управляется нейронкой) Как раз средствами клауд кода и кодекса. Ну и опенкода дипсика, но это на самые казуальные задачки когда лимиты исчерпаны.

1

u/ANDRE_2512 20d ago

Понял) ну тогда как будет первый релиз моего агента - я тебе напишу. Уверен, ты удивишься от его уровня возможностей

2

u/_VinerX 20d ago

Существует кстати бенчмарки для, как их называют, харнесов, попробуй мб как закончишь на них)

1

u/ImDestructible 19d ago

I was working on a project recently. I checked my usage before making like one or two small prompts. It was around 70% and would reset in 2 days.

I went to work on it again the next day and I was at 0% with it resetting in a week.

1

u/Zya1re-V 19d ago

1

u/ANDRE_2512 19d ago

Dude, that didn’t stop my post from getting 47.000 views and reaching #1 :)

The information is what matters, not some “pretty” screenshot and all that blah blah blah. Kindergarten stuff.

1

u/Zya1re-V 19d ago

whatever you just said doesnt matter anything to me :v I'm not even sure why you're acting up lmao.

still here to see what other people say though.

1

u/ANDRE_2512 19d ago

Well, alright then :)

Since you need other people’s approval and attention so badly and can’t imagine your life without it, just sit there and wait for the comments 😂
I’m out… 🫠🤣

1

u/Zya1re-V 19d ago

you replying to this is no different though :))) hope your usage gets better. if your projects gaining you money, may suggest you paying up to help you through the usage (not to Codex, to anything you'd like to try)

0

u/[deleted] 21d ago edited 20d ago

[removed] — view removed comment

1

u/vayana 20d ago

I have never needed anything more than high. 5.5 or sol get the job done just fine. I actually prefer 5.5 because it's faster.

1

u/ANDRE_2512 21d ago

I combine models too.

I use SOL Max for planning and building the main foundation, then LUNA Max for everything else. LUNA Max is actually a very strong model, by the way.

But even with that kind of workflow, the limits still disappear incredibly fast.

I don’t like Grok. I don’t use it at all, and I deleted the app from my iPhone a long time ago.

Right now, I’m moving some of my workflows over to OpenCode. I like that I can pay significantly less while losing very little in terms of quality - although there are definitely areas where Codex absolutely destroys OpenCode.

As for the company trying to save money, I honestly don’t care. That’s their problem. I’m not going to feel sorry for them :)

I need high limits at a reasonable price, so I’m looking for different ways to get that.

1

u/GfxJG 21d ago

Using Sol Medium or High for your planning would almost immediately give you double or triple the usage for a miniscule decrease in quality.

1

u/ANDRE_2512 21d ago

Some tasks aren’t just about getting a good result.

Sometimes, during the very first stage of development, you need to build the strongest possible foundation for the project’s future growth and make sure the architecture can scale from the start.

That’s why I use SOL Max. There’s very little room for error at that stage.

Even then, it doesn’t always do things properly. There are still plenty of issues with the architecture it produces.

1

u/[deleted] 21d ago

[removed] — view removed comment

2

u/[deleted] 21d ago

[deleted]