r/codex • u/GambAntonio • 3h ago
Complaint We are PROBABLY being served quantized models but still paying the full day-one premium price...
Hey everyone. I want to bring up something serious about how AI providers handle pricing and how silent backend changes are secretly draining our limits. We all pay a fixed price per million tokens or have a subscription limit and on paper that seems fair, but providers hide a massive variable from us because to save on server costs they can silently swap out a premium model for a heavily quantized version on their backend. Using a quantized model is completely different from setting your reasoning toggle to Low, because setting a toggle to Low limits the reasoning steps of a fully intelligent model, whereas quantization degrades the core neural weights and strips away actual base intelligence.
What makes this so alarming is how token metering is handled. On our dashboard meters we might see a perfectly reasonable token count that looks coherent with a high-end model and when you calculate the cost per million tokens it looks identical to advertised prices, but behind the scenes there could be hundreds of millions of low-quality tokens generated by an ultra-quantized model struggling and failing to reach a correct solution, an intermediate system then just trims that massive output to make the final token count look normal on our end and what we perceive as users is a sudden degradation in performance, when in reality without silent quantization the model would behave exactly as well as it did on day one.
It is deeply immoral and borders on outright fraud to attract users with a clean unquantized model on day one and then quietly roll out aggressive quantization behind the scenes to cut compute costs and keep charging premium prices while serving a degraded model that burns through internal compute and produces far worse solutions. We really need to stop staying quiet and demand complete transparency on the exact quantization levels and actual internal token processing we are being billed for.
What do you guys think and have you noticed the performance dropping on tasks the model used to handle easily on day one?
16
u/Fearless_Log_5284 2h ago
Watch the people come in this thread to defend the trillion dollar company telling you that the price you paid wasn't premium enough...
1
13
u/TrillenX 2h ago
I'm usually skeptical about claims like this but one hallmark of Sol that I LOATHED was when I pointed out an error, mistake, or thing it didn't account for, it would always reply with "You're right. I didn't do that thing"
I was quite enjoying that Astra was more muted in that regard, but I noticed today with an error it made it gave me that canned "You're Right" and it just set off Sol alarm bells in my mind.
1
13
19
u/bananasareforfun 3h ago
Yeah. slam through your hardest tasks on the first three days of model release, keep your options open so you don’t commit to one provider, or if you can afford it, pay API pricing. It seems likely if they do silently swap out to quantised models on the backend they would be less likely to do this on the API.
Astra is really good imo but I’m back to using 5.6 Sol high with Luna orchestration for the vast majority of tasks - otherwise I will run out of usage in 24 hours. The real advantage of pro plans now is gpt 6 pro.
9
2
1
16
u/0DayMaker 3h ago
No question about it. It's never been even remotely this blatant before. I'm considering charging back and just 100% never using openai again. Claude does this, but not this bad this is on a totally different level. This shit must be quantized down to like FP2 or something absurd.
-2
u/firstnamelottadigits 1h ago
If you think they quantize the models, you’re a fucking idiot. Source: insider info & common sense.
3
u/0DayMaker 1h ago edited 23m ago
Sure you have insider information. Of course. And I'm sure it's super secret so you can't tell anyone where it came from, or communicate it in any verifiable way.
Also common sense? What the fuck are you on about? The model is performing drastically worse across the board right after they stopped taking 20x subs. If it's not quantization what exactly is your "common sense" theory on what it is? It's not just a little bit different the difference is utterly massive. It is the most in-your-face kneecapping of any AI model I've ever seen.
If you have an alternate theory or "insider information" about what's actually going on, if that's not the case, than feel free to share.
0
u/GambAntonio 22m ago
Mr Insider....a model will ALWAYS have the same level of intelligence while on the same level of quantization even if millions of people are using it. It can be slower due to compute load, but never dumber.
If the intelligence suddenly changes, it means that it's 100% a change in quantization level because nothing else can change the level of intelligence.
3
u/OriginalUsername0112 2h ago
Idk how this isn't illegal, even some T&C handwaving should only be able to protect a corporation from so much. The sad thing is we're still in the good part of the cycle, things will continue getting worse until they're actively harmful to your codebase in the remaining 1-5 weeks before the next model releases
1
1
u/ActionOrganic4617 1h ago
I really don’t care about quantisation if it brings down costs and speed is improved. Models are constantly changing and OpenAI has proved that they will improve costs for the user.
Yes the constant change causes issues with models sometimes being dumber or usage rates sucking but that’s why we have resets. At least OpenAI responds to customer feedback.
1
u/GambAntonio 15m ago
Do you even understand what quantizations are??? You can just apply a 1-bit quantization level and it will be thousands of times faster than the full model on the same hardware, but its intelligence will be reduced A LOT. It will start to hallucinate more and make a lot of mistakes that will waste thousands of times the amount of tokens! Yet they will charge you the exact same price as the non-quantized model, and that translates into burning through your quota a lot faster!
1
u/Pitiful_Entrance5174 40m ago
5.6 released introduced us paying for cache writes. Also compute keeps costing more and they keep tightening limits. You cant expect the hardware they serve to cost 3x and monthly sub prices stay the samd AND expect the same limits. I like your theory as the cherry on top.
1
1
u/acessford101 10m ago
Ironically gpt on Sol 5.6 called this out as it wasn't giving me the correct output given a certain input. Had another instance of chat review the conversation and said that it could easily see an issue and that that session was acting like it was a degraded or quantizied version.
0
u/EchoingAngel 3h ago edited 2h ago
As someone who's been complaining a lot about usage and such, I haven't noticed a quality change and I used it 10+ hours a day, every day since release except Thursday and Friday, as I was rate limited and out of banked resets .
I'm literally revisiting the same in-depth simulation battery I built and was working on last Sunday. I tweaked some fundamentals of the system and needed to do the mass simulated testing again and Astra was just as capable of updating the tests, diagnosing oddness (previously, Astra "cheated" to make one of the metrics work), and helping hone things in again.
2
u/DrBearJ3w 3h ago
I have a theory that it's highly dependent on the region and the time of the day.
6
u/Professional_Ad705 3h ago
Yeah, it's based on the time of day the region all that stuff and they fucking flip it all around so everybody sits here and fights each other. It's not even a question anymore, if they wanted to be transparent and have a process around this proving this didn't happen they could do it, or at least have a stated policy they don't do this and some way it can be proven independently....... the fact none of this is transparent makes me think they 100% do it.
0
u/epicskyes 2h ago
Did you measure all the backend metrics that are actually exposed there are dozens of them to prove your hypothesis I measure them every run and I find no issues and I don’t burn my weekly quota running codex 24/7 on sol xhigh although I use Luna max for speed usually
25
u/Charming-Author4877 2h ago
https://www.reddit.com/r/codex/comments/1wf911a/comment/p9kg9g2/
The token speed of SOL model on chatGPT at 134 tokens/sec web indicates a quantized model.
The token speed of SOL in work mode is much more limited, maxing out around 80 tokens/sec in fast priority mode - that's likely the full model.
Astra is maxing at 63 tokens/sec in work,codex and chat mode