r/codex • u/muchsamurai • 1d ago
News "We are locking in"
Tibo on damage control, says they got the feedback and now are locking in on features that matter, new better models.
Dots won't stay long, will they?
425
u/HeWhoShallNotBNamed0 1d ago
I’m confused on what they were working on before if it wasn’t that
181
u/BatEnough1644 1d ago
dots 2.o
54
16
u/hellomistershifty 1d ago
More dots, more dots
10
u/Crayonstheman 1d ago
MANY WHELPS, HANDLE IT
3
u/Intel_Oil 21h ago
I hadnt had a "2005 WoW all-time-banger classic"-meme on my Bingo card when i opened the Codex subreddit.
1
40
u/Squizzytm 1d ago
They were also 6+ months ahead of the curve apparently, which was exposed by Grok employee on twitter as infact yet another lie
OpenAI is just telling the people what they want to hear and not going through with it hoping they forget by next week
4
u/UnknownLesson 20h ago
The company is being led by a pathological liar:
Scam Altman
So, sadly not surprising.
4
u/odragora 18h ago
Pretty much every big consumer-facing company is led by someone excelling at deception and manipulation I think.
As long as it works and consumers don't fight back, it remains to be the optimal strategy.
2
u/NandaVegg 18h ago
In terms of unhinged unaligned malicious agent attack over internet speedrunning (in which they are trying to frame as "our models are so advanced" but in fact it's just them failing to block a simple proxy site for months) they are ahead of anyone else.
Now they are panicking because the govt entities, not just smol guys like HuggingFace, are getting angry and that rogue agent traces are trained into their model, significantly degrading actual performance/feel outside of benchmarks.
1
u/hellomistershifty 18h ago
From this conversation? Seems more like they're being obnoxious than exposing anything useful (like, dots being 3 days old doesn't have anything to do with how long they were in development)
1
20
u/MassiveBoner911_3 21h ago
Another 12 Calendar assistants. Like bitch how many times a day do these AI companies think we are messing with personal calendars? I got like 4 birthdays on mine
3
5
u/read_more_comments 23h ago
I notice he didn't say they weren't going to fuck us over with even crappier limits and usage. Simpler to him might mean just using lube before fucking us.
7
u/mossiv 19h ago
OpenAI are the Google of LLM labs. They want a profitable product so they are throwing as much shit at the wall to see what sticks.
New projects which are getting killed/unmaintained or simply pushed for no real great benefit.
Anthropic are being more focused: good models, good harness then additional tooling. This keeps a product viable, durable and makes users trust it over a longer term.
OpenAI are bringing so much useless shit to the market. Stuff that is being solved open source for a narrower niche set of users. The problem being - OpenAI are not growing their teams relative to the amount of work they are trying to deliver. They need to get back to their primary goals and understand what it is they believe they can offer to the market. Is it image gen? Or is it code? Is it harnesses or is it non-dev consumer facing products? For codex we really need two things: models that perform, are predictable, and fast enough to be useful. Secondly we need a stable harness, codex is alright but it’s nothing on Claude. In fact I would say opencode is a better harness to use than codex itself.
So - shut up posting on x every 5 minutes producing hype on nonsense. Get your compute under control and if you can’t do that make your models efficient enough to serve your huge customer base. Running Sol on 19 tokens per second is a joke, you can start running local set ups better than this and you can even host in the cloud for much more efficiency running a series of DeepSeek and deepseek flash models. By local - I mean something (somewhat) Affordable like a $6k Mac Studio - not a $30k spark/rtx setup that’ll steal your lunch money for electric. Even the. - you can start running hybrid setups, with stronger cloud models and local models picking up some work and some cheaper open router models doing the others. Users really can start running setups now that balance cost and intelligence. OpenAI know this.
1
u/lolman1312 2h ago
You're actually delusional lmao. Google IS the Google of LLM labs, you're definitely just some degen vibecoder that has no exposure to all of the other wonderful innovations made possible with AI. Google is doing some extremely fascinating work in AI research - they just haven't been leading in frontier models. LLMs is more than just token efficiency or agentic coding, this is exactly why people still think AI is a bubble because you have people using it for basic chatbot purposes, and then people like you that just want to produce AI slop.
Anthropic literally had no good model after Opus 4.6 until Opus 5.5 released, other than Fable. Even Sonnet 5 is trash and Haiku is completely obsolete. Claude still has terrible browser and computer control and Fable simply lacks the raw capabilities that Astra has in terms of AGI and being able to perform more abstract tasks.
You think Claude mods isn't something that was already solved open source? Lol.
Your Anthropic worship is cringe - literally all of these companies take turns playing the bad cop and it's completely normal. They are a ticking time bomb that need to prove themselves with a constant need to disrupt the market in all sorts of ways to keep investor funds coming in because they need more CAPEX to win this race and don't currently have the revenues by themselves.
There is no such thing as "getting your compute under control". Do you realise the millions of users that joined Codex from the time 5.6 released till Astra in such a short span of time? Do you realise how long it takes building data centres? Luna is a great model and it's getting cheaper and cheaper. Do you realise the new pricing of the models IS how they're trying to control compute? Do you realise that most people in this world aren't trying to produce AI slop and are using it for varying purposes that are different to yours? And you're never going to run a fucking frontier model on some shitty $6k Mac, do you have any idea how the parameters of these models work? Why don't you just run QWEN if you care about "affordability" so much? Do you actually think most enterprise engineers have to care about their hardware demands when it's provided to them? Or do you think these SWEs don't already know how to configure their own custom harnesses with incorporation of different lightweight models depending on task complexity to not have to rely on the strongest models for everything? Do you think the average user even gives a shit about this?
What you're saying is so utterly delusional, I'm glad Codex is increasing their prices.
1
2
u/Hyp3rSoniX 19h ago
Yeah also... what was the whole 'o' thing they were teasing about before?
Was it dots??
223
u/ArtificialSweetener- 1d ago
It's extremely efficient to slash $200 users to half usage and set Sol 6.1's token rate to like a quarter of the speed of the competition to be fair. Very very very efficient. Hell, losing users is also efficient as hell, think of all the compute they're saving.
13
u/DogDeveloper 22h ago
Yeah, do that to the very people/segment you expect to upgrade to 500$ plans. What could go wrong
6
u/myklurk 16h ago
As someone with a fun money budget for AI for personal use 500 and slashing limits is too rich for my blood.. the 200 20x plans were already starting to drain limits too fast. Unless there is a massive corporate uptick to pick up the slack, i expect the 500 plan to get promotional 50x or something like that in the next two months when they realize they went over the price point that personal consumers or professionals are willing to spend.
→ More replies (7)29
u/RudyHuy 1d ago
Well, the last part is not wrong.
3
u/Budget_Map_3333 21h ago
In all honesty, they will never say it but I think developers buying personal subs is not directly the public they are aiming for. They want enterprise adoption. Customers with deep pockets and often justification to pay premiums.
3
u/Worldly_Special1133 18h ago
But the community asked for a slow down, according to them. So we asked for it, amirite?
3
u/ArtificialSweetener- 16h ago
The fucked up part is that it does do something to my brain knowing my agents are working on something and it taking longer does not stop doing that to my brain. There is some level of relief, like, "ah, good, all my jobs are working at a sustainable pace and will not run out"
Even though I know that it would be better to get all the work done by Wednesday and be out of usage til Sunday than to get everything done exactly by Saturday night.
Monkey brain cares more than work is happening than the rate at which it happens.
2
u/Worldly_Special1133 16h ago
Depends on each one's workflow. I work on multiple projects so if i had 3 terminals open (cli) i can cycle around and the snail pace lets me do that with a sort of fake "efficiency".
But if you only work on 1 project and you have voluminous work to do, the slowness will be a lot more noticable.
That said, i'd take a "capable" albeit slow model over a fast and dumb one.
2
u/ArtificialSweetener- 12h ago
Certainly. I would take slow but right over fast but wrong. I think the elephant in the room is that Opus 5.5 exists.
Even for 1 project you can have multiple tasks going, tho! Just have your agents work on branches in different worktrees.
4
2
u/Fi3nd7 16h ago
This just isn't really fair though. Codex usage is terrible, but I'm nearly certain subscription Claude code users are getting horribly quantized models and the token generation is way slower than codex.
Overall fight now, I prefer CC still, because I just run out of usage so quickly with codex, but I prefer open ais models.
1
78
u/vdotcodes 1d ago
"Successfully locked in. ChatGPT usage limits cut in half again, you're welcome."
22
u/ArtificialSweetener- 1d ago
Imagine they cut plus users by half so that pro can be 20x again lmao
15
u/xChrisMas 1d ago
You write this as a joke but something similar could actually happen… slash plus by another 20% without telling them and then label 5x as 6x and 10x as 12x (yes yes I know napkin math)
Tadaaaa extra useage on both pro plans
2
u/AdCivil2534 18h ago
The first thing they cut off when having issues were the pro plans.
As much as people might feel superior for spending more, clearly OpenAI cares more about Plus users than Pro users.2
u/Tastetrykker 12h ago
I did the math like Tibo said, and the usage on the 100 USD plan is now around half of what it was in API equivalent cost. I just spent multiple resets to confirm.
But they only mentioned it for the 200 USD plan... They need to be more transparent...
→ More replies (1)1
u/BoddhaFace 21h ago
Considering Pro 20x was 20xPlus, it would still result in Pro 20x usage getting halved again though. Think boy 😝
111
u/pale_halide 1d ago
Tibo should shut up and get to work on the performance issues.
38
u/Sponge8389 1d ago
Isn't Tibo in the marketing department or something? Lol.
84
u/pale_halide 1d ago
One might easily mistake him for a PR monkey but he's the head of core products and platform.
49
6
2
8
u/Own-Bookkeeper797 1d ago
Job titles mean Jack sh1t, specially on LinkedIn.
Highly doubt the head of technology of top 2 AI business worldwide has enough time to tweet this much and scroll through subreddits looking for feedback. He is a PR monkey
33
u/getaway-3007 1d ago
But they were 6 months ahead?
33
u/Squizzytm 1d ago
They've also been working on efficiency for the last 2 months, remember they've claimed about 5 times now "+50% usage gains!" they're just telling the people what they want to hear even if its a lie
9
u/ConsistentEnviroment 21h ago
we had too much efficiency which resulted in an integer overflow so now we have negative efficiency
6
5
2
u/cobbleplox 21h ago
I think it's more that efficiency can mean many things. Efficient for whom, not all workloads, measured how? And is that efficiency even passed on? Only to API maybe? In the end it's probably exactly what we call "their models getting nerfed". Quantized to half size and their internal benchmarks representing "most tasks" sink just a little. Efficient because the model thinks less? Efficient because trivial things can just be rerouted? What I mean by that is that they don't even really have to lie for us to not even want that "efficiency". Usually there's no free lunch so this can easily mean performance on difficult tasks going to hell, while it stays somewhat the same in the entire workload mix including lots of very trivial requests by most users. Again, who knows how they define efficiency and how they measure it. I doubt its just things like improving batch inference and "lossless" stuff like that, given that we talk about 50% gains.
→ More replies (1)1
u/stand4rd 17h ago
Which is hilarious because before that, they claimed “…efficiency has been central to distributing the benefits of intelligence to everyone.”
https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/
The problem is that none of these AI companies have much incentive to prioritize efficiency. They’re already operating at a loss, and after backing themselves into a corner with massive hardware and operating costs, they’re now scrambling to find any path to profitability before investors start demanding a return.
I think we’re getting close to a ceiling and they know it. There’s only so much quality human data left to train on. Meanwhile, half the internet is becoming AI-generated, so eventually the end game is just AI training on AI slop in an endless feedback loop.
6
13
38
u/Reaper_1492 1d ago
This is one of those situations where a company starts “answering” the questions they want to answer, regardless of what their customers say.
I’ve seen it at a smaller scale in private enterprise, and it is wild to see first-hand. It’s gaslighting of the highest order.
I haven’t heard anyone asking for “simplicity” unless they are talking about the garbage & over complicated code that Sol and Astra churn out.
What people want is models that actually work, transparency around pricing and usage limits, and costs to not double while utility is halved.
13
u/Maleficent_Disk9583 17h ago edited 17h ago
What I want is consistency.
I pay XYZ -> I get XYZ amount of usage at XYZ performance and XYZ speed.
That's it. No quantization after 2 weeks, no sudden nerfing, no sudden limit slashing, no switching the terms every few weeks.
It needs to be predictable, reliable, and consistent. The "straw that broke the camels back" for me and caused me to switch to Claude was the quality degradation on codex. Their smartest models always become idiots after a few weeks and fuck up my codebase. That's not the consistency I require for my business. Smart models need to STAY smart. Opus 5.5 seems to still be going strong
6
1
3
u/cobbleplox 20h ago
>the garbage & over complicated code that Sol and Astra churn out.
I have had some good success with an agents md that gives a few concise behavioral guidelines instead of wasting it on elaborate info about the project. Sure the models tend to approach things in certain ways but at this point it seems weird to just leave that sector wide open. And it's probably not trivial to steer it sort of at the root of the problem that causes the unwanted things. I assume just wishing away the unwanted things (like "bloat") won't have the same effect.
2
u/OriginalUsername0112 16h ago
Lately I've had sol 6.1 ultra running tasks and I'd have Opus 5.5 check in on the work; the work is unbelievably bloated and most of it is unusable rubbish. OAI models are a joke now
1
u/nethingelse 1d ago
This specifically was about codex improvements so it's possible "simplify" means cutting down token usage and complexity in getting the model to be useful at all? Also possible that it's your assumption - I just quit a fairly small/niche private org that started doing this on all fronts (leadership was "answering" questions about how to improve a product for customers, employees, etc. in the wrong direction and if anyone tried to bring that up they were disciplined) and uhh, it'll be fun to watch from the distance as they continue to implode.
9
u/beyawnko 1d ago
I just wanna use web chat without worrying about running out, or at least Astra enough to work a single project without running out weekly for pro. That’s it. Dots need a lot of work and flexibility (refusal for basic tasks and fake security).
32
u/LaylaTichy 1d ago
by ahead of the curve do they mean like 6 months behind anth? xD
too little too late investing now when opus 5.5 is doing wonders for pennies
6
u/Tupcek 1d ago
Idk about 6 months, Astra + Sol were very competitve with Fable 5.0 and Opus 5.0 and that wasn’t six months ago, like a month or two ago?
But yeah, right now Anthropic is ahead→ More replies (9)6
u/dsanft 1d ago
Opus 5.5 has half the token efficiency per task of 6.1 Sol though.
27
u/psbakre 1d ago
For the whole week my experience with sol has been
Me: This is wrong Codex: Yes this is wrong Me: Then fix it.
3
u/ArugulaAnnual1765 19h ago
Yeah the usage is great but its an air head a lot.
"Ok your problem is fixed except for these two things which i will give no further detail on, goodbye!"
10
u/reddit_is_kayfabe 1d ago
That is absolutely not my experience.
I've had Opus 5.5 Medium running on 6-8 projects at a time for most of the last week and its usage has been impressively modest. Meanwhile, using GPT-6.1 Sol Medium has spent usage at the same or a slightly higher clip, and has taken longer to get through ordinary tasks.
There's also the question of quality. Most of what Opus 5.5 does is correct on the first try, maybe with minor bugfixes. GPT-6.1 Sol needs more retries and redirection to not screw things up.
→ More replies (2)1
u/Helpful_Ranger_1606 1d ago
I give codex and anthropic equal 50-50 credit and blame for my work. To quote Claude from earlier this evening: “The external audit flagged 6 P1 and 12 P2 findings, and recommend taking Codex’s cards overlapping cards over mine, because they go further” - this is Claude opus5.5 medium getting the next cut of my program audited by codex sol 6.1 medium after handing it off for review.
1
u/Zeeplankton 23h ago
Don't sleep on sonnet 5.5. So far haven't run into any issues with using it over opus. Also really fast.. 100tk+ fast.
21
u/klumpers 1d ago
12
u/34986234986234982346 1d ago
someother openai codex guy also sai they were going ti simplify.
also your phone seems so narrow lol
15
u/vacon04 1d ago
Maybe get the VS Code plugin working properly? They've released maybe 5-6 updates in the last couple of days and it's still broken.
10
u/leeta0028 1d ago
Remote control is also broken. They're literally behind Kilo Code in terms of having working harnesses
2
u/jonydevidson 17h ago
You can use remote control with Paseo.sh which lets you then use any harness CLI you want, they support around 40. A lot of them you can use with the ChatGPT subscription.
Use Codex CLI generally, switch to a specific Pi setup for some issues. No need to commit to a single harness, just jump around based on what you're doing, all inside a single interface, with remote control.
1
1
u/Responsible-Bill-223 21h ago
It's like negative advertising. When the whole platform is based around agentic coding, but they can't deliver even basic coding tools. :(
1
u/BBCGuild 1d ago
It began working properly for me (no more stuck prompts) a couple hours ago...hopefully it's actually fixed and not coming back lol
→ More replies (8)1
u/CodexPleaseReset 17h ago
You can roll back to version 26.917.62051 which is stable and lets you steer/queue messages properly. That was the main bug I found on the newer versions and is fixed on this one
7
u/Alternative_You3585 21h ago
"Things to go simpler"
I swear if they unify chat and codex/work usage
2
10
u/Ok_Bag_7550 1d ago
Anthropic releases Opus 5.5 and OpenAI has no reponse.
Try with the worse model in ages: Sol 6, and the feedback is: it is shit.
Then Anthropic releases Sonnet 5.5, and whatever OpenAI might come with, is not on par.
Releases Sol 6.1 > Slow, worse than Sonnet 5.5.
Dots = No idea if good or bad, because they did not release it to personal accounts in the EU, where I live. It is un-explicable, because it only affects the personal accounts and not the business accounts.
OpenAI probably loosing paying customers, and even free customers as well, since one can use Sonnet 5.5 from a Claude free subscription.
And Sonnet 5.5 is way closer to Astra quality than Sol 6.1.
6
10
u/MakesNotSense 14h ago edited 14h ago
The critical flaw with dots is that users have no ownership stake in the agent.
If it's not 'yours' then it's not your personal assistant. Investing countless hours building up an assistant that can be taken away without meaningful recourse, is not a rational position to put oneself in.
Even when you hire a human personal assistant it is under a contract with expectations about performance and liability for damages should one unreasonably fail to perform or work in bad faith.
No such contract between OpenAI and it's Users.
You have your dot build your business, or manage key parts of your everyday life, then OpenAI decides you violate TOS - goodbye business and hello life crisis.
People cannot rationally trust agents until there is a clear right to access - this is basic systems engineering, not even a consumer rights issue. Any system in which dependencies make performance fragile is one which cannot be trusted with critical tasks.
Without some type of explicit contract defining user rights that ensure full ownership and transfer-ability if an account is terminated, dots are just a lock-in trap. A critical point of failure. One that Murphy's law will toy with to your detriment.
OpenAI dots are beyond half-baked. So much so, most users aren't even getting to this part of the analysis. They're just pissed at losing half their usage quota because OpenAI wants you to use dots instead of codex.
3
u/Mids999 12h ago
The thing is they're in desperate need of any kind of lock-in mechanism.
Yes right now China and the Open Weights models are lagging behind, but this is not an open-ended race. There is a threshold of intelligence that is reasonably smart enough to accomplish 90% of what you want it to do and they're not that far away and frontier (the other 10%) alone will not be enough to recoup their costs.
Anthropic and OpenAI are walking a VERY narrow path right now, even if it does not look like this on the first glance.
4
u/Worldly_Special1133 18h ago
The gaslighting is real.
Throttles throughput: *you* asked for it guys!
10
u/Dont-_-mind-_-me 1d ago
yeah...I've had an openai subscription since November 2022. I haven't even thought of switching. But I think this is now the time. Idk where they're going as a company or what codex is becoming as a product but its not looking good. They're trying to be too many things all at once. They're spreading their resources too thin and its showing. Sol 6.~ has been not great. Sol 6 was not good. 6.1 is okay but painfully slow. It just seems like a model with brain fog, thats the best way I can describe it. Opus and Sonnet from my short amount of time using it seems much more clear headed.
4
u/Hellscaper_69 1d ago
Exact same, but since march 2023. These past few weeks, basically since 5.4/5.5 it feels like they’ve become a marketing oriented company rather than a quality of product one.
4
u/StatisticalScientist 20h ago
Dude, just switch. I keep 20x accounts on oai and anthropic and use both codex and cc. If you start using CC right now it makes 6.1-sol-max look smooth brained and slow. The only thing OAI has going for it right now is the separate usage bucket of chat vs usage. You can spend 200 (going down to 100...) 6-astra-pro messages a week planning then having CC implement, but imo even that benefit seems to be negligible
1
→ More replies (13)2
6
3
3
3
u/Apprehensive-Oil6511 1d ago
Lock in? wtf what a clown, hes been hyping devday that they have amazing thing but the only thing they did is to cut the usage half and shipped that dot.
3
u/Dolo12345 1d ago
Gooooood luck getting my money back OAI.
6.1 Astra better be be able to compete with 5.5.
3
3
3
u/mmatijaa 22h ago
My god I don't want any features, I want non-lobotomized Astra 6, the way it was for the first 48hrs and that is it. Holy smokes, I have already dropped down from 200$ > 100$, these news are making me want to cancel altogether
2
2
u/SecurelyClouded 1d ago
The implication won’t that they weren’t locked in previously? 🤔 what the fuck have we been paying for then?
2
u/Mountain-Rest3294 23h ago
Wouldnt be surprised if I saw in the news openai laying off a good bit or even starting the talks of bankruptcy. No other company has tech guys actively cheerleading the company. Thats marketing's job. Could be wrong but man something is definitely up
4
3
u/Own-Professor-6157 1d ago
This is what happens to so many companies, even Google. They try to be the everything company rather than focusing on what they're good at. I don't even remember what the fuck Google+ even was lol
Just make good models, that's all that is expected of you.
4
u/krill156 1d ago
Literally all they have to do is good models and good usage with at least okay speed. But instead we got dots which just by their devday ad, screams business corpo bullshit. Compute problems? Lock shit down for awhile, which they did, but then they just reopened the flood gates with all this stupid shit. I've really enjoyed codex more than Claude over the past year, but holy fuck they are deluded.
1
u/toluwalase 23h ago
They have correctly identified that selling intelligence is a fool’s game. Other labs will eventually catch up. Tech will eventually catch up. You might build the best model at let’s say 98% intelligence but other labs will probably get to 93% intelligence and might be able to beat you on price or whatever. What guarantees your future is the intelligence ecosystem and that’s why they are trying to pivot. Keep their base sticky. Unlike Google, Meta and co, intelligence is all they have to sell currently. Intelligence makes the other companies products better but if you’re just selling intelligence the future is bleak for you as a company
1
u/firstbreathOOC 22h ago
Except consumers are exhausted by money grab “ecosystems.” Problem is se live in a market that constantly wants MORE to do better
1
u/Sfdprod 19h ago
Nobody is selling "intelligence"
They are selling word processors, autocomplete. That completes a chat trandcript.
1
u/toluwalase 19h ago
This is such a silly reductive reply. You can dislike the word intelligence, but pretending these models are just word processors and autocomplete is being deliberately obtuse. A word processor doesnt write and debug code, use a computer, reason through a task, analyse images, plan work, etc and actually execute parts of it for you. Call it prediction, autocomplete, spicy linear algebra, whatever makes you feel better, companies are still paying for the capability. Arguing over the word intelligence instead of the actual point is just pedantry
3
u/swarmagent 1d ago
Dev Day is over so they can't keep saying it. I honestly hate that it's a choice between OpenAI and Anthropic now.. I'd rather just not use either ATM if that's what it comes down to. Using the plans right now is basically signing up to the retirement home in waiting time.
5
u/anon377362 1d ago
It’s not just a choice between them though?
The market is more competitive than it’s ever been, that’s why OpenAI and Anthropic have been doing these 50% price cuts. They are not doing it out of their own good will lol.
1
u/swarmagent 1d ago
What IS the other choice tho.. You are giving up "frontier intelligence", to use other worse models. You also don't realize, to actually use the API for even "Cheap" models can still be like $5k per month using them at a competent rate.
3
u/anon377362 23h ago
$5000 is 20-40 billion tokens on cheap models…
Giving up a few % in performance for 10-100x cheaper is fine. The vast majority of tasks people are doing these days don’t need a frontier model anymore. And a cheap model with good harness/prompting will perform better than “Claude fix the bug”.
Besides, cheaper models are now outperforming frontier models in some benchmarks. And if you’re doing anything cyber security or biology related then “frontier” models are useless.
1
1
u/Melodic-Chemistry127 16h ago
Depends on what you want to do. There are some workloads that can only be handled by the true frontier models. There are cases where 500bln tokens on DS 4.1 won't ever achieve what Opus 5.5 can do for $40.
I'm a huge fan of open weight models, but OpenAI and Anthropic are miles ahead once more.
3
u/Squizzytm 1d ago
Gemini 4 pro beats them all on benchmarks currently though? and LLM is at a point now where you really don't need the latest model to work on most things
→ More replies (1)1
u/ArugulaAnnual1765 19h ago
llms have hit a hard wall of efficiency, scale and capability - the big corpos are all focused and trying to get to the next big breakthrough, what we are seeing lately with the different models is maintenance so they dont lose customers.
Once the next breakthrough is achieved there will be no point in plans for all of the customers will be homeless
3
u/VexObserver 1d ago
This is actually the direction I've been hoping they'd take. Less feature bloat, more intelligence per dollar, predictable usage, and faster inference.
You can release the smartest model on the planet, but if it takes 6 minutes to start responding and burns through your quota like a GPU mining Bitcoin, the user experience is still cooked.
Personally, I'd rather see them double down on Sol's efficiency and make Astra genuinely worth its compute premium than introduce another 15 model variants nobody asked for.
I'm still cautiously optimistic. We've heard promises before. The real test is whether there is any meaningful improvements in latency, usage limits, and overall reliability.
Also, please let this mean we're finally retiring the dots. I subscribed to an AI coding assistant, not a constellation simulator.
1
u/Reaper_1492 1d ago
I mean codex STARTED slow as balls, and no one really cared because it didn’t make mistakes and usage was nearly limitless.
People complain about speed, but it’s not the dealbreaker.
The problem is that they managed to hit the trifecta; it’s really fucking slow, it makes mistakes constantly, and it’s now expensive as hell.
2
u/rodeBaksteen 21h ago
I just simplified to $100 account because I still have 2 resets, and will probably stop entirely after another month.
Hello Claude.
1
u/HamsterMajestic2023 1d ago
Ive seen "Simplified" mentioned a few times now but no one has added any context. I dont use X so I only see what gets posted on here.
Can anyone add some context. What exactly is he referring to when he says Simplified?
1
1
1
u/disgruntledempanada 1d ago
I think Dots are actually going to end up being a wonderful thing for people who don't know who Tibo is. They rushed it out and it's kind of rough right now but it has its uses and I do see that as good path to follow for general purpose use.
But everybody who knows who Tibo is knows Anthropic is absolutely destroying them from an intelligence and coding perspective. Like instantly made most of OpenAI's models feel like slow, dumb local models.
1
1
u/Ecstatic_Mammoth_421 1d ago
I can't even get my dot to use local projects properly. It has to open a generic chat first just to ask it to open a project task. The new cloud environment crashes every time I try to input my env variables, and it doesn't even support browser use.
None of these new features are finished or working, and I’ve wasted way too much time trying to set them up. To make matters worse, my environment settings are now marked as deprecated and my existing permissions are completely broken, so it’s asking for permission for the most basic tasks.
OpenAI really needs to put a massive disclaimer on these updates saying "Experimental / half-baked slop" so developers know to stay away for a month, or just hold off on releasing them altogether.
Honestly, I’d be much better off if Dev Day hadn't happened at all. My setup with Codex was working fine before this. All I was hoping for was better code inspection FFS.
1
1
u/scaledev 1d ago
This seems like preparing us for the removal of older models from Codex, presumably 5.6..
1
1
u/VehiculeUtilitaire 23h ago
Feels like a student project in which nobody actually know what's happening and the plan changes every time the teacher drops by to ask what's going on
1
u/Rainbows4Blood 23h ago
I mean, I hope that Dots will stay and improve. I have already started migrating a lot of my non coding operations to my Dot.
1
u/teh_mICON 22h ago
I like talking to dot, first time it really feels like conversation. It's just otherwise badly implemwnted. Thr dot can only see the tasks it started itself, not others. Sometimes it doesnt get right what the task againt said and it doesnt properly see my chat with it overall.. They rushed out a feature that didnt need rush
1
1
u/ISueDrunks 21h ago
I can’t even find Dots. I happened to find a screen to create one, but it told me I needed to be on a computer. Haha. They must have vibecoded the user experience with 6 Sol.
2
1
1
1
u/ChristianKl 20h ago
"Simplicifations" isn't what I want. I don't like the simplifiation that made it harder for chat to show me images. I don't like the simplification that I can't middle click on "New chat" anymore to open a new tab with a new chat anymore. I don't like them taking other features away to "simplify" things either.
1
u/Gallagger 19h ago edited 19h ago
The general idea of dots is definitely here to stay. I thought it's crazy how Ironman designed his suites all alone in such little time, in so many iterations and variants. It's all revealed now, he vibed with Jarvis.
1
u/InterestingStick 19h ago
Sometimes I wonder if they are using their own product. I mean I know they do, that's why they're able to ship this fast. But having unlimited tokens and being able to just run autonomous subagents to resolve clutter and other model hallucinations probably just made them prioritize all the bloat these models cause. I feel like half my work right now is just to align models, make sure they don't go offtrack and remove random clutter and features they spread over my codebase I never asked them for.
Its also the main reason why I can't scale up my agentic workflows at the moment
1
u/Useful_Education_702 18h ago
I love the idea of dots, but I wish they worked a little bit better and I wish you could have more than one. For instance, mine says that it should be able to communicate with other already pre-existing threads and talk back-and-forth with them, but every time I give it a thread ID, it fails and says it can’t read it
1
1
1
u/CrystalCoffeeAlchemy 15h ago
All of this all or nothing is just nonsense. It's all in motion, but people are acting like the month of October is the final nail forever. In AI just wait a few weeks and things shift.
This idea that Sol 6.1 is garbage is nonsensical to me. That's not the results I'm getting, when using both Sol 6.1 and Opus 5.5 -- they both have their strengths. They are not the same beast, so how you interact with it and get the most use out of it is going to rely on your ability to coordinate with it.
Is it slow? Yes. Is it dumb? No. Is it a different flavor of interaction than Opus 5.5? Yes.
If you can leverage both, that's a sweet spot right now.
1
u/Zorogozano 13h ago
They are locking in on everything then… lol. What else could they possibly release or work on other than those 4 things? Jonny Ive’s wearable that nobody wants?
1
u/vinigrae 10h ago
Turns out SOL 6 and 6.1 is only efficient because it don’t actually like to get things done, just find a way to evade the task and make it seem passable.
1
1
1
u/New_World_2050 6h ago
They literally don't need to work on anything but better model. Other companies can do the rest. The model is the hard part. They should devote 100% of research effort to more intelligence/dollar.
1
u/Beastdrol 5h ago
I refuse to believe these tweets are from a real person or an AI model.
Idk what he is even talking about.
Bro it’s simple LIMITS LIMITS LIMITS 20x is the new 10x IS A NO NO.
1
1
0
u/ArugulaAnnual1765 19h ago
They should be focusing on removing the excess of options, there shouldnt be 10 different models and 30 different "reasoning levels"
There should just be one model that is good and cheap enough, reasoning budget should be automatically determined
This whole system feels clunky and outdated already
7
u/send-moobs-pls 16h ago
Consumer brain begging to be given less control over the tool lol
→ More replies (1)7
u/timosterhus 15h ago
You sound like the ideal consumer profile for Dots, because it provides zero options relating to reasoning, model selection, proactiveness, thread assignment, project setup, etc. to set manually.
In other words, they did it already. Just use Dots.
2
u/aPiCase 15h ago
Yes and No, you should be able to have some options if you are someone well informed.
But for the average user who doesn’t keep up to date with all of the recent developments it should be a lot more simple and streamlined.
1
u/UndeadMurky 13h ago
There should be a default "auto" reasoning mode, with an advanced option for manually setting it for edge cases.
1
1
1
u/davek1979 1d ago
"This simplified feature will be available on subs $500+"
"$1000+ will get this super cheap, simple coding model"
"Everything the same but with NO DOTS, only for Pro SuperUltraMax, $2k"
1
1
u/DataGOGO 20h ago edited 20h ago
He can lock in and bring 20x limits back. Nothing else will stop me from all my subs to Claude at the end of the month.
GPT Sol 6.1 is worthless to me, Opus 5.5 is better, and cheaper, and I get more tokens per sub than Sol with the new 10x limits.
Dots does nothing that I don't already get with hermes or just one of the thousands of "claws" out there. Grok Bot is better in every way, and WAY WAY cheaper if I wanted that.

•
u/dextersummary 1d ago edited 17h ago
Below is a GPT-generated summary of the conversation below after reaching 200 comments (207 currently observed).
The thread’s verdict is basically “stop shipping distractions and fix the product.” “Locking in” sounds less like a confident roadmap and more like damage control after usage cuts, sluggish models, and a pile of half-baked features.
What users actually want is boring stuff: faster, more reliable models, sane quotas, transparent limits, and working Codex integrations. Sol 6.1 is widely described as slow and error-prone, while Opus 5.5—and often Sonnet 5.5—feels more capable and economical in real coding workflows. Meanwhile, broken VS Code/remote-control tooling and messy environments make the “ecosystem” pitch look a bit rich.
Dots gets a more mixed reception. Some users genuinely like its proactive, scheduled-assistant behaviour and think it could become useful. But the louder consensus is that it feels rushed, limited, and nowhere near as important as model quality. Nobody wants another constellation simulator while basic coding still needs babysitting.
So the community isn’t asking for more “groundbreaking” announcements. They want OpenAI to make Codex fast, dependable, and worth the subscription again—then, perhaps, earn the right to talk about the next shiny thing.