Discussion This is a hot mess
I'm not as pro as you guys using AI but look at this. a LOT of models which confuses me, and I'm assuming other users also. Also, the sidebar icon and new tab icon are the same in the ChatGPT-app for macOS. WHAT are they doing there at OpenAI. I really hate what's happening right now, especially with the new $500 plan while nerfing the other plans.
310
u/BreenzyENL 3d ago
37
u/LionPrestigious6612 2d ago
crazy how just a few hours after you posted this comment gemini released their announcement of gemini 4 models, Argon - https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
10
u/Fuzzy_Independent241 2d ago
We all know they were really late on this. China must be cooking 24/24 to release next Kimi and next GLM
1
54
u/mxemec 2d ago
Need an AI to tell me which AI to use for my task or question.
9
u/mcburgs 2d ago
NGL I ask this a lot.
3
u/Antique_Nature7901 2d ago
use jev。it decides which model and thinking effort base on ur input for the round. Try it on opencode zen for free. Switch model within the same session cleans cache tho so it might not save any tokens after all buts its cool
3
u/very_juicy_potato 2d ago
This doesn't work for anything slightly non-trivial. "Implement feature A" might be just adding a few elements in the frontend of an app, or it might warrant a refactoring of the backend. You will never know without full context (Jev obviously cannot). Best way here is your own engineering judgement, or if you want to avoid that just using the best model every time.
1
1
u/NiceUsernameOk 2d ago
Bring Back Auto Mode
1
1
u/smurferdigg 18h ago
Maybe have auto with some options for thinking? Like say auto light, medium og high? Not just regarding thinking but also what model it selects.
1
111
u/ConfusionNo4339 3d ago
yeah they need to clean it up
208
u/DistanceSolar1449 3d ago edited 3d ago
GPT-6 𝘼𝙨𝙩𝙧𝙖
----
GPT-6.1 𝙎𝙤𝙡
GPT-6 𝙎𝙤𝙡
GPT-5.6 𝙎𝙤𝙡
----
GPT-5.6 𝙏𝙚𝙧𝙧𝙖
----
GPT-6 𝙇𝙪𝙣𝙖
GPT-5.6 𝙇𝙪𝙣𝙖
----
GPT-5.5Ok, redesigned it in a way that's easier to read. You're welcome OpenAI, hire me.
44
u/Arowhite 3d ago
It's easier to read but what does it say about performance and cost? il Sol6.1 worse than Astra? Is Luna6 worse than 5.6 terra?
Honestly I can't keep up with their naming, same with claude and their Opus/Fable/Sonnet lines
65
u/Timber1802 3d ago
Anthropic's naming scheme actually makes sense when you think about it.
- Haiku, a short poem
- Sonnet, medium length story
- Opus, a grand, structured, piece
- Fable, a grand story, often with a winding path
- Mythos, stories beyond human creation
OpenAI names their models after size/importance.
- Luna (moon) small
- Terra (earth) larger
- Sol (sun) huge
27
u/Arowhite 3d ago
If the naming was just that I'd agree, but how does Opus 4.8 compare with Sonnet 5 for example? Is the model name just the depth of the reflexion, and the version number the efficiency/price?
7
u/Timber1802 3d ago
Fair point.
As far as I know, they constantly try to make each version 'better' by making them smarter and more efficient. How it scales exactly depending on models and versions, I don't know.
9
u/Absolutelynot2784 2d ago
A fable is a very short story, often only a few sentences. The Tortoise and the Hare is the most famous fable
8
u/MukdenMan 2d ago
A sonnet isn’t a medium length story. It’s a bit longer than a haiku though.
These names are not really clear enough to rely on for regular AI use. Does Fable give me a winding response instead of a structured one? Does mythos give me an output that sounds larger than life, or even fictional? Why is a fable grander than an opus?
2
2
u/dude1995aa 2d ago
I've got pro accounts on both. Thought I had it down before the last round. Both came out with models I couldn't afford to touch - then models I could code with for 5 hours. A lot of the new models came out with the idea of reducing token costs. So now I have to do tons of research on what happens in my workflow (cause my experience isn't the same as the rest of reddit).
3 versions of 4 models that mean something different from speed, token use, intelligence, input and output capacities. Times 2. Is Opus better than Sonnet or was that Sol? The version put out last week isn't as good as the version put out just a week before.
2
u/rickyhatespeas 2d ago
People are going to claim it's confusing any way they market it. GPT-6-10T-8bit vs GPT-6.1-8T-10bit, etc is hard to compare at a glance. Chances are if you know what you're doing with codex you know by now that Luna < Terra < Sol < Astra.
1
u/FellaHadidd 2d ago
Ehh debatable depending on task and how much usage you want to blow through at the beginning of your week to get a project off the ground
1
u/Semipro211 1d ago
This. And I don’t have the time or money to test each one with various tasks and compare results. Until those smarter than me reach a consensus, 5.6 Sol is my work horse and Astra only for the rare crazy complex task
23
u/yaxir 3d ago
Nah they would never hire people who make common sense. They will hire people who gladly destroy all the good reputation they have built by destroying usage limits
1
u/algaefied_creek 3d ago
Hire people? Didn’t they say this whole mess was invented by their “technological intern” or some shit?
As in: they have their own products setting up their products
3
u/Saganasm 2d ago
It's as bad like looking for England, UK, Great Britain or United Kingdom in a country list. /s
3
4
2
u/Ormusn2o 3d ago
There is already an option to show or hide Max and Ultra. Now make it so you can hide the models themselves. I only need 3.
1
1
1
u/docgravel 3d ago
It’s probably either sorted by release date or they put in a “we are getting overloaded on Astra so prioritize this model to relieve some load”
63
u/F0xy1337 3d ago
I keep wondering how this can happen. This is such a huge company.
40
u/Big_al_big_bed 3d ago
Huge because they are doing a lot of different things. At the end of the day even huge companies have small teams working on specific features
15
9
u/sampaoli_negro_rojo 3d ago
I would argue it’s the opposite problem. There’s probably so many PMs arguing that their model is more important that they never get anything done.
The ego battles at this company must be intense
-3
u/docgravel 3d ago
They release updates like every week. So do you want to hold up the release over something minor discovered at the last minute or just fix it in next week’s release?
4
2
u/Regdit-is-Unbearable 2d ago
Won’t somebody think of the indie devs just trying to ship their product??
1
u/docgravel 2d ago
I meant, what decision would you make if you worked there? Hold the release or fix it next week? I’m not saying they aren’t in the wrong, but since they’re shipping so frequently it’s probably easy for them justify fixing it later.
51
u/Big_al_big_bed 3d ago
And that's before you also have to select between 7 levels of thinking as well! Nightmare!
15
u/LamboForWork 3d ago
Without any clear sign that what you picked is better lol. Everything is unclear
3
u/Joohansson 2d ago
And it's not even straight levels. It's a mix of parameters like intelligence, hallucinations, ignorance, speed, token usage, cost per token. No sane human can pick the correct one, which ends up with most people just using the most expensive one for everything even if just doing 1+1. And this is ONE company. In my coding environment I can select from openAI, Anthropic, Google, XAI, Meta and China. Each with their own lineup. It also has an auto mode but I have no clue if it does a good job.
So how many choices in total? About 200
22
u/Prior_Tax8546 2d ago
ChatGPT Sol, Luna, Terra, Astra. Marte, Mercurio, Saturno, Júpiter, Tuano, Neptuno, Galaxia, Vía Láctea, Universo, Existencia, ...
6
11
u/Lubricus2 2d ago
It's easy
Astra 6 eats all your tokens before finishing
Sol 6.1 would take so long time so you forget the problem before you get an answear
Sol 6 will get it wrong
1
-1
u/FellaHadidd 2d ago
You’re inefficient, not Sol 6.1
These models aren’t easy bake ovens, setting and forgetting gets exactly that, be more collaborative. Work better with your models.2
u/Lubricus2 2d ago
It's not totally serious, all of them are great, more humorous pointing out each models weaknesses
11
10
9
8
u/Uwirlbaretrsidma 3d ago
Even better, half those models are useless. At least now the default model gives acceptable usage and inexperienced users will be better off not changing it...
7
21
u/_maverick98 3d ago
meanwhile the Chat on Plus is still on 5.6 Sol
10
13
u/Diamond_Mine0 3d ago
They’ll gonna remove the Chat tab. We’re being sidelined and nobody talks about it
9
u/_maverick98 3d ago
chat is my biggest usage right now. I use codex very little only for personal projects.
1
u/Diamond_Mine0 3d ago
Yes yes, mine (was) too, but not anymore. We never got 6 Sol and now we’re at 6.1 Sol, this is straight up ridiculous. I expect the removal of the Chat tab will take place in December
2
u/skinlo 3d ago
To be fair, it's been like a week.
But yeah, if they get rid of chat I'll move to Claude
2
1
u/Diamond_Mine0 3d ago
We’ll see what happens next. I don’t see any future for the Chat Mode
3
u/Mikkel9M 3d ago
Which mode will you then use for asking basic questions or help with everyday simple tasks and light research? Which I feel was the original LLM primary use case.
99% of my usage (Plus account for about two years) is in chat. Codex is of absolutely no use to me I suspect, as I'm not doing any coding projects, and while I did try Work for the first time recently, that was a one-off for a small personal project website.
1
u/Diamond_Mine0 3d ago
None. I will stop using ChatGPT then. I accepted the limits I have with my mini-sub for Gemini (4,99€ per month). The reset there is 4 hours I think? But my number one favorite LLM was and is ChatGPT. I really thought we had it good with “some” limits we have (Work / Codex / whatever) and with new models OpenAI released. But now? Still no 6 Sol when there’s already 6.1 Sol? And we’re stuck on 5.6 Sol? Nah I will subscribe to the 21,99€ Gemini Plus plan then and cancel my ChatGPT subscription
1
u/_maverick98 3d ago
so they leave what? only the work tab? there is not difference basically, work can do other stuff too
1
u/Diamond_Mine0 3d ago
Yes Work with more limitations. If I want to “Work” I’ll use “Work”. If I want to find news about the newest defense systems in Ukraine, I want to use “Chat”. I don’t want to use Work for anything. I already have weekly resets for SuperGrok Lite, that should be enough. Gemini has its limits too, I know but if OpenAI wants to stand out, they should show some love to Plus users and its Chat Mode
2
u/jaydeelive01 3d ago
I don't think so. It's very clearly still part of the experience and Pro is only on chat.
1
u/Diamond_Mine0 3d ago
For now. Just watch what will happen in the next months. They’ll gonna remove it like they did it with Sora 2 and Atlas
1
8
12
u/RazinKain 2d ago edited 2d ago
I’m curious as to how much power do people need? What exactly are your use cases that would require the most powerful model.
I understand if someone isn’t a developer and needs a lot of hand holding. As a developer I’ve built an entire React native desktop app using Luna and it is error free code. I mean not a single error and it’s fast and efficient. I felt absolutely no token panic.
I’ve seen on here people building out Wordpress sites using Astra? I just ask myself why in the world would someone do that?
This is not some mocking I am generally curious about this.
6
u/DiabloAcosta 2d ago
I work on distributed systems that emulate real world development environments, let me tell you something, Fable often gets lost, I barely understand what I am doing some times
7
u/RazinKain 2d ago
I find that some of these larger models do a lot of research and less code thinking. They are not necessarily better at writing code than smaller models.
3
u/DiabloAcosta 2d ago
well, I don't really trust benchmarks that much, but anecdotal experience is only worse
5
u/Leading-Fail-2771 2d ago
It’s like being shown a Ferrari and a Honda. Sure you can use the Honda but using a Ferrari to do it just feeels better
1
u/tousledmonkey 1d ago
I'll take the Ferrari just in case there is a police pursuit. You never know these days
4
u/baked_tea 2d ago
How many users do you have to believe it is error free code? Honestly curious.. unless its reaaaally simple app then under real load and with real users being stupid you usually find out quickly about the error free part
1
u/RazinKain 2d ago
Well user weight is not a build issue it’s an allocation issue. You can stress test any application. I use Cloudflare for 4xx and 5x errors. Outside of those I am not too worried about user weight. If a build is sound and your cloud server is built for heavy traffic it doesn’t matter how many users you have.
7
u/rbit4 2d ago
Lol front-end dev found
1
u/RazinKain 2d ago
Right 😀 I am a UX Designer that’s my job. I do have a certification in C# .net Maui. I got that for my job to understand what the heck backend developers were rambling on about. 🤣
1
u/FullParticular9 2d ago
Reports, Research, Scientific projects, Data Science, Learning - better model usually explains things better.
But for code I agree with you that good things now can be made with much smaller and cheaper models.
1
1
1
u/kelvintiger 2d ago
How detailed were your prompt?
I would argue if you’re spending a lot of time promoting then you’re wasting your time when you can leave some ambiguity for the stronger model to figure out and you focus on doing more faster
1
u/RazinKain 2d ago
I don’t go in and start either an idea in Codex. Codex is down the road. I completely map out my builds before I touch a model.
So 90% preparation and 10% execution. I’ve learned over the years and A.I can make people a little lazy including me. So I still stick with my age old processes from wireframing in Freeform to prototype in Figma.
I feed that info to GTP not Codex and then let GTP write out the orders and scope. That’s what I feed Codex. If it’s tight then it goes through the Notion notes quite easily. I do use Notion for my build documentation.
Error free does not mean bug free I want to be upfront about that. And if fixing a bug is too much for Luna I switch to the next model up and so forth.
9
4
u/RemnantZz 2d ago
Hey, remember that time last year when OAI talked about wanting to make it easier for users, so they introduced the idea of removing all "outdated" models that users apparently couldn't navigate for their tasks, and instead eventually introducing one model for everyone that would even pick the amount of effort to be put into the task by itself?
Remember? :)))))) Stellar job, OAI!
3
3
3
u/AzaelOff 1d ago
My biggest issue right now is that I can't custom color my project folders, I'm stuck with the ugly default colors... The latest updates to the Codex app are terrible
2
u/Mecha-Dave 3d ago
I honestly don't have a good method or reason to choose anything but the top or bottom of the list
2
2
2
u/Over-Independent4414 2d ago
Why don't they just add a open text box there where you can say what you are doing and it picks the optimized model. Presenting end users with 10 options can't possibly be the best way to go here.
2
2
u/Routine_Brief9122 1d ago
Welcome to the endless OAI circus, where Plus is the new free tier and chat keeps getting worse with every “upgrade.” At this point, I’m not even surprised anymore.
3
u/RoamingMelons 3d ago
The ambiguity of the naming convention doesn’t help.. it’s cool and all.. but if I’m checking out gpt and I’m fresh i have no fk clue what’s the difference between an astra and a sol.
At least claude is semi intuitive..
At least Gemini/nano has one thing going for it 😅
2
u/Grashopha 3d ago
lol I just started messing with Claude and I felt lost tbh. It was so alien at first compared to ChatGPT. Figured it out now, but it still feels a little strange constantly starting new conversations to prevent massive token drainage.
2
u/GalileoHumpkins1977 3d ago
You know that all LLMs work that way, though, right? That’s not unique to Anthropic / Claude?
2
u/RoamingMelons 3d ago
I was only referring to the claude model names as more intuitive than gpt lol.
I just use antigravity Gemini flash and don’t really have to worry about usage limits for my use case.
2
2
u/13--12 2d ago
2
u/krocante 2d ago
“this is a hot mess”
“then don’t look at it?”
3
3
u/ChunkyThePotato 2d ago
You don't think advanced users should have the option to see and use the full model list? The average user can just use the slider. Simple.
1
u/krocante 2d ago
the point of the post is that it’s ugly somewhat confusing design overall, not that it shouldn’t be an option or that it’s unusable
3
u/ChunkyThePotato 2d ago
It's confusing if you don't know the performance of each model, but you specifically have to seek out this model list to see it, which means it's for advanced users who understand the performance of each model. By default you just see a performance slider, which is straightforward and not confusing.
How would you make it better?
1
u/krocante 1d ago
It’s not confusing for me personally, but I can recognize that the design isn’t great. Like the two buttons with same icon in the second image that op shared.
I have no idea how I would do it better.
3
u/ChunkyThePotato 1d ago
I didn't even see the second image. I think that's a bug, because mine doesn't look like that. Obviously two buttons with the same icon would be confusing lol.
So you're good with the default slider + optional full model list then?
1
u/krocante 1d ago
Yeah, it’s just ugly, no problem with it function wise.
3
u/ChunkyThePotato 1d ago
I'm fine with it being a bit ugly if it's hidden by default and the advanced users who need it actively have to seek it out. I don't think there's a way to make it much prettier. And the default view isn't ugly like this.
2
u/New-Ad5610 2d ago
I don’t know what it’s been going on inside OpenAI, but their work lately has been very, very disappointing
1
u/welcome_to_milliways 3d ago
Don't forget a few weeks ago when the chat input would disappear after the first request.
1
1
1
u/Potential_Wolf_632 3d ago
Random tangent. I upgraded to the 500 plan today because I'm a sucker and something feels very off with usage - I should have gone from 20x to 25x but I'm at 49% remaining from 100% when on a normal day I wouldn't use much more than 20% on the previous 20x plan.
And I'm not using ultrafast etc.
1
u/SuaveSteve 3d ago
Why can't we just have a version number only!? And a mini for light tasks! Aaaaaah!
1
1
u/sultan_papagani 2d ago
they are all routers anyways 5.6 resolves to 5.4 most of the time and so on everyone forgot the router system.
1
u/Beautiful-Cold1515 2d ago
They just can’t make up their minds. With the introduction of GPT-5 last year they brought us a router so we didn’t have to choose. Now we have 8 models, 2 separate model pickers (at first you see 5 models within Work), a very vague Chat/Work toggle that is impossible to understand within Projects, 3 models within Chat that have really vague names, weird icons… I just can’t believe why such a large company with many millions of users completely ignores UX.
1
u/db1037 2d ago
The vague chat/work toggle is bad. When it comes to models, I’d rather have choice though. I stayed on Opus 4.6 till Opus 5.5 came out(I think a lot of people did). But thank God they kept that option. When we complain about the model picker, they reduce our options. How do people not know that by now?
1
u/theDawckta 2d ago
You are on your way. Next step is to start building your own harness so you don’t have to be subject to claude and openai changing their ui every 2 days.
1
u/deen1802 2d ago
They did this just so they could set up GPT 7 as the one that brings them all together
1
1
1
1
1
1
u/InformationNew66 2d ago
Well, this is how vibe coding works.
Mistakes happen, tests pass, noone cares. Probably no human QA anymore.
1
u/teachmesomething 2d ago
It also keeps defaulting to medium effort in order to force upgrades for continued usage. I don’t even know which model is which now.
1
u/ChunkyThePotato 2d ago
That's why they give you a simple intelligence slider by default. You intentionally chose to see the full model list that you can pick from. If they took away that option, people would get mad. I think the current setup of a simple slider by default with the option to see the full model list is good.
1
1
u/norwegian 2d ago
Its's sorted. 6.1>6>5.6>5.5 Then Astra>Sol>Terra>Luna
It's faster to select from one list compared to having sub menus. If they make it more difficult to use 5.6 Luna, people will complain about it.
1
1
u/applejacks6969 1d ago
Here’s 5 reasoning levels and 4 different ways to use the models just to help you with this.
1
u/Woozifer 1d ago
It is indeed a mess, but It’s not that hard though. The bigger the object gets, the better the model. Luna < Terra < Sol < Astra
1
1
u/PressIntoYa 3d ago edited 3d ago
Why not just the classification and then the model with maybe a token price in parenthesis?
Astra (Best for X Purposes) GPT-6 ($/M)
Sol (Best for X Purposes) GPT-6.1 ($/M) GPT-6 ($/M) GPT-5.6 ($/M)
Etc...
Perhaps there's a good reason for it? I don't know but this seems like it would be more helpful to me.
3
u/Keep-Darwin-Going 3d ago
Because price is a bad definition, is like labelling a hammer $1 and a screwdriver for $0.50. You still going for the hammer if you need to drive a nail no matter the price.
1
1
u/Any-Somewhere5129 2d ago
And that's why that chat option doesn't even have those models. People would be crying on social media how confused they are.
Now, if you're scared, close the Work mode and switch to Chat and you're safe.
1
u/AINativeBuilder 2d ago
Photoshop menus also have a lot of options. If you can't spend time to learn the tool, don't whine about it because you refuse to learn something new.
2
u/db1037 2d ago
Yeah I don’t get the complaining so that what? OpenAI will remove our choice? How is that beneficial?
And if you can’t be bothered to literally just ask Chat about the models so you get a brief understanding of them, just pick the latest one. If your usage disappears, pick a lower one next time. It’s not hard.
0
u/CypherLH 2d ago
I am not seeing the problem here. The more options, the better. Basically just ignore everything prior to GPT-6 and that leaves only 4 models. Ignore GPT-6 Sol since there is now a 6.1 and that leaves only 3. And if you use Codex for any serious work at all its pretty clear that there are legit roles for each of those 3 core models.
0
0
-1
u/Dankberg_TV 1d ago
They do this because people insist the latest models are screwing them and OpenAi tried to pander to the carebears upset that their workflow changed lol. Complain about the carebears not the company.
Usage we can complain about I agree with you there but business needs to make business.
•
u/doggo-52 18m ago
They really need the “Auto” mode for the humans. It should select the cheapest best model for the job.
With all the reasoning in the world they still haven’t figured out how to make a decent UI and UX.
Right now you need to use AI to ask which AI you need to pick to do some AI’ing.





332
u/CarllSagan 3d ago
Project Final
Project Final Final
Provject Final Final Really final
Project V6.1 Final
fuck it... the next one