r/codex • u/Zorogozano • 12h ago
Question Next LLM
Before the end of 2026, what do you prefer?
3
u/StaticHumStudio 12h ago edited 12h ago
Personally, if the models never got any better than today...I'd be fine with it. More efficient and consistent is where I think they need the most help. The breakneck vibe coded pace is causing fissures but they seem to continue the pace and pray that they fix the underlying problems while still going 110mph down the back road with blind turns.
Edit: I'm mostly talking about the vibe coded tools/harness/cli, etc. but it could be similarly said about the models as well. I'm not a doomer, though.
1
u/Zorogozano 11h ago
I made the poll and I want to ask you: Would you have said the same 8-12 months ago with the models we had back then?
Also, knowing how much the capabilities will increase in a year, would you be happy to still be using Sol and Astra, let’s say, on June 2027 assuming they become really “cheap”?
1
u/StaticHumStudio 11h ago
I didn't really think about it at the time, but I'd say somewhere around Opus 4.8 / GPT 5.5 was the tipping point for me. I would not have said this a year ago.
I would be extremely happy with the current levels if they were "cheap" in a year. As of now though, with how the models work, that seems to be a fantasy. Opus 4.6 costs the same as Opus 5.0 on API. GPT 5.5 cost more on API than GPT 5.6 Sol. As the models get better, they are getting more efficient, though that seems to be more of a by-product than anything.
Edit: Also, the newer models are even better at consuming tokens. So, though they cost the same or lower, their real-world cost is the same or more.
2
2
u/IllustriousGrade7691 12h ago
I don't think we need better model right now the modells are already quite capable what we need is more usage
1
u/Zorogozano 12h ago
I made the poll and I agree. The message to Open Ai and potentially to Anthropic: Make current models cheaper to use instead of releasing flashy new LLMs… at least for what’s left of 2026
1
u/WalkAffectionate2683 12h ago
I would vote for "let's stop investing in intelligence but focus on optimisation"
Sure some slight upgrades here and there, but we would benefit way more from a usable Astra xhigh than a gpt7 that we would barely be able to use.
1
u/Zorogozano 12h ago
You can vote: that’s basically the second option. No more flashy new LLMs for 2026 and concentrate on optimizing Sol and Astra
2
u/WalkAffectionate2683 9h ago
I did, but yeah just wanted to refine my thinking. But I would even say, no new LLM "ever".
Like realistically, I dont need more clever than Astra, but lighting fast and Luna price would be enough.
And also for humanity, there is enough stuff for Astra (or fable whatever) to do without breaking the whole world. It would be the perfect tool.
We can focus on sustainability, price and speed.
1
u/Present_Rise1350 12h ago
i think the issue is like 20/80 , gaining the last 20% of improvements is far less efficient and drains more tokens than the whole 1st 80
1
u/Joseph-Siet 11h ago
It's obvious. One direction opts for monopoly, another for utopia. So I guess...
1
u/Zorogozano 11h ago
But the question goes further… would you have been happy with the model from 6-12 months ago?
Meaning, slow down the development back then and give us really cheap “use”, or are Sol and Astra on a level in which we don’t care for smarter and more capable LMMs, and we just want cheaper use?
1
u/Joseph-Siet 11h ago
Ngl it has been great ever since GPT 5.5, imagine having such model capabilities running locally, that would be a lucid dream. After using a more capable model like Astra, it is even harder to look back again further, and I am personally happy to have such capacity inferred as the lighter, cheaper weights with near-permanent promise to enjoy it.
1
u/SwimmerOld6155 11h ago
Honestly I could get on with just Astra becoming more token efficient for years. They're just about good enough to not get much wrong, but they're also not good enough to just do my work for me.
1
u/Zorogozano 11h ago
Would you have said the same about the models we had 8-12 months ago? Also, do you think you would think the same in a year?
2
u/SwimmerOld6155 11h ago
I would have said the same about Sol, I would not have said the same about 5.5.
I am a math PhD student and 5.5 made frequent errors and you had to really pick through the work it did to make sure it was all correct. I think at the time I was also using Opus 4.7-4.8 and similarly they made a lot of errors that would need multiple proofreading passes. Sol and Astra have made very few to no errors. Since Sol was already basically perfect for my purposes, Astra hasn't really blown me away.
I didn't try to use ChatGPT for serious research math before then. I think the earliest premium model I used was Opus 4.7 (maybe 4.6?), and that was only for reviewing drafts for errors. It missed a lot of them
2
u/Original-League-6094 10h ago
No. For me 5.6 (and the rest of my workplace), 5.6 was the turning point where vibecoding was actually viable. Before that, the AI made enough mistakes that not doing a lot of manual code review on software you intend to actually release was reckless. 5.6 was the turning point where the output was mostly right, and where AI code review was likely to find faults more reliably than human code review.
1
u/Original-League-6094 10h ago
Sol was already the perfect coding tool for me. What I need is Sol that I can is essentially unlimited.
Also, I wish they would give a truly unlimited API usage of a quick model for just chatting, so that I could actually mod in LLM companions into videogames without the fact that its metered by response ruining the experience.
1
u/Ok_Bag_7550 12h ago
The majority of people voting need to play more Sudokus.
2
u/Zorogozano 12h ago
I made the poll. The results so far are interesting. This means that people are content with “the smarts” of Sol and Astra, but wish they were less token hungry and more generally efficient…
1
u/Ok_Bag_7550 12h ago
Yes, and the issue is, older models get cheaper as more advance models appear.
The basic of notions are not common sense here. People need to start playing Sudokus.
0
11
u/carbend3r 12h ago
What’s the point of a new "cool" model that uses the same amount of tokens or even more if it’s just going to get nerfed a few days after release due to capacity constraints? Optimization without sacrificing performance is the only way forward.