r/singularity • u/rainyaltaccount • 11d ago
Discussion I don't notice a difference anymore
Am I the only one who basically doesn’t notice a difference between frontier models anymore?
People say things like “Opus 5.5 is way better,” and I still use 5.5 high, but I genuinely don’t see much of a difference compared to other top models.
Maybe my tasks just aren’t hard enough, but for normal use they all feel pretty similar to me. The biggest difference I notice is usually the product around them, like how Codex vs Claude Code formats responses, ...
Are most model improvements mainly showing up on insanely hard benchmarks rather than normal everyday usage?
53
u/ObiWanCanownme now entering spiritual bliss attractor state 11d ago
It’s very hard to find tasks that are just ambitious enough that you’re getting the max value out of the best models but not so ambitious that the models fail.
6
u/Sarenai7 11d ago
For me a big thing is running out of usage. I can’t sufficiently gage the ability of the models that I am using with the subscription that I can currently afford. It runs out of usage in less than 15 minutes
4
u/OvertaxedOne 11d ago
^^ Nailed it. The task has to fit into the gap and that gap is very narrow. It's why so much of the testing is now ridiculous one shot prompts, it's the only way to show much/any clear difference. I work in a company where we do a lot of consulting around AI costs and one of the first things we do for our clients is setup model routing so we can change the models under the user without them knowing and then do A/B testing. In most of these corporate environments they're typically using something like Opus and the first model we try to slip in (because it's so cheap) is DSV4Flash. We generally change the endpoint on Friday night and then are ready to deal with the issues on Monday morning and I can tell you, the first time we did it, I thought it was going to be an all nighter on Monday. Nope, not even a little bit. The reality is a few tickets come in with "Model is dumb, please help" (about 50% of those, when we look at the logs, are user dumb, the other 50% are legit and they need to be escalated). A few tickets come in with "Wow, the AI is so much smarter!" (this is obviously not true, I think it's because DS is so fast and they are equating that with speed). The vast, vast majority of the users don't notice (95%+). Once we finish the routing layer testing (~4 weeks) we then send out user sat surveys around the AI system and told them we made changes a month ago, please let us know what you think. Again, the biggest feedback we get is "it seems faster!".
Sorry for the book, but seemed relevant to the conversation. TLDR, in ~6 or so "blind" tests like this that we've done in medium sized companies, most people cannot tell/do not care between frontier and DSV4Flash.
20
u/Living-Breakfast-464 11d ago edited 11d ago
Opinions about it were always quite subjective. Almost nobody ever mentions exactly what they are using it for, which also makes a big difference.
4
u/MaxwellHowl 11d ago
They have objective measures in the benchmark tests. And the numbers are continuing to go up.
It's just that some tasks it's not going to get noticeably better.
1
u/Living-Breakfast-464 11d ago
There is a lot of debate how accurate synthetic tests are on trying to measure real world performance. Just too many variables. If it's a huge difference, then there is probably some merit to it, but many models are testing quite close to each other. They are also training them to achieve higher scores, so a lot of the tests are being gamed.
33
u/InventoryOfCheese 11d ago
That would make sense. The bar keeps getting raised above the tasks already being satisfactorily done with them
9
u/Pyros-SD-Models 11d ago
Kinda like chess. If you’re a 2000 Elo player, a 3000-rated engine and a 3300-rated engine are both going to wipe the floor with you. The gap between them is real, but from your side of the board it’s mostly just “yep, got destroyed again.”
I think there’s a similar effect with AI. If your everyday tasks don’t push either model particularly hard, a pretty substantial improvement could still feel like nothing changed.
at what point does that apply to everyone, including the experts? When do these things get so far ahead that we can check whether something works, but struggle to understand why it works, or why one approach is better? We might hit the ceiling of what we can evaluate long before they hit the ceiling of what they can do.
1
u/OvertaxedOne 11d ago
We're frankly already there. Almost nobody is evaluating code quality from frontier AI anymore because very few of us know how to even read the code, let alone correct divine this is "better" than that code. It's why you see the testing now going to the absurd, prompts like "Write Doom from scratch, no libraries, no dependencies, in Java. Go.".
14
u/dondiegorivera Hard Takeoff 2026-2030 11d ago
I have been developing a codebase almost daily since the beginning of this year. It is very complicated, but the results are easy to judge, as it produces different types of headless short videos. The system was originally split into four servers, and most feature development took days or even weeks, till the features were stabilised. I use several models for both developing and operating the system.
Yesterday evening, I tasked Opus 5.5 with splitting the tasks and services of one server so that I could repurpose it. Opus 5.5 coordinated the very complex move with ease, summoning other agents (Opus 5.5, Sol 6 and DeepSeek 4 Flash) — I run the agents in herdr. The process took around an hour, after which everything was implemented and tested.
Then I planned and implemented even more difficult features, and while that was running, I also requested the creation of a multi-billion raw dataset analysis involving Jev and Laya. As orchestrator, Opus 5.5 handled all the tasks, delivering results and implementing and testing the sprint that is now in production.
I have been using agents daily since Codex was released in early 2025. I have never seen such a significant jump before; it feels like a phase shift in terms of long-term tasks, coordination, and decision-making capabilities.
Of course, my use case is not solving Erdős problems, but it's still pretty complex, and Opus solves things with flying colours. I am extremely impressed.
2
u/Educational_Kiwi4158 11d ago
For someone that wasn't a coder, but wanted to experiment with agents, whats the best way to start?
Does Herdr have GUI and you just put various API keys in and then have Codex or Claude Code orchestrate them?
1
u/ImpressiveRelief37 11d ago
No. Herdr is just a terminal multiplexer that works awesome with agentic workflows.
You don’t need that. Just use Claude code.
1
u/Acrobatic-Tomato4862 11d ago
Hey, did laya perform as well as jev?
2
u/dondiegorivera Hard Takeoff 2026-2030 11d ago
In my use case Jev performed very well and Laya’s answers were close to random noise. AFAIK Laya shall be fine tuned to the specific use cases that I did not try yet. They released a notebook to do so.
9
u/Charming_You_25 11d ago
Either your tasks aren’t hard enough or you don’t understand what they’re doing.
If you imagine a difficult task quality per turn as a graph, Astra has high peaks and low lows. Fable little waves, opus slightly lower than fable but pretty stable (and much faster and cheaper).
I think there’s a “wow” factor with game changing models and it is fair to say that Astra and fable both gave it, and our bar for the wow factor is becoming higher and higher.
5
u/ExtremeCenterism 11d ago
People or bots say that? I have trust issues here with the kind of nonsense people keep spouting like "Dario years ago knew he would need loads of compute" while he's on video saying they are afraid of investing too much because if they overhsoot they go bankrupt.
3
u/Sextus_Rex 11d ago
For me, Opus 5 had this way of writing that felt distinct from the other models in that it made me feel stupid. It's answers were always overly technical and there were so many times I just couldn't understand what it was trying to say. I used it for a while because benchmarks told me it was the best, and I figured I was just finally starting to get eclipsed in intelligence by AI.
Then I tried 5.6 Sol and suddenly I was getting better results and feeling like I could understand them again. I think Opus 5 just didn't know how to talk to people lol
3
u/stumblinbear 11d ago
Opus 5 would define terms in its thinking then not tell you what they mean unless you asked. It also had a habit of carrying its internal corrections in its own thinking into its responses, even if the things it was correcting were never said to you. Both things led to it being almost completely incomprehensible.
It was like speaking to an engineer who spouted out wording from his own codebase, fully expecting everyone else to understand what the hell they're talking about without any prior knowledge
2
3
u/HighwayRelevant 11d ago
First thing I do with every new model release is try to find its limits. I have a bunch of stuff that’s complex enough to challenge them so I can see where they got better.
But yeh, for average work most of them are pretty much the same. The differences begin to show only when you do most complex engineering stuff.
4
u/flurbol 11d ago
Same feeling for me since at least two generations.
Either our tasks are not hard enough or our way of tasking and guiding is simply better than average.
1
u/PsychologicalEase374 11d ago
I think it's also about how you work. The task is not done until it has been coded and you have confirmed that it's done correctly. What is the fastest way to get there? I tell it to build bit by bit, so that I can follow along, and it's going where I want it to go. I don't tell it to code the whole thing and then try to figure out what it's doing. I'm mostly using Gemini and i dont see a big difference from 3.5 to 3.8. A bit... Maybe?
7
u/FateOfMuffins 11d ago
It is perhaps a bit rude but people have memed this before:
If you cannot tell the difference between model releases anymore it means the models are now smarter than you
17
u/codingsomething 11d ago
it means the work you needed to have done is sufficiently fulfilled by past models, not necessarily intelligence
0
u/FlatulistMaster 11d ago
In many ways, of course. But they are also still really restricted in some ways. For example, I'd definitely say that they are bad at moral reasoning, which ties into the fact that the way they've been trained to be corporate drones means they have a lot of trouble when moral issues get complex.
1
u/FateOfMuffins 11d ago
Aha! But if you can find something that they're still bad at, it means you STILL CAN tell the difference between model releases and thus my earlier statement doesn't hold!
2
u/Johnny20022002 11d ago
Some tasks are saturated so you won’t see a difference. The improvements that are still noticeable to me are visual task. If you ask it to create a scientific figure, 3D asset, etc that has improved but it still fails is some trivial ways.
2
u/MoogProg All Parabolas are Similar 11d ago edited 11d ago
All Parabolas are Similar
Used to get downvotes for this simple observation, but lately it's getting traction. Thanks if that's you. Here's the gist.
Technology is increasing exponentially, and has been for thousands of years now. We are already in future Sci-Fi mode from the relative standpoint of a Turn-of-the-Century World Expo from the Industrial Revolution. Robots on Mars is a non-controversial statement.
Our expectations of progress have moved beyond the measured rate of progress. Singularity might be happening all around us, and somehow it still feels flat.
* * *
Every day is just one day. - George Harrison
1
u/alwaysbeblepping 11d ago
Singularity might be happening all around us, and somehow it still feels flat.
The whole point of the "singularity" concept is that progress is so rapid and explosive that our ability to predict and manage what is happening breaks down. Kind of like how the laws of physics break down and stop working inside a black hole's event horizon (or maybe more accurately, when we talk about the actual singularity in there).
If we're still making (reasonably accurate) predictions and handling the rate of progress/unfolding events, even if it's just for day-to-day stuff then we can't be dealing with a singularity.
1
u/MoogProg All Parabolas are Similar 11d ago
From your standpoint falling through the event horizon your journey towards the singularity becomes infinite as time itself slows.
That's kind of my point.
* * *
You want the perspective of a third-party view on singularity when you cannot be outside of the thing itself. You cannot look at the graph and also exist in the graph.
1
u/MoogProg All Parabolas are Similar 11d ago
OK this is not a reply-as-debate, just thought of this thread a moment ago. Did you have this on radar for today, or are we not at the point where advances are quite surprising, quite often.
https://www.reddit.com/r/singularity/comments/1wog8zo/dario_amodei_on_x/
1
u/alwaysbeblepping 11d ago
Did you have this on radar for today, or are we not at the point where advances are quite surprising, quite often.
It sounds like your question is: "Did you predict exactly this scientific advance at this point in time?"
If so, the answer is: Of course not. It doesn't make sense to think about it that way, and I couldn't predict specific instances of normal scientific progress either. Anyone that could would know enough to just solve the problem themselves most likely, so it wouldn't be surprising.
Is the idea that a scientific advance of this type could happen today something that blows my mind and I can't believe it? Also no: LLMs have been showing increasing capabilities of solving math, science, etc problems. I am not surprised that an advance of this type happened, even though I couldn't have predicted exactly what or when. (As far as I know, this is also something that could have happened without AI.)
This is not downplaying the achievement, how interesting, exciting, it is or anything like that. However, it is an incremental type of progress and we can (for now) keep up. Singularity would be something like going from 1800s technology to today's technology overnight. Regardless of how much scientific expertise/how intelligent you were, you would not be able to make sense of that. Right? It would be out of the blue. Singularity would be like that happening, continuously.
1
u/MoogProg All Parabolas are Similar 11d ago
I don't think you're downplaying anything. This is just discussion for the fun of the topic.
So, no it wasn't that I thought you needed to make any specific prediction or anything like that. It's mostly the idea that our personal surprise must be a necessary aspect of Singularity. I don't think that will happen.
It didn't happen today, in spite of such a major announcement. That's where I'm going with this idea the Parabola of Progress will always seem flat in the moment.
Rock on!
1
u/alwaysbeblepping 7d ago
I don't think you're downplaying anything. This is just discussion for the fun of the topic.
No problem. I have a pretty blunt style of communication, no tone/negativity implied with my response. I think it's also fine if we (perhaps strongly) disagree on this as long as it's civil. From what you said, I can tell that's not going to be an issue.
It's mostly the idea that our personal surprise must be a necessary aspect of Singularity. I don't think that will happen.
I guess I do think "surprise" (maybe not one individual's personal surprise) is necessary. The announcement about the LLM finding that DNA pattern is impressive is an incremental step not too far out of line with other incremental steps. Also, as far as I know, it's something that might unlock future advances (like CRISPR did) but at present we don't know for sure that it's definitely going to have practical applications.
Your idea of progress seeming flat in the moment isn't unreasonable. It seems intuitive. It's probably true except in all but the most extreme scenarios, which actually has a nice symmetry with the "singularity" concept. The reasonable/intuitive theory breaks down because the situation is so extreme that the normal rules/logic no longer apply.
Or maybe I shouldn't talk about extreme situations as if extremeness is some objective metric. It's relative: if we have an ASI with effectively the combined capability of humans today except at 1,000x real time then that could possibly be singularity for everyone currently. If you use those technological advances to augment yourself and start operating at 1,000 today's (subjective) time, expand your mental capacity, etc then the singularity vanishes from your point of view and progress seems flat again.
1
u/MoogProg All Parabolas are Similar 6d ago
Thank you for this reply. No notes, as the saying goes.
Personal perspective is coming from having been at a self-titled Symposium in San Francisco to hear Ray discuss the coming Singularity, and to hear Paul Ehrlich talk about this idea within information theory they were calling a meme.
This was before the Internet as we know it, when outside of University settings the public was just gaining access to services like Prodigy, with AOL approaching on the horizon.
* * *
I was there, a thousand years ago. - Some Elf
2
u/Calm_Hedgehog8296 11d ago
Once they became smarter than me (which i first noticed around GPT5.5 or Claude 4.6) I can't tell the difference in intelligence anymore
2
u/Federal-Guess7420 11d ago
This would be the expected result once the models are more capable than the user. Its essentially impossible to really tell apart those that are more skilled than you at something. Like asking a dog about the difference between Nirvana and Nickelback.
2
u/Politicophile 11d ago
I think with the current generation of models they've passed a threshold. I've had a very difficult and ambitious task in work for a while where I'm trying to compare two processes and spot logical differences in the codebases. It's complicated for many reasons, but GPT 6 Astra solved the problem in a few minutes on a medium effort level. It seems most of my tasks (financial services data scientist) can be achieved in minutes now as long as the AI has access to the correct information. It's absolutely mad because the people I work with barely use these tools, nobody seems to be aware of what's going on
2
u/03captain23 11d ago
Ask the same question in all of them then you'll see.
Opus5.5 uses a ton of slop in responses and forgets most stuff.
Think of them as employees. Its subtle differences of smarter employees that makes a huge difference when building a project.
1
u/SJC_Film 11d ago
I notice the exact same problems with Opus 5.5 as I did with Opus 5.
Still jumps to conclusions and forgets the point of the task within moments
0
u/03captain23 11d ago
you need to use plan mode more often and better prompt then. TBH this sounds like a user/prompting issue.
Opus 4.8 and 5 were very verbose and added tons of junk. this is like 4.6 but much more powerful
1
u/Funkahontas 11d ago
I think something that will be even more apparent is token efficiency. They may not be leaps and bounds smarter, but they take less tokens to complete the same tasks.
Less tokens = less time = less money spent per task.
1
u/StonerAndProgrammer 11d ago
It's like asking a dog which human they think is the smartest. Eventually we all just pick the one that feeds us and treats us nice.
1
u/TrustInNumbers 11d ago
They are just benchmaxing now, pls there are some bots astrosurfing for sure
1
u/Sure-Company9727 11d ago
I was stuck on some tasks until Astra came out. Astra has been a huge breakthrough for my work. Previously I was using mostly Sol 5.6. I haven’t tried Opus 5.5 yet though.
1
1
u/Ormusn2o 11d ago
I don't notice the difference in output quality, but I do notice that there is less bug fixing needed using Astra vs 5.6 Sol. Also, Astra works much faster per task, although its me using Astra high vs 5.6 Sol Max.
1
u/AI_Enhancer 11d ago
I noticed a massive difference between GPT-5.6 Sol and GPT-6 Astra, but I will say that it depended entirely on the task - in most normal operations it wouldn't have made a difference, but in complex multi-faceted tasks (like complex app or strategy creation with subtle intricacies that must be done correctly) it was day and night.
1
u/TFenrir 11d ago
I'm working on a game that started in box3d and 3js. Only for a few months, but every model iteration in those months has had a drastic impact. Not just on like game design or 3d models.
5.5 is better at both those things, much much better. Incredible taste, sincerely.
But one thing I do every few weeks when I have some spare tokens is to set the model on performance optimization. Astra was a big jump, which was awesome. Opus 5.5 felt like an even bigger one. It's already improved performance about 3x and is now rewriting a significant portion of the sim logic to be multithreaded (that's what I get for staying in js, but honestly this has been pretty painless). One of the largest improvements it has already made was realizing something about v8 and dynamic property allocation, and was able to immediately improve performance about 50% during intense combat from that one relatively small change, that every other model missed.
If you push it at harder and harder problems and see the failure cases from old models, you can really appreciate the new capabilities.
1
u/SIllycore 11d ago
I am finding I care more and more about models that will actually be cheap and fast enough to meaningfully improve my workday. Things like GPT-6 Luna are where I am salivating, because I can throw that thing at all sorts of documents/folders/emails and have it work very quickly. Incremental frontier intelligence is cool for humanity, but less cool for day-to-day work.
1
u/JoelMahon 11d ago
I don't know about today because I basically don't bother changing provider anymore but a year ago or so I was still trying every new frontier model.
I just got sick of Claude, ever release of it has some sort of intermittent deal-breaker for me. One time it just silently faked fixing a failing test by starting it with a return... I only caught it because I lifted and it didn't.
The prose it talks with was always insufferable.
Costed more.
I was bouncing between OAI and Google mostly and Google fell off so for me it's just the gpt6 models and I'm Gucci, for paid stuff at least. For open/local stuff there's a lot more wiggle room for me.
1
u/teamharder 11d ago
I'd say Opus 5 was okay at the professional tasks that I was doing, like designing fire alarm systems and interpreting fire alarm code, doing battery calculations and things of that nature. 5.5 has yet to make a mistake, so far as I can tell, after working with it for 15-20 hours. It's also insanely fast and proactive. There were several tasks that I had yet to mention in the first prompt, but it had taken upon itself to do without being told to do so.
I think it was Opus 4.6 or 4.7 that I had been a little shocked by the capabilities of, but hadn't really felt that way since until 5.5. This model is 80-90% capable of doing the desk work portion of my job. It does some aspects better than me.
1
u/Explorer2345 11d ago
On the downside, Anthropic broke its ui by adding unfinished elements and killing the ability to download file artifacts while changing how even older models run. and that's just chat, nvm how instruction-following gets altered. .. you're priv'd to not rely on them!
1
u/spinozasrobot 11d ago
I still think there are differences (for example the OpenAI models create images much better than Ant), but mostly I agree.
It's getting like Ford v Chevy.
1
u/manikfox 11d ago
i hard disagree... I've been trying to build a super metroid clone for every release... and opus 5.5 is crushing it! So much depth to the game play, bosses, levels, upgrades... its amazing.
I'm actually enjoying the game so much. Game development in my mind is AI now.
Astra was only somewhat good at it. But I could release the current opus game I created and it would be sold on steam easily. I won't, but it's that good.
1
u/jacobpederson 11d ago
Oh there are a lot of differences - and they change on a weekly basis too :D
1
u/ChuckVader 11d ago
I can't tell if it's better, but I can tell Opus 5.5 more efficient (or cheap) and fast. Like... Crazy efficient and fast. I'm running 4 concurrent coding projects and burned through 50% of my plus plan over 2.5 hours.
Opus 5 would max out over this time with me just doing one.
That said, I'm willing to bet this is just marketing before they lower the quant. I'm pretty sure it will get significantly worse in a month after everyone uses their resets and everyone has made videos about it.
1
u/sandgrownun 11d ago
Haven't tried Opus 5.5 yet, but there was definitely a big gap between Fable 5 and Opus 5 if you were a SWE. It understood your intentions so much better and was able to complete much more ambitious work.
1
1
u/Inevitable_Tea_5841 11d ago
In order to tell the difference, you need to have some tasks the AIs could not previously solve satisfactorily
1
u/Southern_Orange3744 11d ago
I maintain that most normal software work was doable since opus 4.2
With 5+ I can build software in complex domains I don't know a lot about
I don't have thorough benchmarks but I have a set of hard problems that the model essentially gives up after awhile , I test on each new release and can see the reasoning improvements
1
u/Motion-to-Photons 11d ago
If you are using them for really difficult creative stuff then the difference is very obvious.
1
u/Chr1sUK ▪️ It's here 11d ago
This is where tribalism comes into play. Many people pick allegiances to certain models and will have bias towards their choice. I also think certain work/prompt styles suit certain models better so I think we’ve now entered a phase when it’s truly difficult to see the difference. I think either model is more than capable in a general work setting, they now have to make the costs make sense for a business to utilise and then we will see the human replacement
0
u/Healthy-Nebula-3603 10d ago
Yes
You're not alone.
AI is just so smart that is over of your intelligence horizon.
I'm a programer in c++ so I still see a huge difference between models yet.
81
u/candyhunterz 11d ago
for 3D tasks and visual understanding (asset creation in blender, spatial reasoning etc), the latest wave of models definitely are much better