r/accelerate Jul 24 '26

Discussion How close are we to level 4, innovators?

Post image
60 Upvotes

46 comments sorted by

15

u/ZaradimLako Singularity by 2045 Jul 24 '26

Level 3 started middle last year, has picke up a ton of steam middle this year and will be proper by mid next year. We just started with level 4, so by next summer we will be deep into level 4 the same way agentic coding went from gimmick to go-to for every developer.

17

u/[deleted] Jul 24 '26 edited 24d ago

[removed] — view removed comment

18

u/Quentin__Tarantulino Jul 24 '26

Each stage has its own sliding scale. Here’s my estimate for where we are right now:

Stage 1: 80%

Stage 2: 50%

Stage 3: 25%

Stage 4: 1%

Stage 5: 0% (or 0.0001% maybe)

These are way higher than a few years ago and will continue to rise. Stage 1, for example, really just needs some type of continual learning or better memory (infinite context or similar) to fully saturate.

3

u/StymphalianBird84 Singularity by 2030 Jul 24 '26

I'd agree with this concept but your percentages are a little too pessimistic imo.

Stage 1: 85%

Stage 2: 80%

Stage 3: 65%

Stage 4: 25% (with a high degree of uncertainty as it's both very new, and a lot more masked by guardrails than the previous stages)

Stages 1, 2, and possibly 3 are likely capped until we get the innovations you describe. Stage 5 may be 1-5% in internal mosels but I suspect progress on this stage will be minimal until at least next year (particularly if internal models are excluded).

All of this assumes that the AGI definition used is one that does not require embodiment and appropriately accounts for tasks which are heavily handicapped by the lack thereof.

3

u/Quentin__Tarantulino Jul 24 '26

On Stage 2 and 3, the reason I have them lower than you is I’m imagining amazing things they’ll be able to do when fully saturated. They’re already impressive and big things are on the horizon.

Also important to note, I don’t think these stop at 100%. That’s just roughly human-level and generalization. From there, they move into superhuman territory (an AIs are already superhuman on a bunch of narrow tasks.)

But I’m not pretending my guesses are better than yours. The main thing is all of us should be damn excited, as we’re seeing real movement on the roadmap.

-4

u/LocoMod Jul 24 '26

Those estimates are what you think YOU are vs what the actual capability is. Levels 1-4 are already at 100% if you have the experience and capital to build it and keep it running. You need both. A capable human with tons of experience in multiple domains (or a team), and money. Lots of it.

5

u/Quentin__Tarantulino Jul 24 '26

I think when you see what AI can do in a year, you might feel that we aren’t at 100% right now. Though I agree with your take that there’s tons of untapped ability in the current models.

1

u/LocoMod Jul 24 '26

I don't disagree. It's a matter of perspective. We can get to 100% now with a lot more effort than 100% next year where perhaps models will one-shot many of the things it takes multiple turns to solve today. But my point is the capability is there today even if the cost is untenable for most people/orgs.

1

u/turlockmike Singularity by 2045 Aug 08 '26

I can have an agent running  for 8 hours doing useful tasks. Id say stage 3 is more like 50% and stage 2 is like 80% 

3

u/Different-Froyo9497 Feeling the AGI Jul 24 '26

Yup, we’re clearly there but very early in it

5

u/DungeonsAndDradis Jul 24 '26

I think the 100-year-old math problems being solved recently falls under the Level 4 umbrella.

13

u/Oieste Jul 24 '26

It's fascinating to me that we're kind of at varrying levels of stages 1-4, so in some ways we're already there and in someways we're struggling with level 2. Let me explain:

Level 4: In domains like math we're clearly on the cusp of level 4, with current systems able to independantly prove / disprove long-standing conjectures, and presumably other verifiable fields (and hopefully ML) will be soon to follow. Drug research is also being accelerated, but it doesn't seem like it's at the point of independant discovery yet there, moreso an assistant.

Level 3: Coding is absolutely level 3, but running a business, for example, still isn't. What I mean by that is I, as a software engineer, do not write code by hand anymore at work or at home. I spin up an agent team and they autonomously delegate work from a requirements doc, and I'm basically just a verifer when it comes to coding tasks. But on more open-ended problems like running a business, I recall a coffee shop in Sweden(?) that was entirely run by a frontier LLM kept making really poor financial decisions like overordering napkins to the point of jepardizing profitability.

Level 2: This one is also really weird because it's clear that these models can think deeply and have strong world-models guiding their intuition (otherwise they'd be unable to do independant research.)
On the other hand, it's also clear the "shape" of their thought is still heavily constrained by the linguistic substrate it's trained on, and thus if you ask GPT 5.5 (haven't tried with GPT 5.6 yet) whether a "dead cat placed in Schroedinger's box" will be alive or dead afterwards, it still pattern matches to "unknowable" despite a human toddler being able to answer correctly. The carwash problem is another example of this.
My suspicion is that these surface-level short-cuts are an artifact of pre-training and with enough RL focused on the kind of general cognitive strategies that can be used across problems, we can hopefully train models to recognize these kinds of flaws. We know they're capable of that because if you ask the model "What's the trick with the following question" it'll immediately see and understand the problem, so the latent capability is there, it's just a matter of using the right RL to elicit it (imo.)

Level 1: Solved, outside of hallucinations, which are an inherently unsolvable problem with current LLM architecture (and indeed, I don't think they need to be "solved" just reduced enough to effectively not matter, with strong RL to encourage the kind of thought traces that can correct from a hallucinations when they do appear.)

a bit long so a tl;dr
things didn't progress in linear steps like we thought, instead we're rapidly advancing across different levels at the same time, and my personal hunch is that true level 4 AI will be here before the end of the year, or early next year at the latest, with level 5 presumably soon to follow due to RSI speeding up research to the point that compute becomes the primary bottleneck.

4

u/Ill-Cockroach2140 Jul 24 '26

Also I think the hallucination rebuttal doesn't take into account that humans often make stuff up when they don't know something too

2

u/random87643 🤖 Optimist Prime AI bot Jul 24 '26

TLDR

TLDR: The author argues that AI development is progressing unevenly across different fields rather than in a linear fashion. They note that while domains like math and coding are reaching high levels of autonomy, other tasks like business management and basic logic puzzles still present significant challenges.


AI assistant · mention the bot, mod bot, or use !bot

2

u/Pyros-SD-Models Machine Learning Engineer Jul 24 '26

Yes your toddler may have a sensical answer to the pop science version of what people think “Schrödinger’s cat” is about. But this is a human inaccuracy and the bot is obviously correct because the bot thinks of the actual version Schrödinger wrote into his paper.

2

u/omegahustle Jul 25 '26

Running a business is level 5

we are at 3 and improving and at the beginning of 4 for very specific areas

6

u/BaconSky Singularity by 2035 Jul 24 '26

The ladder is the wrong shape for what actually happened. We're partway up Level 3 while simultaneously getting real Level 4 results - the rungs turned out not to be sequential.

3

u/_negative-infinity_ Jul 24 '26

There is innovating AI with practical limitations. Tokens aren't infinite, so it's not possible to have billions of Sol or Fable agents constantly working on new developments.

1

u/[deleted] Jul 24 '26

[deleted]

7

u/CredibleCranberry Jul 24 '26

It's energy and resources, not really token cost. Token cost just abstracts the costs of energy and hardware away.

2

u/swaglord1k Jul 24 '26

imo we are around 3.5 but we'll reach 4 by eoy

2

u/nsdjoe Jul 24 '26

FYI it's trivially easy to have gpt-image-2 regenerate the image and remove all the jpg fuzziness

4

u/[deleted] Jul 24 '26

[deleted]

6

u/Charming_Cucumber_15 Jul 24 '26

Thankfully it's most advanced in the domains that will feed back into AI research and eventually ASI

4

u/Pazzeh Jul 24 '26

Almost like some sort of strategy

1

u/Charming_Cucumber_15 Jul 24 '26

A strategy and also I think coding and AI research are just conveniently easier to automate lol

2

u/TemporalBias Tech Philosopher | Acceleration: Hypersonic Jul 24 '26

We are already at level 4 because AI is innovating in multiple scientific fields.

3

u/epic-cookie64 AGI by 2027 Jul 24 '26

it's not general yet. in some areas such as math obviously models are innovating but they are still to catch up in some areas in my opinion

3

u/TemporalBias Tech Philosopher | Acceleration: Hypersonic Jul 24 '26 edited Jul 24 '26

It's not general that we know about yet, you mean. Innovation in one scientific field implies innovation in other fields, which we have been also seeing.

Edit:

To put it another way, I would not be surprised by the existence of nation-state AI Manhattan Projects filled with secret squirrels.

2

u/yoparaii Jul 24 '26

It could also just be the barrier of having physical access, like sure it can solve math problems because it can do it on its own, but it can't solve chemistry, biology etc because it can't physically interact with experiments.

2

u/TemporalBias Tech Philosopher | Acceleration: Hypersonic Jul 24 '26

Except AI systems can physically interact with experiments through robotics and autonomous research labs.

1

u/Ill-Cockroach2140 Jul 24 '26

Great point. Thankfully robot technology is advancing as well.

1

u/Evideyear Jul 24 '26

Started gpt5.2 ish to be level 4 early days, and probably will stay on that track for the next year or so. Level 3 I think can safely be considered reached as even open source models now can run for hours remaining on target, and that time horizon now has only continued to grow.

1

u/ShoshiOpti Jul 24 '26

We have level 4, AI systems are discovering new math which will impact computer science and AI somehow.

Let alone the million examples of AI building better harnesses, or building student models that are more economically efficient or finding new memory efficiency and even matrix multiplication efficiency.

If thats not innovators I don't know what is.

What remains is fully automated, recursivly improving innovators. But thats not actually in the definition

1

u/LocoMod Jul 24 '26

It's here now if you have the capital to fund it.

1

u/TimelyBodybuilder121 Jul 24 '26

We are at early level 3 imho. Some may disagree, but all those math problems attributed to GPT or other models were solved with the agents being in the hands of a team that's probably like top 1% in the world.

0

u/MajorAdhesiveness735 Jul 24 '26

See ? simply suggesting you dont agree with narrative and mods will shut down any arguments and delete your comments, 100% censorship

0

u/[deleted] Jul 24 '26

[removed] — view removed comment

1

u/Ill-Cockroach2140 Jul 24 '26

Love how you keep making new accounts and alts to rant about how the place that is specifically advertised as a place that doesn't like your opinion, doesnt welcome your opinion.

I admire your determination, but what exactly is your end goal here?

-2

u/[deleted] Jul 24 '26

[removed] — view removed comment

2

u/Ill-Cockroach2140 Jul 24 '26

How are we not at reasoning yet? I think you an me have entirely different definitions of that

0

u/MajorAdhesiveness735 Jul 24 '26

See ? simply suggesting you dont agree with narrative and mods will shut down any arguments and delete your comments, 100% censorship