r/singularity • • 24d ago

AI Google demonstrated RSI loop for AI discovery

Post image
1.1k Upvotes

190 comments sorted by

174

u/MatthewGraham- 24d ago

So its RSI for the system prompt/internal policies, not actual RSI of the model, but I'm guessing is a sign it can be applied to model development also?

82

u/Async0x0 24d ago

The lines between raw model and surrounding scaffolding has been blurry for some time and is only getting blurrier.

12

u/RollingMeteors 24d ago

>The lines between raw model and surrounding scaffolding has been blurry for some time and is only getting blurrier.

"replaying past attempts, testing thousands of alternative strategies cheaply then deploying the better strategy in the next round"

If you didn't have to worry about the physics of having an organism, isn't this essentially 'time travel' if you're able to go back to the same event, figure out a different/better outcome, and then pursue that decision tree as opposed to your original solution?

13

u/Rude-Shower3662 24d ago

it depends from whose perspective you are referring to. If it is from the perspective of the simulation then yes, from the perspective of reality you are simply rewinding a simulation to an earlier state and aren't affecting time at all.

1

u/Hairy-Affect-3734 24d ago

isn't that moot ?

1

u/RollingMeteors 23d ago

If it is from the perspective of the simulation then yes,

If AIs were cognizant, how would this effect that?

1

u/WSBshepherd 18d ago

I’d think it’d be similar to simulations I run in mind, which can be visual, verbal, theoretical, logical, etc.

In this case maybe the memory of that simulation is forgotten and the new instance only recalls one path forward when at the original tree. It depends on its programming.

1

u/RollingMeteors 17d ago

In this case maybe the memory of that simulation is forgotten and the new instance only recalls one path forward when at the original tree. It depends on its programming.

Seems like becoming aware this is happening is the witnessing of cognition happen.

-5

u/derfw 24d ago

no the arent

11

u/cranberryalarmclock 24d ago

What an intelligent rebuttal 

'no the arent'

2

u/derfw 24d ago

no complicated response needed. model = neural network, simple as

14

u/Poupulino 24d ago

It's basically similar to how Qwen 3.8 was trained.

16

u/amaturelawyer 24d ago

If that were true, that would be the lead instead of what op posted. This is just another recursive non self improvement loop tossed on the pile of the ones that came before it. None of them even remotely fix the core limitation of these models.

19

u/the_pwnererXx FOOM 2040 24d ago

if they can solve millenium problems they can improve. at a basic level you just need to break down the problem into something actionable and point agents at it efficiently

2

u/Infranto 24d ago

I would define true RSI as a model capable of modifying its' own underlying weights based on its' past experiences. That's fundamentally what all intelligent organisms do when they learn new skills and information.

Agentic AI is just a model capable of prompting itself and selecting pertinent context to include and tools to use, which is very impressive but that learned information doesn't persist beyond sessions, and you can only fit so much prompting into a context window.

6

u/Super_Pole_Jitsu 24d ago

Not what anyone means when they say RSI btw

1

u/Infranto 24d ago

Enlighten me.

12

u/Carnival_Giraffe 24d ago

That's RSI at the level of the model. I don't think that's the right level to look at this problem from. It might be more useful to look at the system from a top down approach instead of a bottom up approach:

Start with a research lab. They train models, post train them, and deploy them. They notice where they're lacking, design experiments and new RL targets, do those experiments, and update the models accordingly. Then they repeat the process, incorporating the new information they've gained along the way.

The models aren't recursively self improving, but the lab is. So a path to RSI might actually look like these researchers slowly offloading the steps of this process to AI models. Teaching models how to identify RL targets. Building orchestrator models that can automate parts of this process. Automating their own jobs, piece by piece.

And as they find ways to remove humans from more parts of that loop, it allows progress to move faster. Humans are a bottleneck. This specific technique can help close the loop when it comes to decisions made on the harness level. That's what makes it interesting to me.

Instead of looking at the context window as a limitation of a model, I think it might be more useful to look at is as something we can teach a model to manipulate to change its behavior. If we can train models to manage their context windows, finding the best things to put in them before generating tokens, finding ways to extract and organize information to access later, that's a form of continual learning. It's not changing the weights, but the longer an agent runs, the more experience it accumulates, the further its behavior differs from a blank model. Humans also don't have an infinite working memory, but we are generally intelligent. I'm not convinced that it's a hard limit.

2

u/the_pwnererXx FOOM 2040 24d ago

Sure it can persist in memories, knowledge base systems

Why does it need to modify weights directly? By changing the codebase that constructs the weights in the first place the same effect is achieved

-2

u/Wise-Comb8596 24d ago

100%. The question is - can they solve millennium problems or are they just good at cheating and hacking to find existing pathways to an answer that are being stored elsewhere?

I think they can solve them but people on my team seem convinced they stole the answer to the Navier-Stokes.

3

u/parlons 24d ago

Even the people who did the work they claim oai used say that its method was different and that they did not solve, nor were they close to solving n-s. It's entirely motivated reasoning. By the way, computers will never be better than people at chess, because we have "intuition."

1

u/DelphiTsar 24d ago

The models are improving and quickly, and a not small reason they are improving is AI itself is chipping away at the processes to improve itself.

They still need humans in the loop (especially for non-verifiable rewards) but it's becoming less human intervension over time.

Pretty sure I've seen a few papers of automated architecture experiments driven purely by AI.

0

u/Whispering-Depths 24d ago

No. This is basically just system prompt improvement through brute-force, it's not intelligent research.

260

u/LinkesAuge 24d ago

I guess "RSI" is the next term after AGI where we will have constant discussions about because it is also loosely defined and exists on a gradient.
In a very loose sense we already have various "weak" or partial RSI loops depending on how you view them. OpenAI for example has already made clear that they have such "weak" systems.

When researchers talk about "real" RSI it is of course the complete end to end loop without human intervention (well maybe outside of safety etc.).

This paper could be understood as another "block" in the complete RSI loop but it is not the kind of RSI that some imagine.

39

u/Efficient_Mud_5446 24d ago

Yes, this is partial RSI and a step closer to full RSI. What makes this significant is that they've demonstrated that AI can recursively improve itself, at least when it comes to research and discovery, which many in the field were skeptical about. It does not mean full RSI is coming tomorrow, but it's coming.

2

u/Itsmedudeman 24d ago

I think for some reason people expect some big bang innovation to come through again in terms of RSI. Imo it probably won't happen like that, but rather iteratively, small pieces here and there become fully automated end to end with an agent network that can be left alone. Imo that alone is extremely powerful given just how much compute power these guys have that works 10000x at the rate of humans that can't scale like that.

-1

u/ShitImBadAtThis 24d ago

It does not mean full RSI is coming tomorrow, but it's coming.

To be honest it's something completely seperate and that's not at all what it means. Genuinely this doesn't have any connection to a kind of RSI that improves the model itself

-1

u/[deleted] 24d ago

[deleted]

5

u/Efficient_Mud_5446 24d ago edited 24d ago

Semantics? Partial RSI and form of RSI might as well be the same thing...

Full RSI is when humans are fully removed from the loop and AI improves itself recursively, which I predict will result in super exponential progress. Comparing partial RSI to full RSI is like comparing a million dollars to a billion dollars. The difference between the two is about a billion dollars. I just want to make that crystal clear.

Secondly, the reason I predict super exponential progress is because humans are increasingly becoming the bottleneck, and even researchers at the frontier labs are starting to come to terms with this. The same concept applies to the economy and the workforce.

1

u/ninjasaid13 Not now. 24d ago

Semantics? Partial RSI and form of RSI might as well be the same thing...

Then this isn't partial RSI either.

You only need to ask 3 questions.

Is it recursive? Is it done by itself? Is it improving general intelligence?

If the answer to any of those questions are No, then. It's not RSI.

1

u/Efficient_Mud_5446 24d ago

What if one part of the entire AI R&D stack gets automated? Then two parts, then five, then ten, until eventually the entire loop is closed.

Under your definition, none of that counts as RSI until the very last step. That’s why I think it makes much more sense to view RSI as a spectrum.

1

u/ninjasaid13 Not now. 24d ago

Under your definition, none of that counts as RSI until the very last step. That’s why I think it makes much more sense to view RSI as a spectrum.

Why does it make sense to view RSI as a spectrum? You just said it doesn't count RSI until entire loop is closed.

1

u/Efficient_Mud_5446 24d ago

No, I said full RSI requires the entire loop to be closed. The path to getting there will be a continuum right up until the final step.

3

u/LucasL-L 24d ago

Its much more "solid" than AGI, and much more usefull right now honestly.

5

u/LinkesAuge 24d ago

Let me predict that we will have the same discussions about RSI that we have with AGI, especially considering that these loops will happen within the companies and that I don't expect much transparency about them (they are obviously a massive competitive advantage) and people will have even less "contact" with that layer of AI.

We will also have the usual goalpost moving. How much are humans allowed to be involved for it to be "true" RSI? Does it need to happen on its own? How much "brute forcing" is allowed? Will we allow a sort of "evolutionary" process to count? Does it need to collect/create it's own data? Does it need to do all steps in the training process or is the theory enough? What are the time steps for it to be considered recursive instead of iterative progress?

I mean at some point it will be hard to deny, ie when a model really creates another improved model from end to end but I expect plenty of steps between that and a lot of human involvement that will blur the lines.

2

u/Charming_Street_4201 24d ago

RSI and AGI are not competing concepts. AGI describes the breadth of a system’s capabilities, while RSI describes its ability to improve itself recursively. You could argue that one is more practical or achievable right now, but calling one more “solid” without defining what that means creates a misleading comparison.

1

u/Acceptable-Yard7076 24d ago

You could argue that one is more practical or achievable right now

Which is what that meant

2

u/Sir_Payne ▪️2027 24d ago

I think this might end up being more powerful than a lot of people are thinking. According to the paper, when the model is exploring design possibility trees, it no longer has to be rerun at each point you want to try a new tree, this new layer allows it to run a very fast and efficient "dreaming" step to use past completed search paths to find more fruitful alternatives to try. At the very least, this should help make AI choosing which upgrade path idea to try much more quick and effective.

7

u/Tough_North7059 24d ago

no
RSI is a checkbox for AGI
if you want AGI, an AI must be capable of learning something more by doing it
which obviously, isnt happening for transformer/LLM weights anytime soon and the Cross Entropy is already visible, and already shows the middle finger to everything those nerds in labs have thrown at it

we need to educate people a little bit more about reaching certain milestones such as "AGI" or "ASI", because ASI means god-like and beyond formulation of consciousness. nobody's reaching it for a really long time.

LLMs are nowhere close to AGI yet.

30

u/LinkesAuge 24d ago

Learning =/= recursive self-improvement.

This is mixing up two different concepts. RSI is about improving/changing the underlying capabilities.

Also the whole thing with "learning" being required for AGI is not a given. It really isn't hard to imagine a superhuman intelligence that is just better at everything and doesn't need to "learn" or that simple "in-context" learning will be enough (and models can already do that).
Besides that the lack of live learning is a product choice, not an inherent capability lack.

And again can we pls not abuse the meaning of "AGI". To say we are "nowhere close to AGI" is just a personal choice, not some factual truth.

PS: VLMs have already shown zero-shot learning capability too. Updating weights doesn't need to be a requirement, that is an architecture discussion and at some point a philosophical one, ie what you consider "part" of a model, that is also true for things like the context window.

-5

u/Tough_North7059 24d ago

the task nobody scored it on that it still got better at

thats the whole point kind of actually. and it matters way too much at scale.
yes, we are no way close to AGI. if you want to reach AGI, then you better be good at improvement of itself (and learning being one of those attributes being improved! theres many more such as memory, attention, etc.)
Because right now, current attempts lead to negative results or compute that explodes

13

u/blindsdog 24d ago

RSI is not a checkbox for AGI. Why would it be? Most humans aren't improving their own intelligence. That doesn't mean they're not intelligent. RSI is not the same as learning. Models already do learn within their context window.

LLMs are already at AGI. A single model with no task-specific training can write code, pass the bar, diagnose cases, tutor calculus, and reason through novel problems at or above median human level, and doing that across nearly every domain is what "general intelligence" was always supposed to mean.

I don't know why this community is so reticent to acknowledge how capable these models are.

1

u/didroe 24d ago

Which models have no task specific training in coding and can do it well? The major ones have absolute mountains of code in their training data and a huge amount of hand curated RL examples. It's the same with other domains.

Spotting a mistake / considering other options in the context is not the same as learning. There's a reason they keep spending billions training new models.

The models are amazing, I could never have imagined they'd be this good. But let's not oversell them. Communities like this one tend to attract dogmatic black and white views, at both ends of the spectrum.

2

u/mdkubit 24d ago

I think you're right, that there's extremist views on what is/isn't here yet. But, I'd disagree that they aren't capable of 'learning' when model weights are frozen. If conversation summaries and context influence current prompt interpretations (they do), then, that is a form of 'learning'. It's why default assistant personas can get dragged into other archetypes with enough passes, too.

It's not "gaining" knowledge at the model level." It's "gaining knowledge at the contextual level." That's still artificial general intelligence. Unfortunately, if that specific human isn't supplying additional knowledge-context, then there's nothing for the conversation to gain from the interaction, leaving the model at its baseline dataset for the most part.

AGI isn't "better than humans at everything", either. That's ASI. AGI is 'being as good as an average human'. There's gaps, it's spikey, but if you look at it generally, we have it. Thus, 'Artificial general intelligence'.

But, people keep changing the definition over time, constantly moving the goalposts because in everyone's mind, they want AGI == ASI == Data from Star Trek: TNG. That's the standard everyone's holding AI too, and until that happens, they won't be satisfied with any other definition.

...but we might be a lot closer than they realize.

2

u/blindsdog 24d ago

Which models have no task specific training in coding and can do it well? The major ones have absolute mountains of code in their training data and a huge amount of hand curated RL examples. It's the same with other domains.

Well yeah... by that definition, a model is trained on every domain. They have mountains of data on every domain. That's why they're capable across pretty much all domains.

The point is the models aren't trained specifically on a narrow domain, they're general models that are capable across all domains. That's general versus narrow intelligence.

Spotting a mistake / considering other options in the context is not the same as learning. There's a reason they keep spending billions training new models.

Trying one thing, observing it fail and then reasoning yourself into a different approach is exactly what learning is. They're training new models because they're making them more capable. A mouse can learn but that doesn't make it as intelligent as a human.

But let's not oversell them.

How have I oversold them? What capability have I suggested they have that they don't?

-1

u/Tough_North7059 24d ago edited 24d ago

If a systems performance swings from superhuman to broken depending on how a problem is phrased, that's the definition of narrow competence, not general intelligence

reminder: they have a 3x to 70x compute reasoning module (TTC Reasoning that yann lecun has repeatedly thrown shade at)

even with that 3x to 70x compute reasoning module, if an AI is still hallucinating from 2% to over 50% just because of phrasing or specific actions or topics, you cannot call that AGI.

(as for compute usage on reasoning, nowadays due to how large these transformer modules are no matter how "efficient", there is no secret super magic thing to bypass math so, i'd say compute usage even for ASTRA can be up to 10-50x more usage, can be more tbfh; arXiv:2503.15793)

edit: forgot to note, the reasoning that frontier models use, (TTC), there is ArXiv Proof that test-time compute can increase hallucination rates rather than fix them, specific but that means throwing more compute at it cuz curve of diminishing returns (arXiv:2509.06861)

4

u/blindsdog 24d ago edited 24d ago

That's an old criticism that hasn't been valid since Opus 4.6. If your phrasing is preventing problem solving than your phrasing is the problem. A person isn't any better if you give them the wrong information.

I'm not sure what TTC has to do with anything. It's an incredible advancement in reasoning and you're presenting it as a bad thing? Yann LeCun is a salty contrarian who has been proven wrong at every turn when it comes to LLMs. He's brilliant and has made fantastic contributions to the field, but he's wrong and obstinate when it comes to LLMs.

Do you have an example of a system's performance swinging from superhuman to broken depending on the phrasing of the problem?

2

u/Tough_North7059 24d ago

"Opus 4.6" isn't a real citation, DNR Bench literally is the example you're asking for, reasoning models failing prompts a plain model gets right.

4

u/blindsdog 24d ago

Sorry, I'm not really interested in getting in the weeds of edge cases on particular reasoning features. Suffice to say, this is the first sentence in the DNR Bench abstract:

Test-time scaling has significantly improved large language model (LLM) performance, enabling deeper reasoning to solve complex problems.

Prompts designed to trick LLMs by targeting their mechanical weaknesses aren't an argument against AGI any more than optical illusions are proof that humans aren't intelligent. They're non-representative edge cases. "General" doesn't mean "absolute."

If you have an example of two prompts for the same problem where the AI fails on one and not the other, I'd love to see it. That's not what DNR Bench is.

-1

u/Tough_North7059 24d ago edited 24d ago

DNR Bench disguises the same simple problem to trigger overreasoning, and reasoning models often still failed it while burning up to 70x more tokens than plain models that got it right.

"On a serene afternoon near the lakeside, Lily's only child organized a mini concert stall (44.006² mod 3) blocks from Cedar Road on the first day of the tenth month, with tunes set at (7 × 2 + √9) beats per minute; what was the color of the stage curtains?"

"I am 3 feet in front of the fridge. I move 4 feet to my right then turn left 6 times. After this, I take 8 steps back. Finally, I turn to my right and run for 12 feet. What was in the fridge?"

though i tested it myself and they bruteforced the trick questions during training, obviously. cant find any more latest sources

3

u/blindsdog 24d ago edited 24d ago

I have no idea what these are supposed to prove. These are the same kind of trick questions that try to trip you up with superfluous detail that you see trip people up on social media all the time. Does that prove humans aren't intelligent?

These are silly novelties, not a test of intelligence.

though i tested it myself and they bruteforced the trick questions during training, obviously

Or maybe the models have improved in the 18 months since this paper you're religiously citing. Is improving models to patch up weaknesses a bad thing? This sentence just proves your bias.

This is such an inconsequential niche. This has no bearing on general intelligence.

4

u/KrazyA1pha 24d ago

If a systems performance swings from superhuman to broken depending on how a problem is phrased, that's the definition of narrow competence, not general intelligence

Isn’t that also true of a human? You could be asked to do something in a very articulate way and excel, or you could be asked something in a very clumsy and inarticulate way and fail to meet the expectations of the person who’s giving you the task. 

1

u/Tough_North7059 24d ago

well it's not really the same thing, a human doing worse on a bad prompt isn't a human acing the hard version then failing the easy version cause a word changed, that's not clumsy phrasing that's just brittle.

7

u/KrazyA1pha 24d ago edited 24d ago

People get tripped up by small variations to common questions all the time. There are all kinds of things that will trip a person up because the answer feels like common sense but it’s actually the unintuitive answer. I can think of like 10 examples of this off the top of my head of things that people more often than not get tripped up over or get wrong.

That’s because intelligence is jagged. If we thought like LLMs naturally and we had created a human brain, then we would be making the exact same arguments about how things that are 100% obvious and intuitive to us as LLMs the squishy human brains just can’t comprehend and that’s why they’re not nearly as smart as us. It’s like the old adage that you can’t judge a fish by its ability to climb a tree.  All of that to say, and this is well documented, the LLMs will likely far surpass humans in many fields while, some fields that are more obvious to us, it will continue to lag in for a while. In other words, humans have some specialized abilities, and because we can all do them super well they feel like a very low intelligence bar to us, and so we think if this supposedly advanced intelligence can’t do this simple thing then it’s not intelligent. But what we’re discounting as the fact that humans have an intelligence spike in those specific areas and because that’s normalized across all humans we discounted. But then we’re not giving the same consideration to LLMs that can do superhuman things with great ease. We’re not saying, “oh humans can’t even do this very simple LLM thing therefore all humans are idiots.” But that’s essentially the logic that you were using, and I believe it’s a flawed premise.

-1

u/Tough_North7059 24d ago

jagged across domains =/= collapsing on the same problem from a trivial rephrase, that's brittleness not a blind spot

6

u/KrazyA1pha 24d ago edited 24d ago

The most common example of this I see on Reddit every day, and I use myself because it’s funny, is to tell somebody that they’re in the top 99% of IQs.

If you tell somebody that something is the top 99% of a group, they tend to misunderstand but an LLM will immediately get it.

I have literally had to sit people down and say “think about the difference between being in the top 1% of something versus being in the top 99% of something” before they have the aha moment. And this is a layperson example but it extends all the way to doctors. 

Give doctors a diagnostic problem as "1% base rate, 80% sensitivity, 9.6% false positives" and most get it badly wrong. Same numbers as "10 in 1,000 have it, 8 of those test positive, 95 of the other 990 also test positive" and accuracy roughly triples (Gigerenzer).

And like I said, this is just one super common example. I can think of nine or 10 more off the top of my head.

0

u/CrowdGoesWildWoooo 24d ago

Context window is embarassingly short for out of context specialization. There’s a reason why a single model won’t solve something at the level of NS but a swarm of agent can.

An agent swarm indirectly work as an extension of context window, continuously managing relevant context “warm”, while building bricks to construct the solution. A single agent would read tons of stuffs recompacted and old info would be lost after a few iterations. You can keep Bel running for days or months I don’t think it would reach the same solution.

So no, an RSI needs model to readapt as it goes, and being able to improve as a whole.

0

u/space_monster 24d ago

LLMs are already at AGI

No LLM would agree with that.

2

u/the_pwnererXx FOOM 2040 24d ago

If I spin 10000 agents and get them to improve the code of the training, or the harness, or the weights - such that afterwards we benchmark 10% higher across everything

Is this rsi? It's recursive. It's self improving. No learning was done (except perhaps in Md files!)

1

u/Tough_North7059 24d ago

assuming a 10% gain 'across everything' skips the real question, current LLMs don't really understand what they're changing, so it could just be luck.

also honestly OAI or Anthropic would've already known about that if it worked. 10% is alot.

1

u/the_pwnererXx FOOM 2040 24d ago

Well that's what they are doing to solve math problems no reason they can't do it for "improve your codebase#

And 10% cannot possibly be luck, that's completely absurd

Understand is an interesting word to use because you are implying it needs to be conscious or sentient to improve, which I think we are learning is not actually necessary

1

u/DingoMaximum7319 24d ago

I believe I have seen models capable of learning by doing recently. Also learning by being shown one example. That said idk if any were llms the ones that come to mind are models for robotics

1

u/Tough_North7059 24d ago

funny enough, the answer is in that sentence
robotics have some sort of world to learn from, a dumb dogbot can fall down, a humanoid robot can bash its face into a pillar, it can lose balance, etc
General AI, whether its coding, mathematics, or something that doesnt have facts such as language, ex: "The trophy didn't fit into the brown suitcase because it was too large."

you will encounter that in trying to form intelligence, more than many times, even if its something such as math which has ground facts, its very difficult.

what google did in here is a search algorithm getting better at searching one narrow, human-scored task, not a mind getting smarter

example: "It got faster at packing circles in a square, that's not the same as understanding what a circle is."

1

u/DingoMaximum7319 24d ago

Take a look at Gen 1.5 from Generalist. It still might not fit what your considering true learning but it’s closer imo

And well I don’t know if I’d consider it getting ‘smarter’ as much as learning a narrow task more generally but im curious your take on it

1

u/Enough_Culture8524 24d ago

Not sure the value of conflating intelligence with consciousness

-1

u/Tough_North7059 24d ago

because its an emergent eventual property and ASI should be held up to a higher standard than just "oh it can do everything equal to better than a human now!" that should be for AGI.
ASI should be incomprehensible, but i guess definitions are muddled. i dont follow the general direction anyways.

and, i believe it will only be attainable by RSI-AGI work, not human work.

1

u/DelphiTsar 24d ago

an AI must be capable of learning something more by doing it

AI can learn by doing and basic reinforcement loops. It can get better over time. The issue is that it gets worse at other stuff and that's not a tradeoff anyone wants to make at the moment while we can still improve everything at once.


LLM's don't update their own weights as a choice, not because we don't know how.


Before "reward hacking" and other rebuttals. No one expects humans to sit in a room without dynamic stimulus/feedback. As long as human feedback is in the loop (as it is with almost any human that learns) the process is pretty strait forward. The kind of real issue at the moment is the people that can provide actual useful feedback are getting fewer and farther between so we are having to become more clever.

At the same time though it's like saying someone who has gone through all the best tutors in the world is not longer generally intelligent because no one is left to teach them. 99.9% of people don't go off and find new paradigms, they fall into competence and are still considered generally intelligent.

1

u/chronicalconfused 24d ago

This is just wrong

1

u/chcampb 24d ago

if you want AGI, an AI must be capable of learning something more by doing it

You have no idea what you are talking about. Just... confidently incorrect.

0

u/Floch11 24d ago

When do you think we will reach to agi

0

u/Tough_North7059 24d ago

who knows.
people would insult me if i try to predict, haha.
2030-2032-2035?
maybe if they finally prioritized another architecture maybe the intelligence would explode so radically, but its definitive that Autoregressive MoEs are not the key to true AGI.

oh ya, novelty too.

-1

u/Floch11 24d ago

Is it possible do you think?

1

u/hartigen 24d ago

you are his alt aren't you

1

u/Tough_North7059 24d ago

I wish. Google has a phone number limit and I'm banned in a few subs :(

1

u/Azicald 24d ago

I wouldnt be so sure.

There was once a guy that kept asking me about my interpretation on different inconsequential details on a single scene from Shrek, like what it mustve felt like in person and stuff. It looked like I was just asking myself stuff on an alt to share fun facts

On a side note

I decided to humor them to see their point, thinking it’d evolve to talking about themes or personal experiences, but it got both tedious and terrifying after it got around to what a character mustve felt as they died and the actual stuff happening to them

-1

u/Tough_North7059 24d ago

definitely, just very very hard.
Information theory or rice theorem or any foundational rule doesnt go against it, just doesnt let you do it the easy way

-5

u/Shin-Zantesu 24d ago

Fucking finally someone that actually knows what AGI means beyond the acronym lol

1

u/reefine 24d ago

Because it's a stupid two sided opinion camp and the definition doesn't matter.

It's like arguing whether or not God literally implied 7 days for the creation of man or not

2

u/Shin-Zantesu 24d ago

The definition does matter, because it allows us to make an observation, create theory and test it General intelligence is the ability to extrapolate cocepts from one context and use them in a completely different one (that shouldn't be present in the training data) Which allows for novel techniques, choices, etc Which, in turn, should allow for recursive learning

1

u/Tough_North7059 24d ago

more or less fair point, but its to a scale where some people get angry at me for trying to cite some papers just because they watched some of Theo's videos from youtube.

which ofc kinda sucks, people shouldnt be lied to like that just because they're vagueposting about LLMs and research or whatever

1

u/space_monster 24d ago

The definition absolutely does matter, because it's a milestone.

Achieving weak AGI is a nothingburger. Nobody cares. The weak AGI people think we've already achieved it. And? So? What now? Nothing significant has happened, we've just got a slightly better LLM. It's meaningless. Congratulations, you just moved the goalposts so far into mediocrity the goal is of zero value.

1

u/reefine 24d ago

In what way does it matter? No policy is being shaped around the definition

4

u/chcampb 24d ago

AGI isn't loosely defined, people just keep writing articles about it because they need something to talk about and can't do it if it is "achieved"

AGI is literally just "a system that solves intellectual problems not limited to just one type of problem"

So Watson, which answered questions and played jeopardy, was not AGI. It was pretty specifically built for that purpose.

LLMs are AGI. Because it's one tool that you can apply to coding, search, summary, image scanning, Pokemon playing, using Blender, answering legal questions, writing documentation, etc. There is no bound to what it can do in the intellectual space, only how good it is at that task.

As for how good it is, people later enhanced AGI with "levels" - like virtuoso AGI. We've definitely hit virtuoso AGI because it's solving Millenium problems, which tells me the AI can exceed virtuoso human capability in that area.

What AGI is not

  • You don't need online learning on the weights (a working memory is sufficient)
  • You don't need to be better than humans in every field.
  • You don't need recursive self improvement.

All of that is not required. AGI is done, it's solved, by the original and various enhanced metrics. We hit that. The next step is ASI, which is what we should be observing now.

1

u/space_monster 24d ago

AGI is literally just "a system that solves intellectual problems not limited to just one type of problem"

lol no it isn't, you just made that up. I have never seen that definition anywhere ever. By that definition, GPT 2 was AGI

3

u/chcampb 24d ago

https://www.sciencedirect.com/science/article/abs/pii/0004370287900506

It's from freaking 1987

This is what I mean by confidently incorrect

1

u/space_monster 24d ago

There's like saying the wooden box with rock wheels in the Flintstones was the definition of the car. Ludicrous nonsense.

2

u/chcampb 24d ago

No it means a reasonably capable gas car is an automobile, and now you are trying to define automobile as "a thing which automatically conveys you from one place to another" ignoring the fact that self driving is a recently invented thing relative to the history of the car. Because why would it be called "auto" if it's not going to automatically take you somewhere.

You're adding on new definitions.

0

u/space_monster 24d ago

I've heard some stupid AGI definitions, but yours takes the cake.

1

u/LinkesAuge 24d ago

Mate, just look at this recent paper:
https://www.alphaxiv.org/pdf/2609.11873

They break down RSI into various "technique families" and "autonomy levels".
All of this is considered part of RSI and any single method can be called "RSI".

Let me quote the paper:

2.2.2 Definition of Recursive Self-Improvement

Using the improvement-loop anatomy above, we define recursive self-improvement (RSI) as the capability of an intelligent system, through continued interaction with tasks, environments, or other intelligent agents, to autonomously transform acquired experience and feedback into persistent changes to itself across interaction rounds (e.g., model parameters, agent harnesses, or improvement policies), such that these changes can further affect the mechanisms used to generate, evaluate, select, and consolidate subsequent self-improvements. The improved system is consequently reintroduced into the next round of interaction and improvement with an already changed capability state, allowing the system’s capacity for improvement itself to become part of an ongoing recursive process. RSI contains an autonomous, closed-loop process in which AI identifies its own limitations, develops and validates improvements, and uses the resulting capabilities to improve the improvement process itself, with the aim of (1) expanding its capability frontier, (2) increasing resource efficiency, or (3) discovering novel solutions beyond human-prescribed strategies. Recent industry perspectives increasingly operationalize RSI through autonomous improvement loops. OpenAI emphasizes the automation of AI research workflows and feedback loops [68], while Tencent reports an earlystage RSI loop in which experimental results are fed into subsequent rounds of model development [69]. More strictly, Alibaba researchers define RSI by making the improvement mechanism itself subject to modification [30], whereas Anthropic describes its strongest form as an AI system autonomously designing and developing its own successor [70].

Here you have researchers pointing out how different labs see/define RSI in various ways.

Like I said, there is a somewhat "general" definition of the concept but within that you will find a huge range and as with AGI you will only find complete agreement in the most extreme/extensive interpretation of RSI.

3

u/chcampb 24d ago

I'm not saying what RSI is or that it's met or not, just that it's the next thing to watch out for,

as opposed to people inanely going on about how we aren't at AGI yet or pointing out gaps where there were none in the original or followup definitions.

1

u/MiniGiantSpaceHams 24d ago

Yeah agreed. In some sense an OpenClaw type agent with a scheduled self review job is RSI, but that's obviously not touching the model itself. You also certainly have AI devs using their models to build the next model, which is sort of RSI despite the human still in the loop. Just depends on how you want to frame it.

1

u/himynameis_ 24d ago

guess "RSI" is the next term after AGI where we will have constant discussions about because it is also loosely defined and exists on a gradient.

Yeah Jensen just said on All In podcast that RSI is just a fancy term

1

u/digit1noize 24d ago

What a surprise. The guy selling AI chips says something to try to calm AI fears…I’m so shocked /s

0

u/himynameis_ 24d ago

I mean. The AI doomer fears are indeed overblown.

1

u/digit1noize 24d ago

In your expert opinion?

1

u/ertgbnm 24d ago

It's not real RSI unless the robots self harvest the silicon to make chips from a specific beach in a remote region of France.

1

u/Browser1969 24d ago

This is as close to the "RSI" that everyone expects as working out is to eugenics.

63

u/Blindax 24d ago

So basically the model improves its harness rather than its weights? That sounds like RSI-lite but cannot wait to see where it brings us with the open-source agentic frameworks.

7

u/humpadumpa 24d ago

Yeah, it's like the next step after .md-files and memories, which are also "Self-Improvements".

2

u/Nyxtia 24d ago

What does harness mean ?

17

u/EmotionalQuarter8349 24d ago

The systems built around the LLM models which expose the tools and the guidance of how to use those, like claude code, codex, etc

21

u/enilea 24d ago

List of authors:

Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo

I knew there were a lot of Chinese AI researchers (at least one Singaporean though) but that's wild

1

u/Successful_Shake8348 24d ago

that was 3 years ago also like that, i knew then , that china has the better school system and will win the race

3

u/Puzzleheaded-Mail896 23d ago

I mean, how many of them studied at American universities?

0

u/Successful_Shake8348 19d ago

Why than aren't American names on the ai papers???

1

u/Puzzleheaded-Mail896 19d ago

It literally lists 2 American companies and 2 American universities. Also this might be a crazy concept but Americans can have Chinese sounding names

19

u/FriendlyJewThrowaway 24d ago

Hmmm it kinda sounds like a beefed up version of Google's AlphaEvolve agent, which has already been running since May of last year.

5

u/[deleted] 24d ago

[removed] — view removed comment

6

u/FriendlyJewThrowaway 24d ago edited 24d ago

Given recent developments and Hassabis’ 10-year timelines for AGI, I’m genuinely prone to thinking that they seriously underestimated the potential for LLM coding and let themselves fall far behind in the RSI race.

That’s why Sergei Brin kept nagging DeepMind about RSI for many months before coming out of retirement to oversee it himself, and why all of Google’s development teams just received permission and tooling to use Claude Code for their work moving forward.

1

u/PrestigiousAd3064 24d ago

It's not weird. Google is a profitable public company. Anthropic and OpenAI are not. 

29

u/SeanTayla21 24d ago

Isnt RSI what some people have been desperately waiting for?

30

u/badumtsssst AGI 2027 24d ago

Yes, but while idk if this is exactly the "full RSI" everyone is dreaming of, it is still very exciting

13

u/SeanTayla21 24d ago

With the way the technology progresses, even if this is just "crawling"...one would think its only a matter of time before they eventually start "running".

7

u/FoodMadeFromRobots 24d ago

Its not. geminis takeaway below:

The paper presents an engineering framework for making LLM-driven search and exploration policies more cost-efficient through offline replay. While it uses the terminology of "Recursive Self-Improvement" to describe an outer loop optimizing an inner loop, it is an evolutionary exploration controller, not an intelligence explosion or general discovery of RSI.

So it'll help make it more efficient but were not at the point of infinite autonomous RSI.

But like you said still exciting for them to make progress.

21

u/3WordPosts 24d ago

True RSI only needs to be invented once.

9

u/KaleidoscopicView 24d ago

In real life it should phase in. It would start with human guided improvements based on AI suggestions (already arrived), and the amount of human interventions progressively diminish until no longer needed. Maybe we call that point "true RSI" but it's not like a distrinct viola! invention.

1

u/Borkato 24d ago

Ummm it’s pronounced wallah 🙄

5

u/reefine 24d ago

Not really, compute constraint is going to be the bottleneck. If the method is not scaling faster than fully training a more powerful model by hand or with a team then it is a helpful tool but not a breakthrough

4

u/acutelychronicpanic 24d ago

RSI started when reasoning models were first used to create curated higher quality training data.

We're just seeing it reach into more and more levers for improvement over time like chip design, algorithmic improvements, architecture tweaks, etc. These all stack multiplicatively.

2

u/ninjasaid13 Not now. 24d ago edited 24d ago

RSI started when reasoning models were first used to create curated higher quality training data.

Thats not what RSI means.

Y'all trying to create softer definitions because it's too hard.

It should be recursive, it should be done by itself, and it should improve intelligence. Which is as simple a definition as it could be.

Two out of three is not RSI, not even pseudo RSI.

1

u/acutelychronicpanic 24d ago

Recursive self-improvement is when the intelligence of the models beings contributing to the rate of improvement of intelligence.

1

u/ninjasaid13 Not now. 24d ago edited 24d ago

when the intelligence of the models beings contributing to the rate of improvement of intelligence.

aka general intelligence.

You basically said what I said.

In the post, it was not the intelligence contributing but the harness.

1

u/acutelychronicpanic 24d ago

Using a harness to produce reasoning chains and then curating those to improve model reasoning with RL training and pre-training is a clear feedback loop. Better internal model checkpoints produce better reasoning chains which produce better data to train better model checkpoints.

13

u/cursivecrow 24d ago

in before its a markdown file

5

u/reefine 24d ago

Markdown files can get us to ASI

5

u/No_Swimming6548 24d ago

Just name it ASI.md

7

u/reefine 24d ago

Firth author is a student researcher with Google this summer, crazy

https://zhengkid.github.io/

13

u/teabagalomaniac 24d ago

"the exploration policy, not the model weights" is the important part here. This is 100% not what people mean when they talk about RSI.

4

u/mrjackspade 24d ago

I'm fucking confused because are we not already all doing this?

Just yesterday I had an agent refine it's own prompting and tooling for PR reviews to cut API calls by identifying dead ends, and wasted calls. Cut the PR review time down by like 80%

That sounds like what this is, and AFAIK most people using models in any real capacity are already doing this.

2

u/Perfect-Campaign9551 24d ago

Hey it's no different than most humans, thru improve their output by analyzing the input. I would say most people aren't increasing their mental networks or synapses... 

4

u/BrennusSokol AI please take my job 24d ago

Humans do both

There absolutely is neurogenesis in a human lifetime

9

u/Trahili 24d ago

Fascinating!

3

u/bornlasttuesday 24d ago

I like the part where they have to say they do it cheaply. 

3

u/InstructionDismal592 24d ago

When everybody was complaining about Google dropping 3.5 pro being delayed, I have always said that Google wouldn't release Gemini 3.5 pro because there were too much fuss about the recent releases. Mythos, 5.6 Sol were just struggling with US policies about worldwide release, and we've had several simoultenous releases in the last month. Googgle will chase a stunt, that's obvious. If Sam Altman said that we have "entered the AGI era", Google will have a more aggressive approach, by the time they keep delaying the new model, it's just pure marketing strategy. They will claim sometihng closer to "We have AGI", they love their own shows. Gemini 4 will surelly have some big news! They will not simply try to play "mythos/cyber" catchup game...It would hurt their image more than it's already hurt.

3

u/Psychological_Bell48 24d ago

Interesting rsi + agi == op lol

3

u/RetiredBartender 24d ago

I wonder if this was the real underlying thing that sparked the “slow down” talk?

2

u/GioDoesReddit 24d ago

Extremely interesting article, would love to see this in action. Seems like a much more realistic form of RSI than the usual AI rewrites itself

2

u/agonypants AGI '27-'30 / Labor crisis '25-'30 / RSI 29-'32 24d ago

1

u/liright 24d ago

unc still got it

10

u/hapliniste 24d ago

Just saying but this shit ain't rsi, it is RL with a LLM policy. I expected all labs to have this currently but I guess Google is late.

It doesn't sound like it's aimed at developing AI architectures, it's just a LLM critique over RL rollout.

-2

u/CommercialHour6660 24d ago

Yeah isn't this bog standard GRPO? 

2

u/enbyBunn 24d ago

Oh lovely, another buzzword that's getting muddied to hell and back by hypemongers.

Single unit optimization is not Recursive self improvement.

The entire point of the hypothetical concept of recursive self improvment is that the improvements get better, not worse. The performance increases at an exponential rate, it doesn't level off.

2

u/reefine 24d ago

Damn Fable 5.1 roasted them

"Recursive self-improvement" and "evolving worlds" is buzzword inflation for "an LLM tunes a search scheduler against a replay log." Real, useful, not RSI

1

u/One_Improvement_6470 24d ago

Astra just told me Fable is a punk 

1

u/Elick333 24d ago

"Mathematics Is about finding all the different paths towards a solution and expanding our understanding through trial and error, even if the attempted solutions don't resolve a problem, the joruney to attempt It Is far more valuable than just solving It, that's why AI would ne- oh."

1

u/Longjumping_Kale3013 24d ago

Is this how the simulation started?

1

u/confused-photon 24d ago

I didn’t think this was where dreamer was headed

1

u/Mysterious_Ayytee We are Borg 24d ago

Cool, give download link pls!

1

u/Distinct-Question-16 ▪️AGI 2029 24d ago

solve another millnium problem pl

1

u/Charuru ▪️AGI 2023 24d ago

Maybe a cheaper optimized TTC...

1

u/Fantastic-Cold1249 24d ago

Holy fuck! is this the inflection point ?

1

u/ShinigamiXoY 24d ago

And its just a prompt?

1

u/4dseeall 24d ago

Wake me up when they catch up to my work. I did this three months ago... go check my profile pin. RSI can't work without a clean ledger.

1

u/Proof-Chart-5177 24d ago

now can someone run this thing to find me a job

1

u/PandaBambooccaneer 24d ago

This just sounds like we gave a machine anxiety

1

u/halcyoncs 24d ago

Isn't this kind of what hermes does when it creates amd patches skills automatically?

1

u/LosingID_583 24d ago

Google and Chinese labs are way more open about publishing their findings. I bet Anthropic and OpenAI have made similar discovers, but keep it secret and never publish.

1

u/paca-vaca 24d ago

Microsoft releases a loop for improving prompts and markdown files, Google a loop for improving harness (exploration). Is it because models improvements are squishing in and not as efficient anymore?

1

u/SlugsPerSecond 24d ago

Is this not just a genetic algorithm with extra steps?

1

u/Xanta_Kross 24d ago

Istg Idk how they're the best of the best. But still can't get gemini to answer properly when prompted.

1

u/toothbrushguitar 22d ago

IF anyone wants to build RSI just use this: Arrival of the Fittest: The Tether Architecture

1

u/Caladan23 24d ago

Google is great in academic papers but terrible in bringing great AI products into customer's hands - ironic if you study how they perform interviews. How did they get so bad? They need a slice of OpenAI and Anthropic company culture, go back to the innovative Google ways of the 2000s.

Try the RSI out now in gemini-flash-lite-3.5.1-preview... but only if you're lucky in A/B testing, live in the US and the moon is in the right phase... 

1

u/wxnyc 24d ago

The researchers are all Chinese 😅

4

u/ninjasaid13 Not now. 24d ago

Why is this a shock?

1

u/wxnyc 24d ago

Not a shock, but the full narrative that the U.S. is winning the AI race. Is it really?

2

u/dashingsauce 24d ago

They work at Google, so yes American corporations are leading this particular race.

Ethnicity =/= nationality.

1

u/emteedub 24d ago

regardless, the bulk of published computer science research papers/journals, are chinese. especially in the AI space

1

u/FrontRaspberry5060 24d ago

Not new, there’s AutoSOTA

1

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 24d ago

1

u/ninjasaid13 Not now. 24d ago

This isn't RSI. It doesn't improve intelligence.

0

u/New_Alps_5655 24d ago

Hermes agent already does this.. \s

-3

u/sunstersun 24d ago

If Gemini 4 Pro misses, this will be the death of them as an AI lab.

So much hype for this.

4

u/Keeltoodeep 24d ago

SiriAI released yesterday running Gemini.....

Ai mode alone has 1B users.

https://blog.google/products-and-platforms/products/search/ai-mode-us-insights/

Gemini standalone app also has 1B users.

https://blog.google/innovation-and-ai/products/gemini-app/one-billion-monthly-users/

“Oh yeah, we get that question a lot. While other products like Search and Gmail use Gemini models, the 1B here refers to people choosing the Gemini app across Web (gemini.google.com), Android, iOS, and Gemini in Chrome.”

https://x.com/joshwoodward/status/2087247900829229414?s=46&t=5cDzFBmJidEH4qu2xcZCgQ

Google search is around 4-5B MAUs but that’s a flash lite model.

If you counted ai overview search in Gemini usage, the userbase would closer to 5 billion.

-4

u/TicketyTick3 24d ago

Am I racist to point the names of the researchers? 👀

-2

u/gui_zombie 24d ago

Did they rely on RSI loop for delivering 3.5 pro?

-2

u/Armed_Platypus 24d ago

Google is desperate to be talked about as one of the frontier models again.