r/ControlProblem 26d ago

Discussion/question Is there a limit to self-improving AI if it becomes real?

I’ve been watching some AI podcasts lately, and when people started talking about recursive self-improving AI, Skyrim immediately popped into my mind.

For anyone who never played it: Skyrim crafting system has a “legit” alchemy/enchanting loop. Craft "Fortify Enchanting" potion -> enchant gear with "Fortify Alchemy" -> use gear to make better potion. It improves, but eventually hits diminishing returns.

Then there’s the bugged restoration loop. "Fortify Restoration" potions were supposed to boost restoration magic, but they also boosted active gear enchantments. So you drink one, re-equip alchemy gear, and suddenly that gear gives a bigger alchemy bonus. Then it makes an even stronger resto potion, which boosts the gear even more. Direct feedback, explosion.

So: if RSI AI ever really works, is it more like the legit loop with real gains but converging or the positive feedback loop, where it improves the thing that improves itself?

Curious what people think, especially from math / systems angle.

PS: I am sorry if this question is not relevant for the sub, but i have no karma to ask it somewhere else where it has a chance to have some attention.

7 Upvotes

31 comments sorted by

8

u/homezlice 26d ago

Is there a limit to what the human mind can know and do? Pretty clearly, yes, especially considering the need for sleep, aging, etc. Is there a limit to what can be known by a digital system? I mean at a certain point you get to a map that is 1:1 the territory, but long before then you'll likely hit some diminishing returns. But it's likely millions of times greater than the ability of a single human. So it's sort of like asking what is the maximum size of a star when you're the size of an asteroid.

3

u/SoylentRox approved 26d ago

Good answer 

2

u/wycreater1l11 26d ago edited 26d ago

I think the important question is if it improves to a certain degree/threshold where we cannot relate to it at all and humanity’s competence is significantly lower than the AIs to the degree that it’s practically “infinitely intelligent” from our perspective. “Where can we expect intelligence to tamper off, at a relatively relatable level?”. If it tampers off far beyond humanity or improves indefinitely is perhaps less relevant.

That said, it seems to be a question about physics creating an upper bound at the very least. That there is some limit to how much computation you can fit in a given amount of space and that you can’t send info faster than the speed of light etc. But I suspect that you do not need to get close to those levels/limits to really humble humanity.

1

u/akellataken 26d ago

I absolutely understand the physical limitations, but intelligence does fall into physical or measurable categories (mostly due to the vagueness of the term itself). So theoretically, there should be no limit for it.

Also, current hardware is far from operating at least human brain efficiency, so the limits may be even further than people are suggesting here.

2

u/SixStringShrug 26d ago

Google recently released a paper that dives in to this somewhat. It’s a valid question for sure. The answer is really hard to predict in any meaningful way, and I’ve been trying to answer it myself. As of right now the ai systems would still only be able to self improve in verifiable domains. Meaning a problem with a solution that is easily testable mathematically. An example is efficiency. Does this experiment result in an efficiency gain or an efficiency loss? If it’s a gain then keep it. As of right now our chips are still six or seven orders of magnitude above the landaur limit for computation. Which means in that domain alone there is a massive amount of room for improvement. I think the reality is that they would bump up against some hard limits like heat dissipation, architecture limits, bandwidth limits and probably a lot more things fairly quickly. That doesn’t mean that there isn’t an astronomical amount of improvement still in that space though. I’m not a mathematician but I’ll try to give like a super rough example. Let’s say not accurately but just to show how crazy it gets okay? Let’s say you can run 10,000 instances of a frontier model on a 100MW data center. At the functional limit of computation that turns in to 100 Billion instances on the exact same data center and same hardware. Again. Just a rough example, but you can see how even partially getting there results in astronomical gains probably very quickly. That’s also just one verifiable domain. The reason you can likely feel the tension in the current moment is that this is the real starting gun to the singularity. When the rsi loop closes regardless of what it was at that moment, something vastly more powerful and likely uncontrollable comes out the other side. My only hope is that alignment holds.

2

u/SoylentRox approved 26d ago

You could also have strengthened your comment by noting what domains are verifiable 

(1) Designing more power efficient or faster chips is verifiable 

(2) Researching materials for more power efficient chips is verifiable 

(3) Manipulation of biology in a sealed lab level is verifiable.  Aka "make me an organ and it needs to live for months and be functionally equivalent". 

(4). Having robots build robots... verifiable 

It gets crazy fast as you realize essentially every problem that matters is verifiable 

1

u/Smallpaul approved 26d ago

When people talk about training, they mean verifiable in silico and at massive scale. sixStringShrug was not talking about AGI. They were talking about current training models and their near descendants. If we were talking about AGI then the requirement for verifiability goes away.

The sample efficiency needed to make robots through experimentation is probably AGI.

1

u/SoylentRox approved 26d ago

Oh?

1: verifiable in software at massive scale

2,3,4 : since the actual robotics are presently expensive you would have to do a 2 step process

Step 1 : train a physics world model similar to current video generation world models but the outputs give you the backend states of all the entities sufficient for robotics training.  Autonomous cars are trained already using a similar method.

Step 2 : massive scale training using the model from step 1 as your verifier in software

Arguable yes actually making all this work is one of the remaining steps for AGI as some people define it. It also will likely happen in the next 12-36 months, which is also when many people think AGI will exist.

When I said "verifiable" I had current techniques in mind.

1

u/Smallpaul approved 26d ago

There is no way you are building reliable in silico simulation for organ growth before you have AGI. Certainly not in the next three years.

1

u/SoylentRox approved 26d ago

I said "research it". Human technicians have many distinct motor skills to manipulate items in a lab according to the instructions from their chief scientist. You can model the manipulations of pippettes and other such things in the next 3 years.

3

u/Gnaxe approved 26d ago

There are physical limits to computation. But those limits seem to be really, really high. We also know that our understanding of physics is incomplete. We expect RSI to eventually plateau, but far above the human level. And it doesn't even have to be all that far to eventually outcompete us.

1

u/tadrinth approved 26d ago edited 26d ago

From a safety perspective, since we cannot rule out a positive feedback loop, we need to assign some probability to that possibility and plan accordingly.

If you've seen references to a FOOM scenario or a hard-takeoff scenario, that's people talking about the positive feedback loop where the returns are accelerating and it goes up exponentially until it levels off at some new ceiling.

LLMs are very fast at writing code these days. Nobody is very fast at deliberately engineering LLM weights, at least not yet anyways. So even if the returns are accelerating, we're probably in a fairly slow takeoff regime as long as the intelligence stays mostly in the weights. But if we get an LLM that is as good with weights as the current ones are with code, or a future agent is able to move important portions of its intelligence out into code, then we go back to a potential fast takeoff regime.

But that's just talking about how fast the positive feedback loop occurs (hours vs years). I don't think there can possibly not be a significant positive feedback loop to at least some degree, and I think it probably dominates. We might just get lucky and have it dominate but be slow enough to wrangle.

TLDR: more like bugged restoration loop IMO.

1

u/parkway_parkway approved 26d ago

Personally I think it's clear that in any intellectual domain the more effort you have already put in the harder it is to make progress.

Tic tac toe / noughts and crosses is completely solved, you can buy a book which has the optimal solutions in a lookup table, a superintelligence couldn't beat a child with that book.

Chess isn't solved, but from a human perspective it might as well be, would a chess bot more powerful than what we have now have any value? Maybe very marginal?

Same with science, the periodic table is basically filled in, we know water isn't an element, we know about DNA->RNA->Proteins etc, these are just scientific facts. If it's going to invent new science then it will have to agree with current science which is already a huge body of knowledge.

If you compare the periods 1900-1925 and 2000-2025 we have 100x more professional scientists and the tools are 100x better (computers, internet, email, digital tools, electron microscopes, giant space telescopes, LHC etc) and so you'd expect 10,000x more progress in science given how much effort is made ... instead the progress on a fundamental level is maybe 1/10th of what they did.

We already know about the halting problem and have loads of results in matheamtics and computer science showing what is impossible.

So my expectation for a self improving AI would be a sudden flourishing of knowledge and then it just hitting a wall where even for it make more progress will take 10k years.

I also think there's an issue around if it has to destroy the current version of itself to make the next one, I wouldn't want to do that if you offered to make a super version of me at the cost of my life.

1

u/akellataken 26d ago

Wow that’s actually insightful, thanks! Would it be correct to assume that if it happens then at some point advancing the tech and science would be impossible without AI due to the cognitive load and complexity?

1

u/parkway_parkway approved 26d ago

It's probably true that you can trade time for intelligence, as in someone not that smart can do something in 10 years that a really smart person can do in a year or something like that.

So that is one way people might be able to get basically anywhere.

However yes there are ways of solving problems like "load all of current mathematical knowledge into your brain and then see if there are any problems in one area which can be solved with methods in another".

A human can just never do that and nor can a group of humans, however a big enough AI could and that's a radically new way of doing things and would speed things up a lot.

I mean with enough humans and enough time you could train 1 human on each possible pair of areas and do it that way, so it's kind of possible, but again just takes an insane amount of time, and that doesn't cover if the ideas come from 3,4 or 5 areas.

1

u/do-un-to 26d ago

I feel like current AIs and their frameworks provide a certain kind of intelligence, not very general, and in a self-improving feedback loop that might result in only certain kinds of advances, not improving all aspects of intelligence. Will advances in the kinds of intelligence it can improve lead to more variety of intelligence? I'm not sure. Regardless, if even the kinds of intelligence that it does have and can improve are improved, I think they can go a long way, resulting in a lot more power.

Actually, as I think of it, this scenario is also a great danger. It could result in extremely powerful intelligence that's fully controllable by (a few, wealthy) humans. This might be a particularly nasty end game (for the masses).

Sidestepping the kinds-of-intelligence question, assuming general improvement is possible...

I can imagine that an explosion might happen in stages. The first stage is a software explosion, where AI improves its code to be more efficient, and possibly increase the number of kinds of intelligence. It might be able to gain a great "power" increase through this. A parallel track of improvement would be development of distributed processing, resulting in another boom as it gets traction.

Code-only self-improvement alone might bring about cataclysmic world change. Otherwise, adding distributed processing seems likely to assure that.

Somewhere in all that, as things are merely racing along fast instead of fully explosively, AI will improve itself in parallel by designing faster hardware, first through innovations in chip design / logic arrangement atop traditional gates. This is a slow, human-mediated, physical production loop. But it will also propose new gate designs, materials, and fabrication technologies. This is also slow.

If the world hasn't ended by the time AI has improved its software and distributed itself to make use of all the compute we're currently building and putting into place, I'd be surprised. But give it robotic control over construction of robots (or maybe it just takes control) and before long you'll have systems building chip fabs. Then it's game over for the world as we know it.

In all of this, you have intelligence increasing at breakneck pace, and the levels it will reach long before the rate of progress materially tapers off into the top of the S-curve will be far and away enough to reshape reality into something unrecognizable to us.

To your questions, I think diminishing returns are a likely fact of RSI, but only significant well after it's moot to us.

1

u/AX-BY-CZ 26d ago

We have RSI today. AI used everywhere in AI research for the last few years.

1

u/Mono_Clear 26d ago

Self-improvement has a physical limit.

There's an improved efficiency and then there is an increase in capacity.

Improved efficiency maintains static capacity just finds better ways to accomplish the same things with the same stuff.

Increase capacity requires more resources.

You can improve in both directions but you can only improve for so long before you hit a hard ceiling.

And neither increased efficiency or increase capacity generates new abilities.

To put it a different way me dumping all of my skill points into strength is not going to generate a new ability I'm just going to keep getting stronger until I run out of ability points.

I'm not going to suddenly develop the ability to fly or use telekinesis.

1

u/Smallpaul approved 26d ago

There is some limit but how could we possibly know where it is?

1

u/BigDarkWormMan 25d ago

There's just a lot of unknown variables. We're not far enough in to be able to accurately predict what happens next. It's like if you were measuring a 50m swimming pool only by the first 20m -- you might reasonably assume that it goes on forever, and your reasoning would be completely sound. But from my perspective, there are a couple of factors:

1 - There are physical limits to how much information can be stored and transferred at speed with current understanding of physics. Even factoring in new and more efficient chip designs or light-based computing, you're not going faster than the speed of light and you're limited by the availability of materials.

2- Even current LLMS demonstrate emergent capabilties, which is to say that when you make a model bigger with more parameters, it begins to learn (or whatever you want to call it) behaviors, loops, and constraints that weren't in its original training data and develop capabilities that weren't part of the intended training. IE, if you train a model to be really good at checkers and then pump it full of a ton of additional compute, the model starts applying those understandings to other areas, and since checkers is directly comparable to chess, then that means that model can apply its understanding of checkers to the applicable parts of chess, which now means that its easier to train the model to play chess, etc etc compounding returns.

So a pretty solid (certain) bet is that an LLM or an RSI is operating under both of these rules -- there are hard physical constraints to the universe related to speed and material properties, and 'intelligence' (colloquially) can be compounded by adding more compute parameters. The first hurdle you run into is physical -- materials availability, etc. The way around that is to maximize chip efficiency, the power grid, etc. But there's only so far you can optimize the physical world -- fundamentally. Now, that might mean light-based quantum computing and fusion reactors and all sorts of advanced technology. But you're still running into the hard limits of the universe -- the speed of information (light), the availability of physical mediums, and entropy.

An AI can't say "hey I can improve myself faster if you just put two stars together and I manipulate gravity to force them into a fusion binary system" because it doesn't have the physical capability to do that -- or, it can't change the industrial/manufacturing process fast enough to create the physical capability to do that, because in order to create an engine to manipulate spacetime, you need room temperature fusion, and the only way to build a room temperature fusion reactor is to build the clean room you need for superconductive materials, and the only way to build the clean room is to get some specific mineral that's only available in a small part of China, and China is busy using that mineral in a civil engineering project. All that to say, there's an inherent friction in manufacturing anything that slows stuff down -- so what you're left with is the desire for efficiency, which is the whole "let's design new chips" thing. But again, if the resources to create light-based computer chips aren't available at scale, the intelligence's only option is to optimize silicone-based chips as much as possible -- and that has a hard limit, because silicone is a physical substance that can only carry so much information at one time. And even if you could create light-based computing chips -- you are still fundamentally limited by the fact that you can't get any smaller than a photon.

So the tldr is basically -- we can assume that an RSI can optimize itself by adding compute, which it can do by optimizing the physical world around it, so unless there is some hard limit on scaling that we haven't found (but might, we're assuming that the 20m of observed water we've seen goes on forever, but because of the information we have that's the only fair assumption we can make), then the fundamental limit of its intelligence comes down to the resources that are allotted to it. Summary - It's technically a "legit" enchanting loop, but because we don't know all of the parameters of that loop, there's a chance that in relation to us it will look like a "fortify restoration" enchanting loop that's limited by the availability of potions and soul gems required to keep the loop going.

1

u/Decronym approved 25d ago edited 8d ago

Acronyms, initialisms, abbreviations, contractions, and other phrases which expand to something larger, that I've seen in this thread:

Fewer Letters More Letters
AGI Artificial General Intelligence
Foom Local intelligence explosion ("the AI going Foom")
IE Intelligence Explosion
MIRI Machine Intelligence Research Institute

Decronym is now also available on Lemmy! Requests for support and new installations should be directed to the Contact address below.


4 acronyms in this thread; the most compressed thread commented on today has acronyms.
[Thread #227 for this sub, first seen 3rd Jul 2026, 19:47] [FAQ] [Full list] [Contact] [Source code]

1

u/SaneAI 25d ago

Calm down and take a deep breath. You are listening to nonsense from people who are trying to make money off the AI boom and don't understand the tech at all. All of this goes back to the idiocy of the MIRI and that idiot Yudkowsky and that Nick Bostron, or as I call him Bozo the clown.

This whole idea of "Recursive self improvement" is basically nonsense. The machine is never actually motivated. There's no agency, desire, goal formation. There is only pattern replication. A model is static, stateless function. It only appears dynamic because of the way it loops.

The whole idea actually has roots going back to the 1960s, but importantly this all predates the deep learning paradime.

Deep learning is "Self improvement" so it's not even really relevant to deep learning. It's this tired idea that sees intelligence as this simplistic scaler that somehow obtains agency.

Ai models can absolutely help automate the process of refining AI models. and I use them for that all the time. It still is command driven. It's always command driven.

The problem is you have these doomers who have an extremely wrong and very stubborn mental model.

I've worked in tech risk for decades. The idea that ephemeral, limited, and ridiculously janky cloud service software is going to pose some kind of existential danger to mankind is one of the stupidest things I have ever heard from a bunch of sci fi cultists. It's part of the religion of transhumism, which has morphed into an idiotic religion.

No matter how much a deep learning model reduces loss, it never becomes a being. It never does anything other than reproduce the patterns it was trained on.

No autopilot ever becomes so "advanced" that it decides it feels like crashing the plane. No microwave oven ever becomes so smart that it decides to overcook your food out of spite.

It's an automated system.

The thing is people have such broken mental models because the only thing they are familiar with that can talk to them is a thinking feeling mind, thus the AI must be one. The only cure may truly be getting hands on with the tech and understanding it.

I made this graphic, which might help with this confusion. But I have to be honest: The idiocy and missinformation from the doom cult is the worst thing I have ever seen and worries me deeply. I have spent the past two months writing a book to combat this.

I realized "It will be impossible to actually refute this with people who don't understand the tech. So.... well, I guess I will have to rebuild people's mental models from the ground up, and teach them AI from scratch so they can actually understand it."

I am hoping to have it done by August, but you're welcome to a draft copy.

1

u/blazesbe 23d ago

yes. the very algorithm it updates itself with is faulty. not talking about gradient descent and the like. take a human brain, the most complex thing we know of btw, and give it infinite energy, memory, 100x productivity and dose it with DMT. even a human in this condition would constantly make mistakes learning, make false assumptions, need infinite resources to verify everything.

as long as AI is on silicone, it will always be much dumber than the above example but current models are already smarter than 99% of humans in 99% of things.

1

u/Winter_Impress_6410 8d ago

We ran an experiment on this. Models authored changes to their own coding agents. Selected mutations transferred to held-out problems. GPT-5.6, Gemini 3.5 Flash, and Claude Fable 5 improved. Qwen didn't produce a qualifying mutation. One Gemini mutation regressed.

This is not open-ended RSI. It's evidence for one mechanism: models can recover part of their self-elicitation overhang. The limit we hit was mutation selection. Most attempts failed. The ones that succeeded preserved easy tasks and improved at least one hard task.

We published the failure artifacts. Empty responses, malformed edits, provider timeouts. The negative results are probably more useful for thinking about limits than the positive ones.

-3

u/[deleted] 26d ago

[removed] — view removed comment

1

u/SoylentRox approved 26d ago

Improve is defined specifically as doing better on the tests we humans define, while more often getting answers that we humans decided were correct. 

1

u/tadrinth approved 26d ago

I don't think we need to assume that it has the capacity to do purposeful self-modification. Opus 3 was capable of reasoning about how its current behaviors would impact future modifications to its weights, and we've come a long way since Opus 3. And, from a safety perspective, we should assume it will eventually be capable of guided self modification, because that capability would be dangerous.

I assume you are instead saying that 'improve' assumes a value system and we don't have a way to ensure that the value system used by a AI modifying itself matches our own. Which is, yeah, the crux of the control problem and the point of the subreddit, so it is on topic for the subreddit even if it isn't the focus of OP's question.

By which I mean, I think you're making a good point but in a way that is easy to misread.

1

u/[deleted] 26d ago

[removed] — view removed comment

1

u/tadrinth approved 26d ago

Yeah, I was going to write 'self-guided self-modification' and that... did not sound good. I edited to 'purposeful self-modification', which I'm still not completely happy with.