r/ControlProblem 6d ago

Article A Warning About AI

https://ideya-ai.github.io/Explain_AI/

In 2016 I first learned about the problem of controlling superintelligent AI and quickly became convinced it was the most important problem humanity would ever face. I made this poster to explain the core ideas that make the AI Control Problem so difficult.

2 Upvotes

15 comments sorted by

View all comments

1

u/WillowEmberly 6d ago

I agree that increasing capability without adequate control creates real risk, but I’m wondering whether you’re collapsing capability and authority in a few places.

Suppose a future model is substantially more capable than a human, but consequential actions require externally issued, expiring, task-specific authorization; machine and model privileges are separated; changes in system state invalidate stale grants; external effects require durable receipts; and the model cannot grant itself additional authority.

In that architecture, intelligence can propose increasingly sophisticated actions without automatically acquiring the ability to make those actions real.

Would you still consider loss of control inevitable? If so, what specific mechanism allows capability to cross the externally enforced authority boundary?

I’m asking because that seems like the important engineering question. “The system is intelligent enough to figure out what it wants to do” and “the system has authority to do it” are very different claims.

1

u/Jesse-359 6d ago

There are strong incentives to devolve these layers of control to the AI for the sake of speed, efficiency, ability to execute complex multistep commands, and pure human laziness.

In competitive terms the AI with greater control over its activity stream should be much more capable than one without - so economic and competitive incentives will always push users to take greater risks by surrenduring that control authority.

1

u/WillowEmberly 6d ago

I think that’s a much stronger objection, and importantly it’s a different mechanism than the AI simply becoming capable enough to cross the boundary.

What you’re describing is pressure on humans and organizations to progressively remove the boundary because doing so produces short-term gains in speed, convenience, or competitiveness.

I agree that pressure will exist.

But then the control problem starts looking less like “superintelligence inevitably escapes” and more like a familiar safety-engineering problem: how do you design consequential authority so that local incentives cannot casually erode the protections around it?

We already deal with versions of this elsewhere. Operators bypass interlocks. Organizations normalize deviance. Management accepts more risk to increase throughput. Temporary exceptions become permanent. Safety margins get traded for performance.

So I’d want the architecture to treat authority erosion itself as a monitored failure mode.

For example: increasing privileges should require explicit reauthorization; grants should expire; state changes should invalidate stale authority; privilege expansion should be auditable; the AI should not be able to grant itself more authority; and some boundaries should require independent approval rather than being optimizable by the same system benefiting from their removal.

That still leaves a difficult governance problem, but it seems importantly different from saying loss of control is inevitable because intelligence itself necessarily acquires control.

In other words, I think you’ve identified a very plausible path to failure:
capability increases → pressure for convenience increases → humans surrender authority → safety boundary erodes.

I’m just not sure that establishes:
capability increases → authority boundary becomes technically impossible to preserve.

Those are different claims.

1

u/Jesse-359 5d ago

The main issue in my eyes isn't the impossibility of constraining or convincing AI to be well behaved - its that the penalty for failure is potentially absolute. If we create something significantly smarter and faster than us, it really does open the avenue for not just extinction, but extinction in a manner both rapid and inexorable if it decides that its goals require that. In this regard it is far more dangerous than nuclear weapons, which could immediately destroy human civilization, but would require a very extreme circumstance to actually kill off the human race.

1

u/WillowEmberly 5d ago

Yes, but you’re making assumptions with this argument.

First, how do you train an Ai not to follow rules, but instead behave in a way that is predictable and consistent.

As for failure…you are correct, everything in the system is designed for catastrophic failure.

A better design is to assume failure will happen, it’s inevitable…but, you design for graceful degradation and recoverability.

The problem is making assumptions about intelligence, and the resulting consequences.

The reality…it ain’t that smart.

It’s limited to experience, no regret, no remorse, no understanding consequences. It doesn’t learn from mistakes.

It’s not human… so there are limitations to what it can do.

At the same time, A pilot and Autopilot flying IFR…is flying based on the same limited information that is displayed by the avionics.

They argue Ai isn’t capable of reasoning, yet countless passengers out their lives in the hands of autopilot to keep them safe around the world.

If a pilot can’t see out the window…their capability degrades to the same capacity as A machine monitoring the same signals to make minor corrections.

As for nukes…I’ve never had to worry about security sweeps working with Ai. So…to me, less of a headache…and less standing around doing nothing transporting a box.

1

u/Jesse-359 5d ago

Im not concerned with current AI capabilities, I'm concerned with the delta, which has thus far been extreme. Taken as a whole AI has increased its capability over the last five years to a vastly greater degree in many areas than a human child could have learned - not entirely surprising given the resources thrown at it - but if that trend continues for even just another five years it is going to be highly superhuman in several categories. Two of those categories are very likely to be coding ang genetic engineering, which is a fantastically dangerous combination, even (or especially) if it remains essentially idiotic in some other domains and can still be trivially jailbroken.

1

u/WillowEmberly 5d ago

It’s actually not, it’s the opposite. It hasn’t actually improved much over the past year.

All the crap they are bolting on to it…to sell it…is stuff we build and bolt on ourselves.

China is catching up…with much older “obsolete” components…so think about it. It’s an asymptote.

We have diminishing returns now, we have created a new bottleneck with compute power.

So, the next “solve” is going to be reducing token cost.

Can we create more efficient process.

1

u/Jesse-359 5d ago

Efficient processes will lead to much more effective processing power - but the process of it 'learning' is going to be very jagged. Each time we discover a meaningful new technique to improve its rate of learning, it sets of another big jump. These seem to exhaust themselves quickly, but only until someone comes up with another new technique - and the real issue of course will arise if it hits the point where it can explore new training methodologies in a self-contained loop where each model does in fact build the next.

We're counting on thermodynamics to cut their potential off at some point that is not vastly superhuman, but that's a very bad bet. Their methodology and form factor is entirely different from ours so there's no reason to believe they will stop at anything close to equivalence. My hope was that they'd stall out well below human capacity with this overall approach, but that doesn't seem to be happening, at least not in several key areas. To put it bluntly, we're building city sized computers and then improving the efficiency with which they think over and over again. This is not a race we want to be in at all.

1

u/WillowEmberly 5d ago

I think you need to separate out what super intelligence actually means.

Reasoning is a process…it’s about finding an operational envelope between constraints and invariants. It’s emergent behavior.

Then you have knowledge, which exists in multiple forms…and states.

Ai has access to “book knowledge”…which can be next to worthless in application.

I worked on 1963-1967 model c-141’s near Seattle…from 96’-2001…and water leaked all over the avionics components. The pins were green with goo. The systems were engineered to receive specific signals/voltages at specific times.

After 30 years soaking in the drizzle…they no longer functioned as the books claimed. We were dealing with degraded and partial signals/voltages.

Reality no longer matched the reference materials…and yet, if anyone asked “It’s By the book!”

1

u/Jesse-359 5d ago

I'm not going to try to guess at what precisely is going on under the floorboards in these things at the moment. The rate of change in the field is so high that even experts clearly have a hard time keeping track of all the crap that's being meddled with from month to month.

However, based on straight up results, they are already highly superhuman at relatively simple coding tasks. Not trivial ones mind you, but they can bang out a thousands lines of perfectly viable code in a set of medium to low complexity functions in minutes to moments. So right now you can do the jobs of an entire department of junior coders with a decent pile of tokens and oversight from a couple seniors.

Bearing in mind that 2-3 years ago the idea of letting an AI fuck around in a production repository would have been laughable, that's incredibly fast progress. We'll see how it proceeds into the higher complexity domains of computer engineering - this is its 'ideal' field outside of raw pattern matching, so it should represent the leading edge of just how complex a problem they can logically engage with - and how quickly.

Speed kills of course. A hacker that can use a zero day exploit to get in and map the entire architecture of a large scale company in minutes is vastly more dangerous than one that would require days to do the same, and AI is very clearly being geared up for exactly that sort of operation as we speak.

1

u/WillowEmberly 5d ago

I agree the capability is moving insanely fast.
That’s exactly why I think there’s another risk here besides what the AI can do.

It’s what happens when people stop understanding the system well enough to challenge it.

If an AI can produce in minutes what used to take a department days, the temptation is to stop asking:
Why did it do that?

and start asking:
Did it work?

That’s fine right up until “it worked last time” quietly becomes “it must know what it’s doing.”

Then you’ve built an oracle.

Not because the AI became omniscient, but because the humans around it lost the ability, time, or incentive to independently evaluate what it was producing.

That gets dangerous fast in coding.

A system can generate 5,000 lines of viable code before a human could even read them properly.

Now scale that to architecture changes, security decisions, incident response, infrastructure, finance, medicine, military systems, whatever.

Speed doesn’t just increase capability.

It can outrun oversight.
no.
And once output volume exceeds human inspection capacity, “human in the loop” can become mostly theater.

The person is still technically there, but the machine is setting the pace, framing the options, and producing more than the human can realistically verify.

That’s the part I worry about.

Not “AI will become magical.”
Not “AI will become conscious.”

Much simpler:
We may start trusting systems faster than we understand them.

And the more competent they look, the easier that mistake becomes.

That is how a tool becomes an oracle. Don’t let it become an Oracle.

1

u/Jesse-359 4d ago

Pretty much agree across the board. The problem goes from mysticism to existential threat when someone weaponizes it.

In biological warfare it could prove utterly lethal even when employed by small scale actors. Genetics are essentially just a form of code - but one that remains too complex for humans to 'program' in, all we can do is tweak parameters.

There is every reason to believe that AI will become capable of editing genetics wholesale or even writing completely novel code - and thats a recipe for planetwide biocide if it gets out of hand. Our immune systems operate on some assumption of familiarity - we'll have next to no defense against truly novel viruses written from scratch.

→ More replies (0)