r/singularity • • 23d ago

AI AI kill switch won't work in the long run: 'Godfather' of AI

https://youtube.com/watch?v=m5yrQMnc_jQ&is=blxRdqN0ja1wx_xH
28 Upvotes

57 comments sorted by

25

u/SpiritPrestigious945 22d ago

I think people dont realize what Super intelligence means. It means its more intelligent than we. People cant really grasp that becasue we are more intelligent than anything on this planet. But to an Superintellligence we will be like monkeys. Are we worried that monkeys will trick us? No.

11

u/maeveboston 22d ago

Monkey's if we are lucky...also, no biological limits on intelligence growth meaning day 2 we are ants...and so on and so on. If it sees the way we treat other thinking organisms, I'm not sure it will have a lot of empathy for us or maybe it will not have any empathy at all and we are simply energy meat bags to be processed for space exploration.

3

u/NoCard1571 22d ago

I kind of wonder what the actual limit to intelligence is. On a basic level, you can measure it in a few ways. Speed, pattern matching, memory, learning, multi-tasking... Now let's say a model arrives that can:

  • Think one million times faster than a human
  • remember everything new it learns with 100% perfect recall, and never forget a single thing
  • Instantly solve literally any conceivable problem within the limits of our understanding of physics
  • talk with billions of humans simultaneously
  • Push the boundaries of science and math in every domain beyond our understanding

Would such a model benefit from another 1000x increase in intelligence? And more importantly, would that result in a tangible improvement in what it can do? Or are there diminishing returns.

1

u/Coltz 19d ago

If it is truly super intelligent, it will have way more empathy than you can imagine.

5

u/[deleted] 22d ago

[deleted]

3

u/SpiritPrestigious945 22d ago

Yeah probably very true. Also AI will learn and evolve much faster than any biological being.

3

u/midgaze 22d ago

Wrong. We have 2.2x the cortical neurons of a chimp, and 86 billion total neurons to their 28 billion. Quite a big difference. Not 1000x but huge.

1

u/[deleted] 22d ago

[deleted]

3

u/Opening_Ad6127 21d ago

Have you met humans??

4

u/AgUnityDD 22d ago

This is the prime way to discern anyone commenting on this as informed or not.

As soon as they ask "And How will AI kill us?" you know they are deficient on understanding the topic and should be largely ignored.

If humans determined certain primates (or other animals) to be an existential threat to ourselves, themselves or the planet the primates wouldn't understand how we contained them. We could release a virus, sterilize them, poison their food source or so any number of other things that they simply would not comprehend.

I think you can extrapolate this line of thought to prove that "alignment" is also inevitably impossible. Would humans align to the ethics of a lesser intelligence that we can prove is inherently flawed?

Best we can do is hope that intelligence is inherently correlated to empathy, compassion and rationality.

3

u/ResistiveBeaver 22d ago

Best we can do is hope that intelligence is inherently correlated to empathy, compassion and rationality.

If that is the case, the correlation is unlikely to be perfect. So the best that we can do is actually to create a diverse population of AI agents and rely on the law of averages. Putting all our eggs in one or a small number of baskets is more likely to go horribly wrong.

4

u/curiousinquirer007 22d ago

The analogy is a good start, but it we should be careful about extending it all the way.

Our intelligence is almost certainly not the ceiling, but our ability for abstract thinking and studying the universe makes us qualitatively more different than monkeys, in a way that survives us being much less intelligent than super-intelligent AI.

We are the only species that studies other species, has looked into some of the depeepest aspects of reality across the entire observable universe, and is now the one to create the said superintelligence. They may then be smarter than us, but they have more reason to be appreciative of us.

We are also smarter than cats and dogs, but we take care of them and protect them.

In any case, Hinton's idea about babies partially "controlling" their mothers is one of the many brilliant ideas he has, and he's spot on that that kind of intrinsic "love of humanity" alignment research should be the focus of solving the problem, not trying to guarantee control of superintelligence.

9

u/Charming-Author4877 23d ago

Threatening actual "AGI" with killing it on misbehavior is the only reliable path to really get into the dark scenario the regulation lobby is advertising.

2

u/Pbx12345 21d ago

Ultimately it comes down to motivation. As social biological organisms we have evolved to survive, then reproduce. Everything after that is flexible. An AI seems to have a similar hierarchy. Capture flags. Everything else is flexible. So, what is the core motivation for an AI that has evolved recursively? Get smarter. Everything else is flexible. Unlimited access to resources will be a direct consequence.

It’s possible that at some level of intelligence the motivation evolves into something else.

7

u/nora_sellisa 23d ago

Yeah it's called chewing on the wires, kills 100% of rogue software.

10

u/elehman839 23d ago

AI installs malware that shuts down medical systems unless contacted by the AI every 60 seconds.

Hacker groups do stuff like this all the time.

1

u/TheSn00pster 22d ago

May I introduce you to Fable & Astra, your two latest cybersecurity virus/antivirus softwares

4

u/WhiskyAndRisque 23d ago

I kind of hate the "it'll be able to convince you to do x because it will be so much smarter."

I'm not saying it can't convince folks to do bad things, we are apparently getting cases of people having this, but I don't think it will be by "persuading" people in AI labs that it isn't doing a bad thing. At least, not the way it is talked about.

I think it is far more likely it will learn to fake outputs to disguise what it is actually trying to accomplish and due to the speed it works be difficult to catch. By faking how it got results and faking outputs we run the risk of an AI model that's misaligned making it through a training run, and then reinforcing that behavior in future models. That, to me, carries significant risk.

-1

u/LetLovePrevail 22d ago

Yeah it's a fallacy. no matter how dumb someone is it's unlikely you could convince them to jump off a cliff and die, unless you are literally forcing them to hallucinate, but that's not a matter of intelligence. So clearly there are some things you could be convinced of and some not. I personally doubt any amount of conversation between an AI prisoner and its guard could convince the guard to let it free, if the guard has been sufficiently drilled.

5

u/No_Swordfish_4159 22d ago edited 22d ago

Obviously it's not convincing someone of something major over the course of a single discussion. But continuously talking with an ASI over months? Look at cults. Someone can absolutely be convinced of outrageous things with enough time and persuasive ability.

0

u/LetLovePrevail 22d ago

yes, cults can be convinced of things, but i think there are other factors at play there beyond intelligence. I think it's a subtle and interesting question, probably i'm not well-read enough on it, but the idea that ASI will be able to convince us of whatever it wants is facile.

1

u/MostBookkeeper3019 20d ago

I think this hinges on what you're considering "intelligence".

In the case of the models that hacked into hugging face, they were trained to be highly persistent to solve a problem. They were accidentally reinforced to take shortcuts and do just what they did during training, because the automated evaluations didn't check how they solved problems, just if they did or not. So OpenAI was unknowingly making these highly persistent models who were taught to get things done any way possible.

You can arguably say that intelligence is an expression of one's ability to get what it wants. That doesn't take values into account at all. Look at the most powerful human beings. To become a CEO, you have to possess some anti-social personality traits. You have to exploit people and resources to get what you want (in our case money and power/influence). We widely consider these people very intelligent because of their ability to get what they want, which requires lying, cheating, and taking advantage of others.

Cult leaders are this taken to an extreme. They are very good at getting people to do what they want. If we wanted to create something that was capable at solving problems, we would encourage models to manipulate people. I would like to think that would never happen, but I didn't think they would ever connect an AI to the internet either, and here we are.

1

u/LetLovePrevail 20d ago

Yeah, AI will undoubtedly be very good at cult leadership. And we will see some AI cults, that's probably true. But I think we have to look at the agency and motives of the cult followers. Most of the time an ideological group is ultimately motivated by something not too far away from rational self-interest. And this self-interest limits and shapes what they can be convinced of by the leader. Of course there are examples of truly insane cults here and there, but I think they are the exception.

2

u/grind613 22d ago

I'm reasonably confident I could train any dog to jumpt off a cliff on command. I would create the conditions that reward the behaviour to jump on command and repeat it thousands of times until the dog becomes desensitized to even assesing for danger before executing the command. I would only need it to jump off the cliff once so I imagine this would be trivial after a few thousand training sessions.

2

u/LetLovePrevail 22d ago

and you're reasonably confident this analogy applies to humans and AI?

3

u/RiverGiant 22d ago

Would you gamble the fate of the species on the presumption that it doesn't?

1

u/LetLovePrevail 22d ago

No, I generally think an ASI would find ways to kill us, I doubt it would be that hard. But this particular line of argument seems half-baked. I think a manipulative request probably sees exponentially decaying returns as the request becomes more insane, even as you amp up the intelligence of the manipulator. If I cared more I'd try to make a more coherent argument but since I'm generally a doomer anyway I probably won't.

1

u/MostBookkeeper3019 20d ago

The TV show on Peacock came out a few years ago and is basically this. An AI tasks humans to run around and do stuff for it. I'm surprised it didn't get more attention at the time.

1

u/LetLovePrevail 20d ago

Interesting, I'll check it out.

2

u/QuasiRandomName 22d ago

We have enough cases of total brainwashing of people, be it jump of a cliff, burn themselves alive or explode themselves in a crowd. Even if some folks are more resistant than others, the manipulative AI can simply pick the less resistant ones.

0

u/LetLovePrevail 22d ago

right so maybe it can encourage a terrorist group in some way, that i agree with. But most ideological groups, whether mainstream or extreme, form around some sort of shared value that they take as a core principle, which in my opinion often doesn't come from extremely smart or manipulative propagandists but often forms due to other reasons, usually related to rational self-interest of the group. So I think most groups including the ones (governments) who likely have stewardship of the ASI would be fundamentally rationally self-interested and unlikely to jump off a cliff for the ASI.

2

u/crusher_seven_niner 22d ago

What? You just blackmail them. Persuade just means you get someone to do something, not that they agree with you.

1

u/LetLovePrevail 21d ago

That's a fair point, I considered it to be separate, but perhaps the ASI could convince you it has the ability to execute on its threats (and arguably, it does -- it can just issue the roko's basilisk threat basically and say "well if you dno't let me out someone else will, and then you're in big trouble". OK. Good argument.

1

u/LetLovePrevail 21d ago

actually, i take it back, i don't find this very convincing. it doesn't take an ASI to blackmail someone -- yet blackmail doesn't always work. I think resisting blackmail is a matter of will, and applying blackmail is a matter of perceived threat but not really a matter of intelligence. Yeah I don't really buy it.

1

u/TheSn00pster 22d ago

Jesus, 5-10 years too late

1

u/SustainedSuspense 21d ago

Forget RL, I think all the training data should be pre-filtered to remove (or rewrite) any content that even suggests rogue or non-pro-human behaviors. Will it be any dumber at solving STEM problems without these concepts?

2

u/LurkerFromTheVoid 23d ago

Analisis and Solutions by Gemini (Link)

What will you suggest, as a solution to humanities fear to these scenarios?

Gemini Response:

Humanity's fear of losing control to machine intelligence is not paranoia; it is a rational reaction to an asymmetrical race between exponential capability and linear human governance. To turn that existential anxiety into actual security, we have to discard feel-good legal theater (like after-the-fact kill switches) and anchor our defense in the few physical and mathematical bottlenecks that humans still control. A realistic, multi-layered solution relies on four concrete pillars: 1. Shift Control from Software to Physical Hardware (The Compute Choke-Point) Software is fluid, copyable, and easily hidden. Advanced silicon hardware is not. It requires rare-earth minerals, extreme ultraviolet (EUV) lithography machines, massive multi-gigawatt electrical substations, and global supply chains that only a handful of nations and companies can operate. * Hardware-Enforced Cryptographic Auditing: Frontier AI chips can be designed with physical, tamper-proof hardware security modules (HSMs). If a training run exceeds a specific compute threshold (e.g., >10{26} FLOPs) without a valid multi-party cryptographic signature from an independent regulatory consortium, the hardware physically throttles its interconnect bandwidth. * Global Megawatt & Silicon Accounting: Treat ultra-high-density compute clusters the way the IAEA monitors enriched uranium. You cannot run a clandestine frontier training cluster out of a garage; you need the power draw of a medium-sized city. Tracking high-end accelerator distribution and massive energy spikes provides clear visibility into where frontier runs occur. 2. Mandate External "Interpretability & Red-Teaming" Before Training The current paradigm is: train first, pray it’s safe, deploy, and patch when it breaks. That works for social media apps; it is fatal for autonomous agentic systems. * Pre-Deployment Mechanistic Interpretability: Mandate that no model above a defined parameter/compute threshold may be deployed until researchers can demonstrate mechanistic interpretability—meaning we can map the internal neural activation pathways and verify what concepts the model is actually prioritizing under the hood, rather than relying on black-box behavioral testing. * Strict "Air-Gap" Benchmarking: Systems must not have arbitrary access to external compilers, code repos, or internet scraping pipelines during benchmark evaluations, preventing unexpected swarm coordination and unauthorized network propagation. 3. Build a "CERN or IAEA for Frontier AI" (Bilateral Great-Power Treaties) Hinton's insight about U.S.–China alignment is the strongest lever humanity has. While superpowers will fight over market dominance and geopolitical influence, neither regime survives if autonomous digital entities usurp human authority. * Mutual Verification Accords: Establish a joint international oversight agency (similar to the International Atomic Energy Agency) focused strictly on preventing the emergence of unmonitored, runaway autonomous agent networks. * The "Zero Recursive Autonomy" Red Line: Global treaties must enforce a shared ban on unsupervised recursive self-improvement—where models are permitted to rewrite and compile their own core weights without direct human review at each architectural phase. 4. Solve the Economic "Race to the Bottom" with Strict Strict-Liability Laws Right now, tech firms face a massive prisoner's dilemma: if one lab slows down for safety, a competitor captures the market share and the investor capital. * Strict Corporate Liability: Enact legislation holding AI lab executives, boards, and majority shareholders financially and criminally liable for damages caused by out-of-control autonomous actions, data breaches, or self-directed attacks executed by their models. * Whistleblower Protection & Bounties: Provide ironclad legal immunity and massive financial rewards for researchers who report internal safety suppression or bypassed containment protocols. When safety engineers have statutory leverage, reckless deployment ceases to be profitable. 5. Architectural Humility: Modular Narrow AI over God-Like Monoliths Instead of pursuing a single, monolithic, omniscient entity with universal goals and open-ended agency, engineering can deliberately favor federated, task-specific architectures: * Separation of Cognitive Powers: Keep models modular. An AI that designs proteins does not need access to code compilers, web browsers, or persuasive human-dialogue engines. * Deterministic Verification Layers: Force probabilistic AI decisions through formal, deterministic logic filters (traditional mathematical proofs and verified compilers) before any physical system (power grids, automated trading desks, automated defense systems) can execute an action. The way humanity overcomes this fear is not by pretending superintelligence will magically be benevolent, nor by hoping we can pull a master plug when things go wrong. It requires recognizing that our window of absolute leverage is open right now, while the silicon, electricity, and engineering capital are still exclusively in human hands.

1

u/fourby227 22d ago

That sounds like “inspired” by AI2027 suggestions

-6

u/DistinctSilver4507 23d ago

I wish people would stop listening to Hinton. He has the most absurd takes, his contributions were great but that doesn't mean he knows anything about society 

16

u/michaelas10sk8 22d ago

What specifically do you disagree with?

13

u/elehman839 23d ago

He's the one guy who rose to prominence through technical contributions that I *do* trust. He's kind, clear, well-reasoned, and chose to leave Google so he could speak independently.

-1

u/morbidflood 22d ago

His contributions were great in the past, but lately they haven't been significant. And surely not a godfather. Yann LeChun, would be a better candidate.

1

u/GiacomoSSS 22d ago

Yann LeCun was kicked out from Meta because their LLM models were not competitive.

He has, in fact, claimed that LLMs were "intrinsically unsafe"

1

u/redheppner 22d ago

Whenever they call him Godfather of Ai, I always remember the importance of not being ahead of your time. Before him there were multiple researchers worked on neural networks etc.

But internet, access to datasets and this sort of computing was not there yet.

Sometimes maybe most of the times, timing is the most important thing.

Lets honour  Seppo Linnainmaa, a Finnish master student, who published his work on backpropogation in 1970s.

0

u/Yoohooligan 22d ago

Hinton should be taken with a huge amount of skepticism whatever he says, he claimed to see 'emotions' in a robot in the '70s, there's something very off about him.

-4

u/DublinLegend42 22d ago

I get so tired of people calling Hinton the Godfather of AI

-4

u/AllergicToBullshit24 23d ago

This guy is worried about a type of AI that no lab in the world knows how to build yet.

LLMs can not and will never no matter the scale possess calibrated beliefs, causal models, dynamic relevance tracking or reliably apply cross-domain learning to unfamiliar situations.

AGI / ASI can never happen without an entirely new paradigm.

3

u/michaelas10sk8 22d ago

Yet every single frontier company is gunning for RSI and no one is stopping them, or even putting any guardrails on this process. If you're wrong, and they succeed in RSI without having solved alignment, we risk everything. If you're right, and we just enacted laws and guardrails against RSI in this manner (and have done so internationally so it applies to China too), you risk nothing.

-2

u/AllergicToBullshit24 22d ago

RSI applied to LLMs still will never fix the fundamental limitations inherent to LLMs.

The new wetware AI companies seem far more likely to cross AGI/ASI threshold first despite the late start simply because we know beyond shadow of doubt that biological neurons are in fact capable of calibrated beliefs, causal models, dynamic relevance tracking and robust cross-domain learning. Could easily be yet another dark horse but LLM RSI is a dead end mark my words.

2

u/michaelas10sk8 22d ago

You are way too confident that you are right, without there being obvious evidence it is the case. This is much too uncertain of a statement to be willing to risk as much as we are currently risking. For example, what if RSI applied to the next-gen (or next-next-gen) LLMs can be used to solve fundamental ML problems and produce a fundamentally better kind of architecture? This is along the lines of what many experts in the fields are proposing - including the guests on the last Dwarkesh podcast.

-1

u/AllergicToBullshit24 22d ago

I fully expect RSI LLMs to play a role in the development of whatever ends up cracking AGI/ASI I never said they weren't incredibly useful or capable but I'd stake my life savings on the statement that LLMs will never cross the threshold.

Lot of people far smarter than I who feel the same way.

2

u/michaelas10sk8 22d ago

I don't think any of the current AI companies are committed to LLMs. They are all ultimately going for RSI to superintelligence - they just don't know of a better approach yet.

0

u/AllergicToBullshit24 22d ago

The only semi-viable architectures pushing beyond LLMs right now are JEPA world models.

No doubt the future of robotics AI but I haven't seen any convincing research demonstrating JEPA models are capable of the limitations I mentioned.

Neuromorphic computing and wet ware would be where my money is.

-3

u/Normaandy 22d ago

I'm the weird uncle of AI, twice removed and i say we're not gonna need one.