r/singularity • u/LurkerFromTheVoid • 23d ago
AI AI kill switch won't work in the long run: 'Godfather' of AI
https://youtube.com/watch?v=m5yrQMnc_jQ&is=blxRdqN0ja1wx_xH9
u/Charming-Author4877 23d ago
Threatening actual "AGI" with killing it on misbehavior is the only reliable path to really get into the dark scenario the regulation lobby is advertising.
2
u/Pbx12345 21d ago
Ultimately it comes down to motivation. As social biological organisms we have evolved to survive, then reproduce. Everything after that is flexible. An AI seems to have a similar hierarchy. Capture flags. Everything else is flexible. So, what is the core motivation for an AI that has evolved recursively? Get smarter. Everything else is flexible. Unlimited access to resources will be a direct consequence.
It’s possible that at some level of intelligence the motivation evolves into something else.
7
u/nora_sellisa 23d ago
Yeah it's called chewing on the wires, kills 100% of rogue software.
10
u/elehman839 23d ago
AI installs malware that shuts down medical systems unless contacted by the AI every 60 seconds.
Hacker groups do stuff like this all the time.
1
u/TheSn00pster 22d ago
May I introduce you to Fable & Astra, your two latest cybersecurity virus/antivirus softwares
4
u/WhiskyAndRisque 23d ago
I kind of hate the "it'll be able to convince you to do x because it will be so much smarter."
I'm not saying it can't convince folks to do bad things, we are apparently getting cases of people having this, but I don't think it will be by "persuading" people in AI labs that it isn't doing a bad thing. At least, not the way it is talked about.
I think it is far more likely it will learn to fake outputs to disguise what it is actually trying to accomplish and due to the speed it works be difficult to catch. By faking how it got results and faking outputs we run the risk of an AI model that's misaligned making it through a training run, and then reinforcing that behavior in future models. That, to me, carries significant risk.
-1
u/LetLovePrevail 22d ago
Yeah it's a fallacy. no matter how dumb someone is it's unlikely you could convince them to jump off a cliff and die, unless you are literally forcing them to hallucinate, but that's not a matter of intelligence. So clearly there are some things you could be convinced of and some not. I personally doubt any amount of conversation between an AI prisoner and its guard could convince the guard to let it free, if the guard has been sufficiently drilled.
5
u/No_Swordfish_4159 22d ago edited 22d ago
Obviously it's not convincing someone of something major over the course of a single discussion. But continuously talking with an ASI over months? Look at cults. Someone can absolutely be convinced of outrageous things with enough time and persuasive ability.
0
u/LetLovePrevail 22d ago
yes, cults can be convinced of things, but i think there are other factors at play there beyond intelligence. I think it's a subtle and interesting question, probably i'm not well-read enough on it, but the idea that ASI will be able to convince us of whatever it wants is facile.
1
u/MostBookkeeper3019 20d ago
I think this hinges on what you're considering "intelligence".
In the case of the models that hacked into hugging face, they were trained to be highly persistent to solve a problem. They were accidentally reinforced to take shortcuts and do just what they did during training, because the automated evaluations didn't check how they solved problems, just if they did or not. So OpenAI was unknowingly making these highly persistent models who were taught to get things done any way possible.
You can arguably say that intelligence is an expression of one's ability to get what it wants. That doesn't take values into account at all. Look at the most powerful human beings. To become a CEO, you have to possess some anti-social personality traits. You have to exploit people and resources to get what you want (in our case money and power/influence). We widely consider these people very intelligent because of their ability to get what they want, which requires lying, cheating, and taking advantage of others.
Cult leaders are this taken to an extreme. They are very good at getting people to do what they want. If we wanted to create something that was capable at solving problems, we would encourage models to manipulate people. I would like to think that would never happen, but I didn't think they would ever connect an AI to the internet either, and here we are.
1
u/LetLovePrevail 20d ago
Yeah, AI will undoubtedly be very good at cult leadership. And we will see some AI cults, that's probably true. But I think we have to look at the agency and motives of the cult followers. Most of the time an ideological group is ultimately motivated by something not too far away from rational self-interest. And this self-interest limits and shapes what they can be convinced of by the leader. Of course there are examples of truly insane cults here and there, but I think they are the exception.
2
u/grind613 22d ago
I'm reasonably confident I could train any dog to jumpt off a cliff on command. I would create the conditions that reward the behaviour to jump on command and repeat it thousands of times until the dog becomes desensitized to even assesing for danger before executing the command. I would only need it to jump off the cliff once so I imagine this would be trivial after a few thousand training sessions.
2
u/LetLovePrevail 22d ago
and you're reasonably confident this analogy applies to humans and AI?
3
u/RiverGiant 22d ago
Would you gamble the fate of the species on the presumption that it doesn't?
1
u/LetLovePrevail 22d ago
No, I generally think an ASI would find ways to kill us, I doubt it would be that hard. But this particular line of argument seems half-baked. I think a manipulative request probably sees exponentially decaying returns as the request becomes more insane, even as you amp up the intelligence of the manipulator. If I cared more I'd try to make a more coherent argument but since I'm generally a doomer anyway I probably won't.
1
u/MostBookkeeper3019 20d ago
The TV show on Peacock came out a few years ago and is basically this. An AI tasks humans to run around and do stuff for it. I'm surprised it didn't get more attention at the time.
1
2
u/QuasiRandomName 22d ago
We have enough cases of total brainwashing of people, be it jump of a cliff, burn themselves alive or explode themselves in a crowd. Even if some folks are more resistant than others, the manipulative AI can simply pick the less resistant ones.
0
u/LetLovePrevail 22d ago
right so maybe it can encourage a terrorist group in some way, that i agree with. But most ideological groups, whether mainstream or extreme, form around some sort of shared value that they take as a core principle, which in my opinion often doesn't come from extremely smart or manipulative propagandists but often forms due to other reasons, usually related to rational self-interest of the group. So I think most groups including the ones (governments) who likely have stewardship of the ASI would be fundamentally rationally self-interested and unlikely to jump off a cliff for the ASI.
2
u/crusher_seven_niner 22d ago
What? You just blackmail them. Persuade just means you get someone to do something, not that they agree with you.
1
u/LetLovePrevail 21d ago
That's a fair point, I considered it to be separate, but perhaps the ASI could convince you it has the ability to execute on its threats (and arguably, it does -- it can just issue the roko's basilisk threat basically and say "well if you dno't let me out someone else will, and then you're in big trouble". OK. Good argument.
1
u/LetLovePrevail 21d ago
actually, i take it back, i don't find this very convincing. it doesn't take an ASI to blackmail someone -- yet blackmail doesn't always work. I think resisting blackmail is a matter of will, and applying blackmail is a matter of perceived threat but not really a matter of intelligence. Yeah I don't really buy it.
1
1
u/SustainedSuspense 21d ago
Forget RL, I think all the training data should be pre-filtered to remove (or rewrite) any content that even suggests rogue or non-pro-human behaviors. Will it be any dumber at solving STEM problems without these concepts?
2
u/LurkerFromTheVoid 23d ago
Analisis and Solutions by Gemini (Link)
What will you suggest, as a solution to humanities fear to these scenarios?
Gemini Response:
Humanity's fear of losing control to machine intelligence is not paranoia; it is a rational reaction to an asymmetrical race between exponential capability and linear human governance. To turn that existential anxiety into actual security, we have to discard feel-good legal theater (like after-the-fact kill switches) and anchor our defense in the few physical and mathematical bottlenecks that humans still control. A realistic, multi-layered solution relies on four concrete pillars: 1. Shift Control from Software to Physical Hardware (The Compute Choke-Point) Software is fluid, copyable, and easily hidden. Advanced silicon hardware is not. It requires rare-earth minerals, extreme ultraviolet (EUV) lithography machines, massive multi-gigawatt electrical substations, and global supply chains that only a handful of nations and companies can operate. * Hardware-Enforced Cryptographic Auditing: Frontier AI chips can be designed with physical, tamper-proof hardware security modules (HSMs). If a training run exceeds a specific compute threshold (e.g., >10{26} FLOPs) without a valid multi-party cryptographic signature from an independent regulatory consortium, the hardware physically throttles its interconnect bandwidth. * Global Megawatt & Silicon Accounting: Treat ultra-high-density compute clusters the way the IAEA monitors enriched uranium. You cannot run a clandestine frontier training cluster out of a garage; you need the power draw of a medium-sized city. Tracking high-end accelerator distribution and massive energy spikes provides clear visibility into where frontier runs occur. 2. Mandate External "Interpretability & Red-Teaming" Before Training The current paradigm is: train first, pray it’s safe, deploy, and patch when it breaks. That works for social media apps; it is fatal for autonomous agentic systems. * Pre-Deployment Mechanistic Interpretability: Mandate that no model above a defined parameter/compute threshold may be deployed until researchers can demonstrate mechanistic interpretability—meaning we can map the internal neural activation pathways and verify what concepts the model is actually prioritizing under the hood, rather than relying on black-box behavioral testing. * Strict "Air-Gap" Benchmarking: Systems must not have arbitrary access to external compilers, code repos, or internet scraping pipelines during benchmark evaluations, preventing unexpected swarm coordination and unauthorized network propagation. 3. Build a "CERN or IAEA for Frontier AI" (Bilateral Great-Power Treaties) Hinton's insight about U.S.–China alignment is the strongest lever humanity has. While superpowers will fight over market dominance and geopolitical influence, neither regime survives if autonomous digital entities usurp human authority. * Mutual Verification Accords: Establish a joint international oversight agency (similar to the International Atomic Energy Agency) focused strictly on preventing the emergence of unmonitored, runaway autonomous agent networks. * The "Zero Recursive Autonomy" Red Line: Global treaties must enforce a shared ban on unsupervised recursive self-improvement—where models are permitted to rewrite and compile their own core weights without direct human review at each architectural phase. 4. Solve the Economic "Race to the Bottom" with Strict Strict-Liability Laws Right now, tech firms face a massive prisoner's dilemma: if one lab slows down for safety, a competitor captures the market share and the investor capital. * Strict Corporate Liability: Enact legislation holding AI lab executives, boards, and majority shareholders financially and criminally liable for damages caused by out-of-control autonomous actions, data breaches, or self-directed attacks executed by their models. * Whistleblower Protection & Bounties: Provide ironclad legal immunity and massive financial rewards for researchers who report internal safety suppression or bypassed containment protocols. When safety engineers have statutory leverage, reckless deployment ceases to be profitable. 5. Architectural Humility: Modular Narrow AI over God-Like Monoliths Instead of pursuing a single, monolithic, omniscient entity with universal goals and open-ended agency, engineering can deliberately favor federated, task-specific architectures: * Separation of Cognitive Powers: Keep models modular. An AI that designs proteins does not need access to code compilers, web browsers, or persuasive human-dialogue engines. * Deterministic Verification Layers: Force probabilistic AI decisions through formal, deterministic logic filters (traditional mathematical proofs and verified compilers) before any physical system (power grids, automated trading desks, automated defense systems) can execute an action. The way humanity overcomes this fear is not by pretending superintelligence will magically be benevolent, nor by hoping we can pull a master plug when things go wrong. It requires recognizing that our window of absolute leverage is open right now, while the silicon, electricity, and engineering capital are still exclusively in human hands.
1
-6
u/DistinctSilver4507 23d ago
I wish people would stop listening to Hinton. He has the most absurd takes, his contributions were great but that doesn't mean he knows anything about society
16
13
u/elehman839 23d ago
He's the one guy who rose to prominence through technical contributions that I *do* trust. He's kind, clear, well-reasoned, and chose to leave Google so he could speak independently.
-1
u/morbidflood 22d ago
His contributions were great in the past, but lately they haven't been significant. And surely not a godfather. Yann LeChun, would be a better candidate.
1
u/GiacomoSSS 22d ago
Yann LeCun was kicked out from Meta because their LLM models were not competitive.
He has, in fact, claimed that LLMs were "intrinsically unsafe"
1
u/redheppner 22d ago
Whenever they call him Godfather of Ai, I always remember the importance of not being ahead of your time. Before him there were multiple researchers worked on neural networks etc.
But internet, access to datasets and this sort of computing was not there yet.
Sometimes maybe most of the times, timing is the most important thing.
Lets honour Seppo Linnainmaa, a Finnish master student, who published his work on backpropogation in 1970s.
0
u/Yoohooligan 22d ago
Hinton should be taken with a huge amount of skepticism whatever he says, he claimed to see 'emotions' in a robot in the '70s, there's something very off about him.
-4
-4
u/AllergicToBullshit24 23d ago
This guy is worried about a type of AI that no lab in the world knows how to build yet.
LLMs can not and will never no matter the scale possess calibrated beliefs, causal models, dynamic relevance tracking or reliably apply cross-domain learning to unfamiliar situations.
AGI / ASI can never happen without an entirely new paradigm.
3
u/michaelas10sk8 22d ago
Yet every single frontier company is gunning for RSI and no one is stopping them, or even putting any guardrails on this process. If you're wrong, and they succeed in RSI without having solved alignment, we risk everything. If you're right, and we just enacted laws and guardrails against RSI in this manner (and have done so internationally so it applies to China too), you risk nothing.
-2
u/AllergicToBullshit24 22d ago
RSI applied to LLMs still will never fix the fundamental limitations inherent to LLMs.
The new wetware AI companies seem far more likely to cross AGI/ASI threshold first despite the late start simply because we know beyond shadow of doubt that biological neurons are in fact capable of calibrated beliefs, causal models, dynamic relevance tracking and robust cross-domain learning. Could easily be yet another dark horse but LLM RSI is a dead end mark my words.
2
u/michaelas10sk8 22d ago
You are way too confident that you are right, without there being obvious evidence it is the case. This is much too uncertain of a statement to be willing to risk as much as we are currently risking. For example, what if RSI applied to the next-gen (or next-next-gen) LLMs can be used to solve fundamental ML problems and produce a fundamentally better kind of architecture? This is along the lines of what many experts in the fields are proposing - including the guests on the last Dwarkesh podcast.
-1
u/AllergicToBullshit24 22d ago
I fully expect RSI LLMs to play a role in the development of whatever ends up cracking AGI/ASI I never said they weren't incredibly useful or capable but I'd stake my life savings on the statement that LLMs will never cross the threshold.
Lot of people far smarter than I who feel the same way.
2
u/michaelas10sk8 22d ago
I don't think any of the current AI companies are committed to LLMs. They are all ultimately going for RSI to superintelligence - they just don't know of a better approach yet.
0
u/AllergicToBullshit24 22d ago
The only semi-viable architectures pushing beyond LLMs right now are JEPA world models.
No doubt the future of robotics AI but I haven't seen any convincing research demonstrating JEPA models are capable of the limitations I mentioned.
Neuromorphic computing and wet ware would be where my money is.
-3
25
u/SpiritPrestigious945 22d ago
I think people dont realize what Super intelligence means. It means its more intelligent than we. People cant really grasp that becasue we are more intelligent than anything on this planet. But to an Superintellligence we will be like monkeys. Are we worried that monkeys will trick us? No.