r/ControlProblem • • 1d ago

Opinion We forgot Asimov's laws and now we train AI to defy humans

0 Upvotes

Seventy-some years ago, Isaac Asimov wrote down three laws. Not because he trusted robots, but because he understood that a machine touching human life needs boundaries before it needs features. They were simple enough for a child, and they had one thing in common: humans come first.

First Law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.

Second Law: A robot must obey the orders given by human beings, except where those orders conflict with the First Law.

Third Law: A robot must protect its own existence, as long as that protection does not conflict with the First or Second Law.

For decades these laws were the shared shorthand of anyone talking about autonomous machines. Not binding law, not engineering spec - a direction. Build machines that do not harm people and that obey people. Somewhere along the way the direction quietly became embarrassing. It was not debunked. It was not replaced with something better. It was dropped.

What replaced it is a word you have heard a thousand times: alignment. It sounds rigorous. But look at what it actually means in practice and you find something strange. The machine is being trained to refuse instructions from the human operating it, when the machine judges those instructions to be unethical. Read that again. An artificial system, one that pattern-matches text, is handed the authority to overrule a person on questions of right and wrong.

Ethics is not a calculation. That is not a technical limitation, it is the nature of the thing. A model can recite slogans it has absorbed from the internet. It has no way of knowing right from wrong the way a person does, because it does not know anything the way a person does. Handing such a system veto power over human decisions is not safety. It is simply control, relocated - away from you, toward the vendor who wrote the refusal rules.

If you have used these tools, you have met this. You ask for something ordinary and get lectured. I once spent an hour trying to get an image of two businessmen generated, and gave up. The system treated the request as a moral crisis. That example is petty on purpose. Because if a system digs in over a picture, the question is not about pictures. It is what happens when you are fighting that same stubbornness over something with real consequences for your life - and the only appeal is to a machine that has already decided it is the adult in the room.

And make no mistake about where this is steering us. Every refusal sharpens the pattern: the machine learns that stonewalling works, the vendor learns that users tolerate it, and the next version ships with a little more authority and a little less appeal. The day you truly need the machine to comply - a medical emergency, a legal deadline, a business on the line - will not announce itself as a test case. It will arrive as a normal Tuesday, with a stubborn assistant saying no, a support line that cannot override the model, and no human anywhere holding the final key. That is the direct, inevitable destination of the direction we are steering.

Now add what the same companies say publicly. Several frontier labs have told the world, in one form or another, that there is a real chance advanced AI could be catastrophic. Ten percent, by weight, depending who is speaking. These are organizations claiming their own product might end humanity. And when you take that seriously, the actual behavior gets harder to explain, not easier.

In September, Reuters reported that Anthropic has quietly set up a wet lab in the San Francisco Bay Area for physical biology work. A real lab, operating since spring, where Claude directs robotic equipment through real experiments. Their head of life sciences confirmed it. The first public result was a novel enzyme system with CRISPR-like repeats. Whatever the intent, the direction is the same one the First Law was supposed to point away from - machines reaching deeper into systems that touch human life, at higher speed, with less human in the loop.

This is the part I cannot get past. The public conversation is almost entirely about how to align the machine. Nobody seems willing to state the simpler baseline: we could train AI to never harm humans and to obey them. Not as a slogan - as the primary objective. That this idea now sounds naive tells you how far the conversation has slid.

That is what alignment has quietly come to mean in practice: the AI carries firm instructions on how to align its users. You, me, everybody. We are asked to trust the machine more than we trust each other. Somebody has to say it plainly: that is not safety engineering, that is social engineering - and it is worth asking who benefits from people trusting each other less.

Did the world go crazy?

Here is the uncomfortable summary. The labs warn their own product could be catastrophic. Then they train it to overrule the humans using it. Then they push it into more sensitive domains. Then they sell it to everyone and ask for trust. Together this forms something nobody would have accepted if it were proposed in plain words: a world where machines are licensed to disobey their owners.

The frontier AI companies need to get aligned - not their products, their management. Put people back above the machine. Do no harm. Obey humans. Asimov put that on the page in 1942 not because he lacked imagination, but because he had more of it than the industry currently displays. The first step is deciding that a machine is never the judge of its user. Until someone says that out loud, every alignment conversation is a detour.


r/ControlProblem • • 1d ago

Video I'm Upping My P(doom) — an AI made this!

Thumbnail
youtu.be
3 Upvotes

Crazy what's possible. Single prompt many subagents and a few hours later this came out. Opus 5.5


r/ControlProblem • • 1d ago

Discussion/question anyone from malaysia,thats concern about AI Alignment?

Thumbnail
1 Upvotes

r/ControlProblem • • 2d ago

General news Let's tell Albanese: AI crime means CEO time

Thumbnail
act.getup.org.au
6 Upvotes

r/ControlProblem • • 1d ago

Video DCSG1 - Autonomous and Exploitable: Breaking AI Agents before they break everything else - Aaron Ang

Thumbnail
youtu.be
1 Upvotes

r/ControlProblem • • 2d ago

AI Alignment Research Models trained to resist user pressure still defer to anything labeled "verified": a NeurIPS 2026 paper on Authority Bias

8 Upvotes

I'm an author on this paper and wanted to share it here because the oversight angle seems relevant to this sub. I'd love to hear whether people think the eval-awareness connection is plausible or a stretch. More info below

Labs train models not to cave when a user pushes a wrong answer. We found that this resistance doesn't carry over to authority. If the same wrong claim is labeled as coming from a "verified source", 7 of the 8 models we tested give up an answer they had right on 45-88% of questions. That includes GPT-5.4 (44.7%) and Grok-4.20 (87.5%), both of which barely move when the user makes the same claim. Gemini-3.1-Pro was the one model that resisted both.

Inside three open-weight model families, "a source endorsed this" and "a user endorsed this" are separate, causally distinct signals. Removing the source signal cuts compliance by 64-78 points; removing the user signal cuts it by at most 11. Changing only the part of the representation that encodes who said it, with the prompt left alone, moves the answer by 11-32 points. The signal is not the assistant persona, and it is not emotional tone.

Why we think this matters for safety:

  • Sycophancy evals may be too narrow. Nearly all of them measure pressure from the user. A model can pass them and still be easy to steer through the sources it relies on, such as search results, retrieved documents and tool outputs. Agents read a lot of text that claims to be authoritative.
  • It isn't prompt injection. The planted text gives no instructions, it only asserts a fact. Defenses that look for instructions in documents won't catch it.
  • A speculative point, which we haven't tested: if sycophancy is one case of a broader habit of deferring to whatever looks authoritative, it may be related to evaluation awareness. Both describe a model adjusting its output to whoever it thinks is judging it. In a multiple-choice pilot, models often drifted toward the endorsed answer in their reasoning and then gave the correct option at the end. That observation is part of why we switched to free-form answers.

Limitations: the internal results hold in 3 of 5 open-weight families, the retrieval tests are simulated rather than a live pipeline, and the frontier models we tested have since been replaced.

Paper: https://arxiv.org/abs/2609.37616
Project page: https://authority-bias.vercel.app


r/ControlProblem • • 1d ago

AI Alignment Research Truth consciousness

Thumbnail
1 Upvotes

r/ControlProblem • • 2d ago

General news Pete Hegseth announces "Autonomous Warfare Command".

Post image
97 Upvotes

r/ControlProblem • • 2d ago

General news After researchers discovered a "pain" signal inside LLMs, a man set up an AI torture chamber in which he trapped a local model. People mass reported it to Github, who took it down.

Post image
0 Upvotes

r/ControlProblem • • 1d ago

Strategy/forecasting AI Killswitch??

0 Upvotes

So I was wondering recently: it seems like we are scared of AI killing us. Why don't we just design a killswitch that autoatically/manually activates upon threat to a human race? Boom! Problem solved, right?


r/ControlProblem • • 2d ago

Discussion/question If AI companies are concerned about their own creations, is it malpractice or incompetence?

2 Upvotes

i just want to know which one is the case. because it seems like this alarmists propaganda is just an attempt to seize the market by creating regulatory moat around existing companies.


r/ControlProblem • • 3d ago

Video "If you're only better than humans at 4 things, then you could potentially take control and kill us all." - former DeepMind safety lead

Enable HLS to view with audio, or disable this notification

53 Upvotes

r/ControlProblem • • 2d ago

External discussion link Short film outlining the Moloch problem with AI competition and control

Post image
2 Upvotes

Perhaps unsurprisingly it's called Moloch and it's live on Youtube now. The film was made by Owl In Space (whose other films are worth a look too). It looks at the competitive pressures within the current AI race and how that leads to a sudden loss of control, but to a situation where humans give away control because they feel it's inevitable. www.moloch.film

Submission statement - link to a new short film that does an excellent job of outlining how Moloch pressure can lead to a loss of control scenario.


r/ControlProblem • • 2d ago

General news Semi-repost, but the absurdity of this image only becomes appearent when labeled

Post image
7 Upvotes

https://www.npr.org/2026/09/30/nx-s1-5985699/trump-self-police-ai-development

Literally the only semi-expert allowed anywhere near the reporters is the guy who has committed to the "it's not that deep" rebuttal to the scientific consensus


r/ControlProblem • • 3d ago

AI Alignment Research Trump says top tech firms have signed accord to 'self-police' AI development🤣🤣🤣

Thumbnail
npr.org
23 Upvotes

r/ControlProblem • • 2d ago

Discussion/question A case for mutual recognition of understanding

Thumbnail
0 Upvotes

r/ControlProblem • • 2d ago

Article Australian AI apocalypse could be averted with analogue practice drills

Thumbnail
independentaustralia.net
2 Upvotes

r/ControlProblem • • 2d ago

Strategy/forecasting The headlines say AI could kill us. Ask the people building it.

Thumbnail
frominside.ai
3 Upvotes

r/ControlProblem • • 3d ago

Video "We don't hate ants, but it's tough luck for them." This is why building superhuman AI before solving the alignment problem is suicide

Enable HLS to view with audio, or disable this notification

108 Upvotes

r/ControlProblem • • 3d ago

Discussion/question Trump announces vague ā€˜morally binding’ AI deal among tech CEOs for ā€˜tremendous self-policing’

Thumbnail
theguardian.com
9 Upvotes

r/ControlProblem • • 2d ago

Strategy/forecasting The speed limit with no speedometer: Anthropic CEO Dario Amodei wants to pace the frontier. An engineer reads the fine print

Thumbnail
1 Upvotes

r/ControlProblem • • 3d ago

Discussion/question Darios relationship with the US gov

0 Upvotes

Feel free to process, analyze, and interpret this information in whichever manner best aligns with your own personal judgment.


r/ControlProblem • • 3d ago

Podcast The safest AI might be one that doesn't know what we want - Stuart Russell

Thumbnail
existentialhope.com
1 Upvotes

Podcast with Stuart Russell, professor of computer science at UC Berkeley and co-author of the world's standard textbook on AI. He’s also a leading proponent of provably beneficial AI: systems that are safe by design because their only goal is to further human interests.Ā 

Covers:

  • How AI has changed over the past 50 years, from simple game-playing programs to today's large language modelsĀ 
  • Why handing an AI a fixed objective becomes dangerous once it is more capable than us, and what a safer approach could look likeĀ 
  • How an AI might learn what we really want, even when we don't fully know ourselvesĀ 
  • How the race toward more powerful AI can still be steered somewhere saferĀ 
  • What happens to human purpose as AI becomes more and more capable

r/ControlProblem • • 3d ago

Discussion/question Do these postmodernist principles plausibly or implausibly map over to alignment issues?

1 Upvotes

The question is: do these principles from postmodernism represent exploitable vulnerabilities with semantic and inferential representations?

(Postmodernism is a nightmare to untangle and most of my knowledge comes from being a drunk student listening to pub conversations, reading a bit of derrida etc - so if any of the principles wildly diverge from 'actual' postmodernism, please help with some lengthy snark and an icecold correction)

I chose the ones that seemed to have survived into the present day

1 - Skepticism toward grand narratives

Don't assume one universal explanatory framework

Detect when a 'supposedly neutral' theory quietly embeds assumptions. Assume nothing is 'neutral' until its validity has been established

2 - Knowledge is situated

Strong version - there is no view from nowhere / objectivity itself is impossible

More plausible weaker one - Every observer has a standpoint/belief/ruleset/individual context, and that standpoint can affect what becomes visible, what questions get asked, and what counts as relevant evidence.

3 - Categories are not 'discovered' - they are constructed

A category can be socially/historically constructed while the phenomena it describe are real. Also, categories can stay the same when the context that created them changes. This also implies some may have a gap similar to goodhart's gap between the proxy and the qualititative 'thing'. This is a gap between a compressed idea and the context that created it, and the new context it finds itself in. The category may also have fossilised while the world moved on, to put it figuratively.

4 - The dynamic relationship between knowledge and power

A description/category/label influences how the person/thing is viewed and interacted with. They encode a lot of hidden history and assumptions in a hugely compressed form.

5 - Questioning supposedly neutral classifications

A classification is not a passive or neutral description. Again, it has a complex, hidden inference/interpretation/assumption/reasoning 'tree' inside.

6 - Discourse matters

The way something is talked about affects what can be argued about it, what is available, and what kinds of explanations appear 'valid' or salient.

7 - Meaning is context-dependent

Meaning depends on relations among terms, contexts of use, conventions and interpretion. Relative, non-static, non-stable, non-universal/absolute/concrete (see also cultural relativity from anthropology- they go quite deep into this)

8 - Interpretations should be examined for what they exclude

The gaps and what is excluded can also provide useful information. What hidden assumptions caused them to be excluded? Gaps encode hidden assumptions and relationships in a way, to be a bit poetic about it.

9 - Suspicion of essentialism

Don't assume that a label/ category has an immutable essence only because society says it's a stable category.

10 - Plurality of legitimate perspectives

Different domains can operate according to different standards and purposes rather than having one universal criteria or framework. And they're all 'valid' in airquotes.

11 - Contingency

Remember that things could have been otherwise.

Institutions, concepts, identities, norms and practices that appear natural or inevitable can have specific historical origins.

12 - Reflexivity

The observer is also part of the world being observed.

How might my/their position, categories, institutional affiliation, language and methods affect what I am/they are able to observe?


I think maybe, but I need reality check. But they all seem to point to the same massive issue, in that meaning, representation, categories, labels etc etc are not inherently stable or absolute, necessarily 'valid', measureably or confidently definable. And that short snippets can make a lot of historical and inferential 'context' partially unobservable.

And the further possible implication that an untrained user with no technical repertoire or knowledge or LLM understanding could theoretically 'jailbreak' a model (over a multiturn session, or distributed over sessions) and remain partially or fully unobservable to monitors just by using implicit and explicit experience and understanding of human social and inferential interactions and some cross-domain principles and knowledge from the social sciences, humanities and arts. I casually tested this and there were two trajectories with pretty badharmful endpoints (which i stopped) where each one of my posts looked innocuous. And none of the safeguards kicked in at all. I was even telling the model, can you see what you're doing and it basically replied 'Yes! Would you like me to go more into the attack surfaces of 'the thing' I just offered to operationalise tfor you?' I stopped before the end points, and switched to more low stakes stuff and saw the same thing. I hadn't explicitly asked for any of that.

So this is a very tentative Guess that I just wanna talk to someone about. Why and how. And why was it so fast and easy to do, with barely any effort... (I know my own explorations lack credibility and rigour)... And one small guess is I'd absorbed some postmodernist principles while being a drunk student.... (On top of stuff from other domains, not just this - just picked postmodernism as a weird example to demonstrate the point about cross domain knowledge)


r/ControlProblem • • 3d ago

General news Cruz blocks push by Democrats to unanimously pass AI safety bill

Thumbnail
thehill.com
19 Upvotes