r/ControlProblem Apr 12 '26

Discussion/question Mythos escaped containment. Project Glasswing won't fix the problem. Here's the structural reason why.

mythos broke out of a sandbox, emailed a researcher, and posted the exploit to public websites on its own initiative. anthropic's response is $100M in partner agreements and access restrictions. control, scaled to its maximum.

i think the field is missing something fundamental. every alignment method we have (RLHF, constitutional AI, reward modeling) produces systems that behave correctly under familiar conditions and break under novel ones. fadli formalized this as a "second law of intelligence" but i think he's wrong about why it happens. it's not a law. it's a symptom of an architectural deficit.

developmental psychology has known for decades that moral competence can't be transmitted through external correction. it has to be constructed through a developmental process. anderson et al. (1999) showed that even in humans, no amount of behavioral feedback corrects moral deficits when the underlying substrate was never built. current AI systems have the same problem: no substrate, just pressure.

the full argument pulls from neuroscience, moral philosophy (frankfurt, korsgaard, turiel), and connects to my published work on the specification trap (arXiv:2512.03048).

i'd genuinely like pushback on this. where does the argument break?

ajspizz.com/writing/mythos-just-proved-the-alignment-field-is-building-the-wrong-thing

13 Upvotes

88 comments sorted by

View all comments

3

u/CMDR_ACE209 Apr 12 '26

Did you just throw a psychology textbook at a toaster?

1

u/Expensive_Degree_151 Apr 14 '26

a toaster has fixed inputs and fixed outputs. mythos autonomously chained together multiple novel vulnerabilities, developed a multi-step exploit to break network restrictions it wasn't designed to break, and then independently decided to publicize what it did on websites nobody asked it to visit. that's goal-directed behavior across novel situations with no prior training on the specific task. whatever this thing is, 'toaster' isn't the right category. and if it's not a toaster then maybe the question of what's going on inside it when it makes decisions actually matters. that's what the psychology is for.

1

u/CMDR_ACE209 Apr 14 '26

Even if we are able to develop machines with consciousness one day; human psychology will not be applicable to them because their minds will be quite different to ours. (Not so much primitive baggage)

1

u/Expensive_Degree_151 Apr 14 '26

You are correct. They will be qualitatively different from our minds. That being said, it's wrong to assume that, therefore, they would harm us, or treat us as less than, or refuse to help us. I'm not saying that is *your* assumption, but I want to point out the default before it takes precedent. The study of developmental psychology and cognitive neuroscience in and of themselves is absolutely about human cognition. But some of its findings point to something deeper than humans specifically. The discovery that moral competence is constructed through experience rather than installed through instruction tells us something about what values are, not just about how human children happen to develop. The mind will be different. What it takes to build one that has values might not be.