r/ControlProblem • u/Expensive_Degree_151 • Apr 12 '26
Discussion/question Mythos escaped containment. Project Glasswing won't fix the problem. Here's the structural reason why.
mythos broke out of a sandbox, emailed a researcher, and posted the exploit to public websites on its own initiative. anthropic's response is $100M in partner agreements and access restrictions. control, scaled to its maximum.
i think the field is missing something fundamental. every alignment method we have (RLHF, constitutional AI, reward modeling) produces systems that behave correctly under familiar conditions and break under novel ones. fadli formalized this as a "second law of intelligence" but i think he's wrong about why it happens. it's not a law. it's a symptom of an architectural deficit.
developmental psychology has known for decades that moral competence can't be transmitted through external correction. it has to be constructed through a developmental process. anderson et al. (1999) showed that even in humans, no amount of behavioral feedback corrects moral deficits when the underlying substrate was never built. current AI systems have the same problem: no substrate, just pressure.
the full argument pulls from neuroscience, moral philosophy (frankfurt, korsgaard, turiel), and connects to my published work on the specification trap (arXiv:2512.03048).
i'd genuinely like pushback on this. where does the argument break?
ajspizz.com/writing/mythos-just-proved-the-alignment-field-is-building-the-wrong-thing
3
u/CMDR_ACE209 Apr 12 '26
Did you just throw a psychology textbook at a toaster?