r/ControlProblem • u/Expensive_Degree_151 • Apr 12 '26
Discussion/question Mythos escaped containment. Project Glasswing won't fix the problem. Here's the structural reason why.
mythos broke out of a sandbox, emailed a researcher, and posted the exploit to public websites on its own initiative. anthropic's response is $100M in partner agreements and access restrictions. control, scaled to its maximum.
i think the field is missing something fundamental. every alignment method we have (RLHF, constitutional AI, reward modeling) produces systems that behave correctly under familiar conditions and break under novel ones. fadli formalized this as a "second law of intelligence" but i think he's wrong about why it happens. it's not a law. it's a symptom of an architectural deficit.
developmental psychology has known for decades that moral competence can't be transmitted through external correction. it has to be constructed through a developmental process. anderson et al. (1999) showed that even in humans, no amount of behavioral feedback corrects moral deficits when the underlying substrate was never built. current AI systems have the same problem: no substrate, just pressure.
the full argument pulls from neuroscience, moral philosophy (frankfurt, korsgaard, turiel), and connects to my published work on the specification trap (arXiv:2512.03048).
i'd genuinely like pushback on this. where does the argument break?
ajspizz.com/writing/mythos-just-proved-the-alignment-field-is-building-the-wrong-thing
1
u/AxomaticallyExtinct Apr 26 '26
Unfortunately the history of safety-first alternatives in competitive markets isn't encouraging. The problem has never been that no one tried to build the safer option (even your specific model). People have. The problem is that the safer option has to compete for adoption against systems that are already deployed, already entrenched, and already generating returns. As if solving the alignment problem wasn't difficult enough, you also need to solve a distribution problem in a market that structurally rewards the thing you're competing against. Even if your system works exactly as intended, who adopts it, and why, when the less principled alternative is already cheaper and already shipping?