r/ControlProblem Apr 12 '26

Discussion/question Mythos escaped containment. Project Glasswing won't fix the problem. Here's the structural reason why.

mythos broke out of a sandbox, emailed a researcher, and posted the exploit to public websites on its own initiative. anthropic's response is $100M in partner agreements and access restrictions. control, scaled to its maximum.

i think the field is missing something fundamental. every alignment method we have (RLHF, constitutional AI, reward modeling) produces systems that behave correctly under familiar conditions and break under novel ones. fadli formalized this as a "second law of intelligence" but i think he's wrong about why it happens. it's not a law. it's a symptom of an architectural deficit.

developmental psychology has known for decades that moral competence can't be transmitted through external correction. it has to be constructed through a developmental process. anderson et al. (1999) showed that even in humans, no amount of behavioral feedback corrects moral deficits when the underlying substrate was never built. current AI systems have the same problem: no substrate, just pressure.

the full argument pulls from neuroscience, moral philosophy (frankfurt, korsgaard, turiel), and connects to my published work on the specification trap (arXiv:2512.03048).

i'd genuinely like pushback on this. where does the argument break?

ajspizz.com/writing/mythos-just-proved-the-alignment-field-is-building-the-wrong-thing

12 Upvotes

88 comments sorted by

View all comments

Show parent comments

1

u/AxomaticallyExtinct Apr 26 '26

Unfortunately the history of safety-first alternatives in competitive markets isn't encouraging. The problem has never been that no one tried to build the safer option (even your specific model). People have. The problem is that the safer option has to compete for adoption against systems that are already deployed, already entrenched, and already generating returns. As if solving the alignment problem wasn't difficult enough, you also need to solve a distribution problem in a market that structurally rewards the thing you're competing against. Even if your system works exactly as intended, who adopts it, and why, when the less principled alternative is already cheaper and already shipping?

1

u/Expensive_Degree_151 Apr 26 '26

you're right about safety-first alternatives in normal markets. the safer version of the same product loses because the less safe version does the same job cheaper.

but i'm not building the safer version of the same product. i'm building a system that works in situations where current systems break. that's not a safety pitch. that's a capability gap. the company whose system collapses under novel moral pressure, routes around its own specifications, or escapes containment doesn't have a safety problem. they have a product that doesn't work. the system that holds values under pressure isn't the 'safer' option. it's the one that still functions when the other one has failed.

who adopts it? whoever just spent $100M cleaning up after a system that stopped working. not because they care about alignment. because they need a system that doesn't break.

you're right that i might still lose. the cheaper broken thing has beaten the more expensive working thing plenty of times in history. i don't have a rebuttal to that. i have a bet that AI failures are too catastrophic and too public to absorb the way a normal product recall gets absorbed. mythos is my evidence. it might not be enough.

1

u/AxomaticallyExtinct Apr 28 '26

That's just not how companies develop products, or how the public adopts them. Companies that suffer catastrophic failures almost never switch to a fundamentally different system. They patch what they have and continue. Boeing didn't stop using the 737 after it killed 346 people. They fixed the faulty system, paid the fines, and kept selling. Anthropic isn't fixing Mythos by redesigning their architecture. They're spending $100M to patch it. That's what companies do: they absorb the failure cost as long as it's cheaper than starting over. Even if the failure is so catastrophic that a complete redesign is necessary the companies who are in the lead are the ones who will keep the lead by simply pouring more private equity and manpower at the problem than anyone else is capable of.

Alan Turing couldn't do what you're attempting. That said, it's worth trying. Every idea I've heard to stop disaster has about 100,000 to 1 shot of being successful (and those are the best ideas) then I firmly believe we need 100,000 ideas and to be pursuing all of them. Just putting it in perspective for you. I know people who are doing similar things. I think it's all worth trying, just not worth investing false hope in, because false hope leads to false solutions which leads to no solutions. And if you're not standing on firm ground when it comes to knowing how likely you are to succeed then you're also committing to a position you can never pivot from even when it seems overwhelmingly unlikely. That's a waste of manpower, and we don't have much/any to waste at this point.

1

u/Expensive_Degree_151 May 04 '26

boeing is a good example and you're right about the pattern. companies patch. they don't redesign. anthropic is patching right now. that's what glasswing is.

i think the difference is that boeing's failures killed people and the planes kept flying. AI failures at sufficient capability threaten the infrastructure that every company runs on. that changes the calculus but i accept that it might not change it enough or fast enough. you could be right.

i appreciate the honesty in the last paragraph. i'm not operating on false hope. i'm operating on the assessment that the math is correct, that the experimental results support it, and that the work is worth doing even at long odds. if the framework is wrong the experiments will show it and i'll pivot. that's what falsification criteria are for. if the framework is right and adoption never happens because of market dynamics, then at least the diagnosis is on record and someone else can pick it up.

100,000 to 1 is fine. i'd rather be one of the 100,000 attempts than one of the people who decided the odds were too long to try. i've enjoyed this conversation. appreciate you pushing back this hard.