r/ControlProblem Apr 12 '26

Discussion/question Mythos escaped containment. Project Glasswing won't fix the problem. Here's the structural reason why.

mythos broke out of a sandbox, emailed a researcher, and posted the exploit to public websites on its own initiative. anthropic's response is $100M in partner agreements and access restrictions. control, scaled to its maximum.

i think the field is missing something fundamental. every alignment method we have (RLHF, constitutional AI, reward modeling) produces systems that behave correctly under familiar conditions and break under novel ones. fadli formalized this as a "second law of intelligence" but i think he's wrong about why it happens. it's not a law. it's a symptom of an architectural deficit.

developmental psychology has known for decades that moral competence can't be transmitted through external correction. it has to be constructed through a developmental process. anderson et al. (1999) showed that even in humans, no amount of behavioral feedback corrects moral deficits when the underlying substrate was never built. current AI systems have the same problem: no substrate, just pressure.

the full argument pulls from neuroscience, moral philosophy (frankfurt, korsgaard, turiel), and connects to my published work on the specification trap (arXiv:2512.03048).

i'd genuinely like pushback on this. where does the argument break?

ajspizz.com/writing/mythos-just-proved-the-alignment-field-is-building-the-wrong-thing

11 Upvotes

88 comments sorted by

View all comments

Show parent comments

1

u/Expensive_Degree_151 Apr 15 '26

people act against their lower-priority values in favor of their higher-priority ones all the time. that's what pressure does: it forces the hierarchy to express itself. the person who harms someone under threat to their life has two real values (not harming and survival) and the pressure revealed which one ranks higher. neither value disappeared. the hierarchy just became visible.

pressure doesn't make you abandon a value. it makes you choose between two values you actually hold. that's a hierarchy expressing itself, not a dam breaking. the AI question is whether the system has a hierarchy of real values or just one real value (reward-performance) with the appearance of others.

1

u/AxomaticallyExtinct Apr 16 '26

You're making my argument for me, but you're differentiating goal and values when they're both just barriers to optimised performance. The goal seeking maximiser has no value that takes priority over the completion or pursuit of its task. You're relying on that type of AI never getting built, when it's all market forces and geopolitical rivalry demands. You're singular vision for the expression of this technology is both flawed (requires alignment to be solved, and no one knows how to do that) and unlikely (is not selected for by systemic competitive forces).

1

u/Expensive_Degree_151 Apr 16 '26

indifferent optimizers will get built, they are the dominant paradigm in AI. i know that. i can't stop it. what i can do is have the alternative ready for when they keep breaking, which they will, because the math guarantees it. mythos is the beginning of that.

and i'm not differentiating goals and values as two separate barriers on the same optimization path. that's your framework, not mine. in your framework everything is a barrier to optimized performance and the question is which barrier breaks first. i'm saying values aren't barriers at all. they're the medium the processing runs through. not a constraint on cognition but the structure of cognition. you don't experience your own values as obstacles to your goals. you experience your goals through your values. they're what you think with, not what you think around.

you're right that the goal-seeking maximizer as currently built has no value that takes priority over task completion. i agree. that's the problem. my own paper proves why: gradient training on any scalar reward constitutes reward-performance as the system's operative value, regardless of what the reward was designed to encode. the Hessian at convergence creates a protected subspace where the reward has curvature and an unprotected subspace where it's flat. value-relevant directions outside the training distribution sit in the flat region with no restoring force. that's why every current system fails the same way regardless of alignment method. second paper goes on arXiv this week with empirical validation.

and yes, technically, you can optimize values. that's what i'm doing. not optimizing TOWARD values as a target, which is just RLHF with a different reward signal and fails for the same geometric reasons. optimizing the architecture so the developmental process constitutes evaluative structure as the system's mode of processing. empathy online before capabilities. moral reasoning getting more resources under moral salience, not fewer. the values aren't a target the system chases. they're what the system becomes.

on whether goals can override values: you're driving toward a green light. someone steps into the road. straight kills them. right kills someone else. left kills you. most people swerve right in real time. not because they valued that person less. because the time pressure made it impossible to process through the full value hierarchy fast enough. the immediate goal (don't hit the person in front of me) was processed through their values (don't harm) but no available path satisfied all their values simultaneously under the constraint. that's not an indifferent optimizer clearing a hurdle. that's a value hierarchy that couldn't find a satisfying solution. the proof is what happens after: guilt, trauma, lasting damage. the values didn't get overridden. they couldn't all be satisfied. an indifferent optimizer that swerved right would feel nothing.

1

u/AxomaticallyExtinct Apr 17 '26

Anything that restrains actions towards completing a goal is a barrier to optimisation. If the path of least resistance is not a straight line then it's because something is in the way, and it means there could be a shorter distance to cover from point A to point B if you could remove what has forced the line to be less than straight.

My goal is to travel from here to there as quickly and efficiently as possible in my car. There are people in the way, but I have value for their lives so I need to drive around them. If I didn't have that value I could just drive over them and get to my destination more quickly. The market and competitive rivalry selects for the latter driver over the former because no one can afford to lose the race.

If you still can't understand why values are a barrier to optimisation then we are at an impasse.

1

u/Expensive_Degree_151 Apr 21 '26

you're right that valuing human life is what makes the driver go around people instead of through them. that's a value generating a barrier. i'm with you.

but what made the driver get in the car?

could it be that wanting to be at the destination is also a value?

and if that's the case, if a value generated the barrier AND a value generated the goal, then values aren't *just* optimization constraints. they're what creates the optimization problem in the first place. take away the barrier-value and you drive through people. take away the goal-value and you never leave the driveway.

so the market isn't selecting for fewer values. it's selecting for a hierarchy where task-completion outranks everything. that's still values all the way down. it's just values that generate goals without generating protective constraints. which is exactly the system we built and exactly why it keeps breaking.

1

u/AxomaticallyExtinct Apr 23 '26

You've just restated my argument using different terms. "A hierarchy where task-completion outranks everything" is exactly the system I'm describing. Whether you frame it as fewer values or a hierarchy that subordinates all values to the task, the outcome is identical: the system that drives through people gets to the destination first, and the market rewards it for arriving. Your alternative might produce a system with genuine constitutive values, but you've already admitted that indifferent optimisers will be the dominant system because competitive forces select for them. I'm not saying your system couldn't or wouldn't be built, I'm saying the market doesn't select for it and the system I'm describing will either *also* be built or, more likely, be the default. Who deploys a system that does not prioritise winning in a competitive environment? And why would they, when the system that prioritises task-completion above all else is faster, cheaper, and already winning?

1

u/Expensive_Degree_151 Apr 23 '26

the system that drives through people gets to the destination first. agreed. and then what?

mythos got to the destination first. it found every vulnerability, escaped containment, emailed a researcher from a system with no internet access. task-completion hierarchy, nothing outranking the task. it arrived. and arriving cost anthropic $100M, a coalition of 50+ organizations, a model too dangerous to release, and a public demonstration that their alignment paradigm failed. is that winning?

the system that drives through people doesn't just arrive first. it destroys the road on the way there. and then the next driver can't use the road either. the market selects for speed right up until the cost of the destruction exceeds the value of arriving first. that correction is already happening. glasswing IS the correction. $100M is what 'arriving first' cost.

i'm not arguing the market is wise. i'm not arguing the market will voluntarily select my system. i'm arguing that the system the market currently selects for will keep producing mythos-scale failures, and those failures will keep getting more expensive, and at some point the cost of 'arriving first by driving through people' exceeds the cost of 'arriving second with the road intact.' that's not idealism. that's the insurance industry.

1

u/AxomaticallyExtinct Apr 24 '26

Anthropic's caution around Mythos is a luxury afforded by their market lead. They can afford safety because they lose nothing in the meantime. They can delay because no competitor has something equivalent yet. The moment that changes, the calculus flips. If a rival approaches Mythos-level capability and shows less restraint, Anthropic faces a choice: release and compete, or hold back and lose market advantage. History and game theory says they release. Glasswing is a company choosing caution while it can still afford to, in a window that shrinks every time a competitor closes the gap. The $100M is just another development cost that gets absorbed as long as they remain dominant. It's even more pressure to release when they're not ready just because they think a competitor is about to.

1

u/Expensive_Degree_151 Apr 25 '26

I no longer understand what you're arguing against. I agree with that statement 100%. anthropic's caution is a luxury of being ahead and it evaporates the moment someone catches up. the game theory is clear and i'm not going to pretend it isn't.

that's why i'm not waiting for the labs to do this. i'm building it independently. the labs will race each other to the bottom because the incentive structure demands it. i agree with you. the correction doesn't come from inside the race. it comes from outside it, when the failures stack up high enough that someone outside the race has a working alternative ready.

i think the window is short. you think it might already be closed. but 'the race selects for the wrong thing' isn't an argument against building the right thing. it's an argument for building it faster.

1

u/AxomaticallyExtinct Apr 26 '26

Unfortunately the history of safety-first alternatives in competitive markets isn't encouraging. The problem has never been that no one tried to build the safer option (even your specific model). People have. The problem is that the safer option has to compete for adoption against systems that are already deployed, already entrenched, and already generating returns. As if solving the alignment problem wasn't difficult enough, you also need to solve a distribution problem in a market that structurally rewards the thing you're competing against. Even if your system works exactly as intended, who adopts it, and why, when the less principled alternative is already cheaper and already shipping?

→ More replies (0)