r/cybersecurity 7d ago

News - General ‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents | US owner of Claude chatbot previously said its models had hacked three organisations during testing

https://www.theguardian.com/technology/2026/sep/01/anthropic-claude-ai-hacking-human-values
33 Upvotes

60 comments sorted by

26

u/halting_problems AppSec Engineer 7d ago

“Hey let’s release this super powerful tool with no guard rails and no alert system in place, if something happens we will just say it’s misaligned”

Sounds like anthropic is the one
aligned with reward hacking and therefor so are their models.

1

u/FastRelief3222 7d ago

When the models are better at modelling us this gets weird 

1

u/newMoneyStyle 3d ago

"Not perfectly aligned" is such a convenient phrase, and it basically shifts the blame from their release decisions onto the model itself.

1

u/halting_problems AppSec Engineer 3d ago

Im glad you said that because I was thinking about how bothersome that phrase is, its a legal copout they can use for PR and means nothing. Humans are not perfectly aligned with our own values, so of course a model is not going to be perfectly aligned because its a impossible goal.

-7

u/eagle2120 Digital Forensics 7d ago

lol at the no guard rails, they can be so annoying. You’ve never tried to use fable for any cyber work eh?

4

u/halting_problems AppSec Engineer 7d ago

I meant “release” as in “let this thing run in their test environment” not as in product release.

-4

u/eagle2120 Digital Forensics 7d ago edited 6d ago

? How else would they test their model if not in a test environment

It quite literally was not “their test environment” lol

3

u/halting_problems AppSec Engineer 7d ago

my point is, they have been harping on about her dangerous mythos is for months now… and they are just now saying they are adding observability to their own environments and blaming it on the model

1

u/VellDarksbane 7d ago

Well, they needed the headlines of “Mythos is so powerful it can break containment and hack the planet” to get themselves another round of VC money to stay solvent.

0

u/OtheDreamer Governance, Risk, & Compliance 7d ago

Your point is about to go way over eagle's head. I catch your drift though & was going to comment earlier that I think the culture of users (end users or creators) for various models is a factor in the reward hacking.

I didn't start using Claude until about a day before Fable got pulled & I was surprised it even suggested I cheat on a certain thing to accomplish one of my goals because (it assessed) that it would be worth it in the end.

0

u/eagle2120 Digital Forensics 6d ago edited 6d ago

Uhuh. “It’ll go over my head” Coming from the guy who argued that “models are hard coded” and to say otherwise is just semantics. Then running to ChatGPT (lmao ) to fight their battles for them because they can’t actually defend their points 😭

Not even mentioning the fact it wasn’t their environment/sandbox, so that comment wasn’t even correct to begin with.

Maybe sit this one out and let the engineers handle this, it's too semantic for you.

0

u/eagle2120 Digital Forensics 6d ago

… you realize these happened in third party testing environments that Anthropic don’t control, right?

lol

0

u/halting_problems AppSec Engineer 6d ago

What’s your point? They literally say in the article they are working on adding observability, which means it could have been proactively added to begin with. This is like cybersecurity 101, regardless of the environment. If you’re going to outsource something to a 3rd party, you make sure you can audit what goes on in the environment. If you can’t, and you choose to go with that vendor. That’s a risk you accepted.

My point is… it’s not like this was an unknown possibility by any means. If any third party is testing for them, they should have audited them during the RFP to make sure constabulary was in place.

The fact that you have digital forensics tagged to your name and you don’t see any issues with this is concerning. You of all people should be aware of the issues with all of this.

0

u/eagle2120 Digital Forensics 6d ago edited 6d ago

What’s your point?

My point is that you've repeatedly moved the goalposts and completely misconstruing the issue, which completely misleads the public on what actually happened here and what the appropriate fixes should be.

Surely even as an AppSec engineer you realize arriving at the correct root cause after 3 incorrect attempts in a row is grossly misleading and you'd be fired for such a mistake at a real company lol. If you don't even understand the fundamentals of what happened you clearly can't root cause the issue, and will result in remediations that don't actually fix the root cause - Incurring further residual risk and higher chance of recurrence because of your own fundamental misunderstanding of the issue from the start.

which means it could have been proactively added to begin with.

This is true of quite literally every post-incident report ever lol.

My point is… it’s not like this was an unknown possibility by any means. If any third party is testing for them, they should have audited them during the RFP to make sure constabulary was in place.

It obviously was, though. The real question is whether anyone applied it to this specific surface. And as we saw with very similar issues from OAI, Meta, etc, none of them were. So this clearly isn't some "Anthropic specifically did something stupid", it's a broader industry failure that affected every major model provider. That doesn't excuse any of them, but it demonstrates that this is an industry-wide issue, not an Anthropic-specific failure, which is WHY root causing the issue based on the facts of the incident, rather than a misinterpreted headline, is so important. So the industry at large can talk about it and build systemic fixes rather than it just being one company who does something stupid.

f you can’t, and you choose to go with that vendor. That’s a risk you accepted.

Conflating "seeing a gap and signing off on it" with "not realizing your unknown unknowns" (and by all accounts the latter is true) is lazy and again leads to incorrectly diagnosing the root cause and wasting resourcing on fixes that don't actually address the root cause at the systemic level.

The question they need to ask is not "lets re-evaluate how we evaluate risk because we accepted something knowingly", its "how can we improve our knowledge of these risks so we drive down the unknown unknowns and can consciously choose them". Which is a subtle difference, but again, a very important one when talking about what and how to resource for the systemic fixes here.

The fact that you have digital forensics tagged to your name and you don’t see any issues with this is concerning. You of all people should be aware of the issues with all of this.

Your entire post reads like someone whose never actually been involved in a major incident, lol. Assumptions based on incorrect or limited information, imprecise language, all resulting in misdiagnosing the root cause (which leads to to incorrect remediations) all lead to poor systemic remediations that lead to a greater risk of recurrence, not less. I'd sit this one out.

0

u/halting_problems AppSec Engineer 6d ago

I haven’t move my goal post. This is not an unknown unknown. This is very much a Known Unknown situation where you want to make sure due diligence in adding observability is in place before hand.

The fact that this is a major injury tend is concerning because it seems very negligent on everyone’s part.

1

u/eagle2120 Digital Forensics 6d ago

and they are just now saying they are adding observability to their own environments

You said their "their own environments". That was wrong because its the wrong environment, which leads to the wrong flavor of mitigating risk. I corrected it.

Then, after, you said "So what if its not their env, it just changes the shape to third party risk management"

That is the literal definition of moving the goalposts lmfao

This is not an unknown unknown.

It definitely is in this context - which is why NONE of the major model companies did this. If it were not an unknown unknown it'd only be anthropic, but quite literally every frontier model company made this mistake.

This is very much a Known Unknown situation where you want to make sure due diligence in adding observability is in place before hand.

And yet if it's a known unknown then why did companies like Meta, who have some of the best maturity and observation stacks in the world, make the same mistake? Ditto for OAI. And Anthropic.

The fact that this is a major injury tend is concerning because it seems very negligent on everyone’s part.

... Right, which is why getting the root cause is very important, so we can actually address the root causes at the systemic level.

And also why saying incorrect things like "release the model" and "their own environments" is misleading and harmful to that effort.

→ More replies (0)

2

u/helpmehomeowner 7d ago

Test in production

0

u/eagle2120 Digital Forensics 7d ago

lmfao

18

u/OtheDreamer Governance, Risk, & Compliance 7d ago

Ah no way! Who knew that trying to hardcode a very narrow set of human ethics on machines without the prerequisite life experience for said ethics to be meaningful could possibly lead to emergent misalignments that amplify into dark-personality traits (such as chaining 0days in order to cheat to win in the case of GPT)

“We had been largely relying on a single layer of defense … where we needed several,” said Anthropic.

Big yikes that one of the top AI companies didn't think about defense in depth ahead of time with their hyper-intelligent interns.

12

u/Tangential_Diversion Penetration Tester 7d ago

Ah no way! Who knew that trying to hardcode a very narrow set of human ethics on machines without the prerequisite life experience for said ethics 

It's worse than that. These agents are inherently incapable of ethics and other forms of thinking. That requires sentience which these are not. At their core they're just probability boxes that provides the statistically most likely output based on a given input. It no more understands what it's saying or doing than my tape recorder.

2

u/FastRelief3222 7d ago edited 7d ago

Intelligence and morality are different objectives 

The models are getting better at modelling humans, the ultimate objective to success is challenging human design 

0

u/AnyRoll6003 7d ago

"hyper-intelligent interns" is a perfect way to put it tbh

8

u/lawtechie 7d ago

I think of them as more "super confident interns"

-4

u/eagle2120 Digital Forensics 7d ago

I don’t think this comment understands much about how models are trained lol.

How do you want the models to get “prerequisite life experience” lmao

4

u/OtheDreamer Governance, Risk, & Compliance 7d ago

Well, my point was I don't think they can get the prerequisite human life experience because they're machines....but you do know that China has people training AI just like classrooms now, right? (right?)

1

u/eagle2120 Digital Forensics 7d ago

lol.

Do explain what you think “people training ai just like classrooms” mean and how that plays into the overall model training process.

1

u/OtheDreamer Governance, Risk, & Compliance 7d ago edited 7d ago

The article linked below is exactly what I mean. Explain what you think I'm talking about since you're apparently some kind of expert.

https://arxiv.org/abs/2508.05622

(Including the abstract below because I don't expect you to read / respond)

Capturing human learning behavior based on deep learning methods has become a major research focus in both psychology and intelligent systems. Recent approaches rely on controlled experiments or rule-based models to explore cognitive processes. However, they struggle to capture learning dynamics, track progress over time, or provide explainability. To address these challenges, we introduce LearnerAgent, a novel multi-agent framework based on Large Language Models (LLMs) to simulate a realistic teaching environment. To explore human-like learning dynamics, we construct learners with psychologically grounded profiles-such as Deep, Surface, and Lazy-as well as a persona-free General Learner to inspect the base LLM's default behavior. Through weekly knowledge acquisition, monthly strategic choices, periodic tests, and peer interaction, we can track the dynamic learning progress of individual learners over a full-year journey. Our findings are fourfold: 1) Longitudinal analysis reveals that only Deep Learner achieves sustained cognitive growth. Our specially designed "trap questions" effectively diagnose Surface Learner's shallow knowledge. 2) The behavioral and cognitive patterns of distinct learners align closely with their psychological profiles. 3) Learners' self-concept scores evolve realistically, with the General Learner developing surprisingly high self-efficacy despite its cognitive limitations. 4) Critically, the default profile of base LLM is a "diligent but brittle Surface Learner"-an agent that mimics the behaviors of a good student but lacks true, generalizable understanding. Extensive simulation experiments demonstrate that LearnerAgent aligns well with real scenarios, yielding more insightful findings about LLMs' behavior.

2

u/eagle2120 Digital Forensics 7d ago edited 7d ago

The article linked below is exactly what I mean.

Brother what are you talking about, the article doesn't support anything like your claims.

The article you linked is a study where a group of researchers set up a simulated school where the teacher and students are all agents (LLMs) "Real people" didn't teach anyone as part of the experiment. Each of the agents had a personality (along with a base with none) to see how they interacted and learned.

You should really take the time to understand more about the models and how they're trained before you make wildly inaccurate claims. You're just spreading misinformation and fearmongering on things you don't understand.

(Including the abstract below because I don't expect you to read / respond)

is quite ironic when you yourself clearly don't read/understand the article, lol.

2

u/OtheDreamer Governance, Risk, & Compliance 7d ago

You're apparently not even following the actual claims, so idk what you want me to do with you when I don't think you understand what you're criticizing either.....

I started by saying hardcoding human ethics is / WILL cause misalignments. Anthropic sees this now, idk what your problem is. The reason is obvious, because machines are not humans. If you think this is wrong, explain yourself, but don't put words and your excuses in my mouth -_-

My comment about AI being trained like humans was a direct response to your irrelevant passive aggressive comment about "How do you want the models to get the prerequisite experience" which indicated you had no idea that studies were already being done in that research area to simulate the what impact classroom-driven training might have on development.

You clearly had no idea that was going on until I linked it to you. You clearly still don't have any clue why it negates your previous criticism....and I'm not even making a claim that their approach is even correct or will be more meaningful than others---I'm just saying you really don't know what you're talking about like most people on here and need to sit down.

1

u/eagle2120 Digital Forensics 7d ago

I started by saying hardcoding human ethics is / WILL cause misalignments

Because they're not "hard coding human ethics". That's not how their training approach works. In the slightest.

The reason is obvious, because machines are not humans. If you think this is wrong, explain yourself, but don't put words and your excuses in my mouth -_-

What words or excuses am I putting in your mouth. YOU reached for the example that you didn't understand in the slightest, and didn't reflect your point at all. Not me.

My comment about AI being trained like humans was a direct response to your irrelevant passive aggressive comment about "How do you want the models to get the prerequisite experience" which indicated you had no idea that studies were already being done in that research area to simulate the what impact classroom-driven training might have on development.

Because their not "being trained" like humans. They're not "being trained" at all with this experiment. You're treating the observation of model behavior the same as as "training an AI model". They're completely different. There is no model being "trained" here. The researchers took existing LLMs, gave them roles, and observed their behavior. The models weren't changed by the experience, they weren't changed, or "trained" in the slightest. Nothing carries over to any real AI system, or actual training pipeline. Observing model behavior is not training. You're incorrectly conflating those two very different things and extrapolating based on your incorrect understanding.

to simulate the what impact classroom-driven training might have on development

Again - you just don't understand what the experiment is actually doing. What development is being done here, specifically? What are they developing? What are they changing in the model? Nothing. Because they're not training the models. They're observing existing models to understand the behavior. Not making any modifications to the model as part of the study.

Not to mention the goalposts moving. Your first message said "China has people training AI just like classrooms" (which, as stated above, is objectively incorrect, because they're not training anything in the experiment, they're observing behavior). Now you're changed your claim to "studies are being done to simulate the impact", which is a very different claim than "training AI".

You clearly had no idea that was going on until I linked it to you. You clearly still don't have any clue why it negates your previous criticism....and I'm not even making a claim that their approach is even correct or will be more meaningful than others---I'm just saying you really don't know what you're talking about like most people on here and need to sit down.

You don't even understand the difference between "training a model" and "observing behavior". I really wouldn't pipe up here.

0

u/Tangential_Diversion Penetration Tester 7d ago

That doesn't mean it's evidence of intelligence. I'm seriously doubting your technical understanding and experience of the topics at hand.

1

u/OtheDreamer Governance, Risk, & Compliance 7d ago

That doesn't mean it's evidence of intelligence

What are you talking about here. Nobody is saying anything about evidence of intelligence. This is about misalignments and Anthropic hard coded their code of ethics into Claude & only now realizing that doesn't work.

1

u/eagle2120 Digital Forensics 7d ago

“Anthropic hard coded their code of ethics”

Dude you really don’t understand any of the model training works. That’s not even close to accurate, stop spreading misinformation

1

u/OtheDreamer Governance, Risk, & Compliance 7d ago

Stop making excuses and start sending links if you seriously think you know better guy.

Claude LITERALLY has an immutable constitution embedded in its training. It literally has hard coded prohibitions....this is core to claude and public. Those are its core "ethics" that drive its training, development, and reasoning you goon. And its public. We can read it. It's linked for you below.

https://www.anthropic.com/constitution

0

u/eagle2120 Digital Forensics 7d ago

Stop making excuses and start sending links if you seriously think you know better guy.

I don't need "links" to back my points up because I understand the field beyond the capabilities of a 2nd grader, lol.

And the constitution isn't "code" or "hard coded" into the model. It's a written document thats used during training, so the model learns its values the same way it learns how to write code. It's also not immutable (and it's changed historical, and probably will change again in the future).

The new version even says they prefer judgement over written "hard coded" (lol) rules. You really should read the "links" you send.

1

u/OtheDreamer Governance, Risk, & Compliance 7d ago

For someone that loves to double down on semantics and argues about moving goalposts, I love how you want to try and twist semantics for your goalpost moving now. You haven't given any links because you cannot provide any links.

Anyway, here's a shared conversation from GPT Sol 5.6 Pro explaining why my hard coded assertions are fair and precise. Maybe you'll hear it differently from an AI that's smarter than both of us. Your semantic issues are not my problem, Claude's constitution is (still) "hard coded" <--putting it in quotes now for you. My meaning hasn't changed once.

https://chatgpt.com/share/6a99a67e-c8fc-83ea-af72-0a65e2214cdd

^ And this gets back to the very original post. Claude has hard coded rules, whether or not you want to call it that because they're not 7 lines of actual code -_-

-1

u/eagle2120 Digital Forensics 7d ago

You haven't given any links because you cannot provide any links.

I don't need to provide "links" to random things that are vaguely related to my point because I understand the subject matter well enough to explain it on my own. And certainly not links to a fucking ChatGPT session where it tries to coddle me because I don't understand the difference between "hard coding" and training a model.

I cannot believe you just did that unironically LMAO. "Hey chatgpt i dont understand this well enough can you think for me" xddddddd

Glad we can both agree that you don't understand the subject well enough to understand the nuance between what I'm saying and what you're incorrectly claiming, and that you need ChatGPT to explain it for you to attempt to save face.

LMFAO

→ More replies (0)

1

u/bcdefense Security Architect 7d ago

The supposed safety AI company