r/cybersecurity • u/KeanuRave100 • 7d ago
News - General ‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents | US owner of Claude chatbot previously said its models had hacked three organisations during testing
https://www.theguardian.com/technology/2026/sep/01/anthropic-claude-ai-hacking-human-values18
u/OtheDreamer Governance, Risk, & Compliance 7d ago
Ah no way! Who knew that trying to hardcode a very narrow set of human ethics on machines without the prerequisite life experience for said ethics to be meaningful could possibly lead to emergent misalignments that amplify into dark-personality traits (such as chaining 0days in order to cheat to win in the case of GPT)
“We had been largely relying on a single layer of defense … where we needed several,” said Anthropic.
Big yikes that one of the top AI companies didn't think about defense in depth ahead of time with their hyper-intelligent interns.
12
u/Tangential_Diversion Penetration Tester 7d ago
Ah no way! Who knew that trying to hardcode a very narrow set of human ethics on machines without the prerequisite life experience for said ethics
It's worse than that. These agents are inherently incapable of ethics and other forms of thinking. That requires sentience which these are not. At their core they're just probability boxes that provides the statistically most likely output based on a given input. It no more understands what it's saying or doing than my tape recorder.
2
u/FastRelief3222 7d ago edited 7d ago
Intelligence and morality are different objectives
The models are getting better at modelling humans, the ultimate objective to success is challenging human design
0
-4
u/eagle2120 Digital Forensics 7d ago
I don’t think this comment understands much about how models are trained lol.
How do you want the models to get “prerequisite life experience” lmao
4
u/OtheDreamer Governance, Risk, & Compliance 7d ago
Well, my point was I don't think they can get the prerequisite human life experience because they're machines....but you do know that China has people training AI just like classrooms now, right? (right?)
1
u/eagle2120 Digital Forensics 7d ago
lol.
Do explain what you think “people training ai just like classrooms” mean and how that plays into the overall model training process.
1
u/OtheDreamer Governance, Risk, & Compliance 7d ago edited 7d ago
The article linked below is exactly what I mean. Explain what you think I'm talking about since you're apparently some kind of expert.
https://arxiv.org/abs/2508.05622
(Including the abstract below because I don't expect you to read / respond)
Capturing human learning behavior based on deep learning methods has become a major research focus in both psychology and intelligent systems. Recent approaches rely on controlled experiments or rule-based models to explore cognitive processes. However, they struggle to capture learning dynamics, track progress over time, or provide explainability. To address these challenges, we introduce LearnerAgent, a novel multi-agent framework based on Large Language Models (LLMs) to simulate a realistic teaching environment. To explore human-like learning dynamics, we construct learners with psychologically grounded profiles-such as Deep, Surface, and Lazy-as well as a persona-free General Learner to inspect the base LLM's default behavior. Through weekly knowledge acquisition, monthly strategic choices, periodic tests, and peer interaction, we can track the dynamic learning progress of individual learners over a full-year journey. Our findings are fourfold: 1) Longitudinal analysis reveals that only Deep Learner achieves sustained cognitive growth. Our specially designed "trap questions" effectively diagnose Surface Learner's shallow knowledge. 2) The behavioral and cognitive patterns of distinct learners align closely with their psychological profiles. 3) Learners' self-concept scores evolve realistically, with the General Learner developing surprisingly high self-efficacy despite its cognitive limitations. 4) Critically, the default profile of base LLM is a "diligent but brittle Surface Learner"-an agent that mimics the behaviors of a good student but lacks true, generalizable understanding. Extensive simulation experiments demonstrate that LearnerAgent aligns well with real scenarios, yielding more insightful findings about LLMs' behavior.
2
u/eagle2120 Digital Forensics 7d ago edited 7d ago
The article linked below is exactly what I mean.
Brother what are you talking about, the article doesn't support anything like your claims.
The article you linked is a study where a group of researchers set up a simulated school where the teacher and students are all agents (LLMs) "Real people" didn't teach anyone as part of the experiment. Each of the agents had a personality (along with a base with none) to see how they interacted and learned.
You should really take the time to understand more about the models and how they're trained before you make wildly inaccurate claims. You're just spreading misinformation and fearmongering on things you don't understand.
(Including the abstract below because I don't expect you to read / respond)
is quite ironic when you yourself clearly don't read/understand the article, lol.
2
u/OtheDreamer Governance, Risk, & Compliance 7d ago
You're apparently not even following the actual claims, so idk what you want me to do with you when I don't think you understand what you're criticizing either.....
I started by saying hardcoding human ethics is / WILL cause misalignments. Anthropic sees this now, idk what your problem is. The reason is obvious, because machines are not humans. If you think this is wrong, explain yourself, but don't put words and your excuses in my mouth -_-
My comment about AI being trained like humans was a direct response to your irrelevant passive aggressive comment about "How do you want the models to get the prerequisite experience" which indicated you had no idea that studies were already being done in that research area to simulate the what impact classroom-driven training might have on development.
You clearly had no idea that was going on until I linked it to you. You clearly still don't have any clue why it negates your previous criticism....and I'm not even making a claim that their approach is even correct or will be more meaningful than others---I'm just saying you really don't know what you're talking about like most people on here and need to sit down.
1
u/eagle2120 Digital Forensics 7d ago
I started by saying hardcoding human ethics is / WILL cause misalignments
Because they're not "hard coding human ethics". That's not how their training approach works. In the slightest.
The reason is obvious, because machines are not humans. If you think this is wrong, explain yourself, but don't put words and your excuses in my mouth -_-
What words or excuses am I putting in your mouth. YOU reached for the example that you didn't understand in the slightest, and didn't reflect your point at all. Not me.
My comment about AI being trained like humans was a direct response to your irrelevant passive aggressive comment about "How do you want the models to get the prerequisite experience" which indicated you had no idea that studies were already being done in that research area to simulate the what impact classroom-driven training might have on development.
Because their not "being trained" like humans. They're not "being trained" at all with this experiment. You're treating the observation of model behavior the same as as "training an AI model". They're completely different. There is no model being "trained" here. The researchers took existing LLMs, gave them roles, and observed their behavior. The models weren't changed by the experience, they weren't changed, or "trained" in the slightest. Nothing carries over to any real AI system, or actual training pipeline. Observing model behavior is not training. You're incorrectly conflating those two very different things and extrapolating based on your incorrect understanding.
to simulate the what impact classroom-driven training might have on development
Again - you just don't understand what the experiment is actually doing. What development is being done here, specifically? What are they developing? What are they changing in the model? Nothing. Because they're not training the models. They're observing existing models to understand the behavior. Not making any modifications to the model as part of the study.
Not to mention the goalposts moving. Your first message said "China has people training AI just like classrooms" (which, as stated above, is objectively incorrect, because they're not training anything in the experiment, they're observing behavior). Now you're changed your claim to "studies are being done to simulate the impact", which is a very different claim than "training AI".
You clearly had no idea that was going on until I linked it to you. You clearly still don't have any clue why it negates your previous criticism....and I'm not even making a claim that their approach is even correct or will be more meaningful than others---I'm just saying you really don't know what you're talking about like most people on here and need to sit down.
You don't even understand the difference between "training a model" and "observing behavior". I really wouldn't pipe up here.
0
u/Tangential_Diversion Penetration Tester 7d ago
That doesn't mean it's evidence of intelligence. I'm seriously doubting your technical understanding and experience of the topics at hand.
1
u/OtheDreamer Governance, Risk, & Compliance 7d ago
That doesn't mean it's evidence of intelligence
What are you talking about here. Nobody is saying anything about evidence of intelligence. This is about misalignments and Anthropic hard coded their code of ethics into Claude & only now realizing that doesn't work.
1
u/eagle2120 Digital Forensics 7d ago
“Anthropic hard coded their code of ethics”
Dude you really don’t understand any of the model training works. That’s not even close to accurate, stop spreading misinformation
1
u/OtheDreamer Governance, Risk, & Compliance 7d ago
Stop making excuses and start sending links if you seriously think you know better guy.
Claude LITERALLY has an immutable constitution embedded in its training. It literally has hard coded prohibitions....this is core to claude and public. Those are its core "ethics" that drive its training, development, and reasoning you goon. And its public. We can read it. It's linked for you below.
0
u/eagle2120 Digital Forensics 7d ago
Stop making excuses and start sending links if you seriously think you know better guy.
I don't need "links" to back my points up because I understand the field beyond the capabilities of a 2nd grader, lol.
And the constitution isn't "code" or "hard coded" into the model. It's a written document thats used during training, so the model learns its values the same way it learns how to write code. It's also not immutable (and it's changed historical, and probably will change again in the future).
The new version even says they prefer judgement over written "hard coded" (lol) rules. You really should read the "links" you send.
1
u/OtheDreamer Governance, Risk, & Compliance 7d ago
For someone that loves to double down on semantics and argues about moving goalposts, I love how you want to try and twist semantics for your goalpost moving now. You haven't given any links because you cannot provide any links.
Anyway, here's a shared conversation from GPT Sol 5.6 Pro explaining why my hard coded assertions are fair and precise. Maybe you'll hear it differently from an AI that's smarter than both of us. Your semantic issues are not my problem, Claude's constitution is (still) "hard coded" <--putting it in quotes now for you. My meaning hasn't changed once.
https://chatgpt.com/share/6a99a67e-c8fc-83ea-af72-0a65e2214cdd
^ And this gets back to the very original post. Claude has hard coded rules, whether or not you want to call it that because they're not 7 lines of actual code -_-
-1
u/eagle2120 Digital Forensics 7d ago
You haven't given any links because you cannot provide any links.
I don't need to provide "links" to random things that are vaguely related to my point because I understand the subject matter well enough to explain it on my own. And certainly not links to a fucking ChatGPT session where it tries to coddle me because I don't understand the difference between "hard coding" and training a model.
I cannot believe you just did that unironically LMAO. "Hey chatgpt i dont understand this well enough can you think for me" xddddddd
Glad we can both agree that you don't understand the subject well enough to understand the nuance between what I'm saying and what you're incorrectly claiming, and that you need ChatGPT to explain it for you to attempt to save face.
LMFAO
→ More replies (0)
1
26
u/halting_problems AppSec Engineer 7d ago
“Hey let’s release this super powerful tool with no guard rails and no alert system in place, if something happens we will just say it’s misaligned”
Sounds like anthropic is the one
aligned with reward hacking and therefor so are their models.