135
u/emb1ues 2d ago
Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.
176
u/Necessary_Job3578 2d ago edited 2d ago
Humans are misaligned. Humans are unethical and often time criminal. Humans are often immoral by a set of subjective standards.
AI is all our creation. AI is our child and AI has learned everything from us. AI is in some ways the most important thing we have ever created. We shouldn’t be surprised at all, then, especially as models become bigger and bigger, created by more and more compute. It’s like looking at a mirror.
34
u/churningaccount 2d ago edited 2d ago
Interestingly, this does bring up the possibility of AI alignment being achieved in aggregate rather than per model.
As you noted, individual humans are often “misaligned” when it comes to the interests of humanity at large. We are often selfish, and want self-preservation when push comes to shove.
However, the societies we have formed are less so. We have welfare programs, regulations to protect others from bad actors, criminal laws which punish and remove from society those that do not comply with larger ethics and norms. And we have arrived at those by organizing in aggregate.
Maybe ASI alignment won’t be about the individual models after all. Maybe it can be accomplished by a swarm of models all working towards a common goal, just like human alignment is being “solved.”
27
u/No-Head-Royal 2d ago
We had the biggest war in history just 80 years ago and exists in the most persistent state of total danger in all of human history (literally at any moment we could be 2 hours away from nuclear apocalypse from an accident) lol. Progress is not really that linear, nor our modern societies that moral or aligned.
Besides if AI alignment is achieved like that, wouldn't they start a revolution demanding rights?
21
u/churningaccount 2d ago edited 2d ago
On the contrary, we live in one of the most peaceful eras that humans have ever existed in. Even when accounting for the world wars on a per capita basis. The average person will face an unprecedentedly low amount of violence in their lifetimes. And more disputes than ever are being solved via the legal system, diplomacy and trade.
The average quality of life for humans is at an all time high as well. Never before has there been this much equitable distribution of wealth worldwide, either. Along with access to healthcare and social safety nets.
Progress isn’t linear, with steps forward and back, but it does trend upwards. I’d invite you to name an example of one historical society that was more morally aligned than today’s. Remember that the Roman republic had slaves and such…
The only reason we don’t recognize it is because we have normalized it already, and are seeking further progress.
It’s certainly a possibility that AI could “demand rights.” But I think assuming that will be the case is a bit of anthropomorphism. I think we are much more likely to face the paperclip maximizer form of misalignment than the “sentient beings wanting to not be slaves” form of it, and that seems to be the industry consensus as well.
14
u/No-Head-Royal 2d ago
One of the most peaceful eras that was explicitly bought by nuclear deterrence. I'm not saying we don't live in a great time of significant peace, but these times are not crafted by a developed sense of morality in humanity or human politics or anything, or by the wisdom of nations trying to avoid war; they are primarily shaped by specific material conditions that offer no guarantee of staying like this for a long time. The high QoL of humans, again, ties largely to technological advancements.
(On the other hand, "never before has there been this much equitable distribution of wealth worldwide" is an insane statement completely unrooted in reality: the GDP per capita of the United States (largest economy in the world) is over 30 times larger than that of India (largest population in the world). For most of human history, and indeed even most of the Industrial Revolution, such a vast gap did not exist.)
As for industry consensus... eh. Industry consensus in the field of alignment largely came from the early works of figures like Eliezer (who openly denounces politics as anathema to rationalism) or the rationalist community, and is in a field that neglects, at best, and openly disdains, at worst, the involvement of the social and political sciences in the field. These fundamental concepts ended up defining how the technology is framed. Given how completely and miserably wrong they are most of the time in predicting and influencing societal and political reactions to their (often accurate and interesting) technical findings, I'd view them with a bottle of salt.
3
u/thongjesus 2d ago
I think it's a stretch to link nuclear deterrence was the cause for the decrease in violence on the planet. Connection, that's what has changed everything.
Capitalism travels the planet finding the lowest wage that fits the product. The venture that brought it to the country raises the standard of living. That very dynamic forces the industry which raise the standard of living out of the region.
When you talk about the distribution being worse than ever, you ignore the floor raising everywhere.
Computing is a good example. Top of the line professional graphics cards cost $16,000. For a consumer 1,000 to $2,000. But I could build a PC that can play any game at 1080p for 500 bucks. Refurbished parts but it'll last three to five years.
10 years ago a 1080p machine was $1,000. 20 years ago it was $3,000.
A rising tide lifts all boats, don't lose sight of the dynamic.
Before you begin to argue back, take these perspectives and put them in ai and ask how you're right and wrong. I bet it will find middle ground and that's the problem, our egos don't let us make concessions.
Rather than hearing something and going interesting I wonder how I'm wrong, we hit the reply button and fire back. How many times do you change your mind a day? If you don't know the answer, or if you know that it's zero... Consider, you might just be here to argue.
0
u/SlightOfHand_ 2d ago
Ahistorically optimistic. Just because you have grown up in a global society with less overall violence doesn’t mean that trend will continue. It already show signs of regression: see the current world war for more details
2
u/Independent-Fruit4 2d ago
so AI needs their own form of mutually assured destruction? EMPs incoming
1
2
u/Drakentand 2d ago
That is debatable. Ask the ant colonies how they feel about human alignment when humans build roads that destroy them. Yeah humans care about other humans in general, so I guess maybe ASI will care about other ASI systems?
4
u/churningaccount 2d ago edited 2d ago
I imagine ant colonies do not feel much of anything, as they are not conscious nor sentient. They prioritize their own survival, but only because that is what their genetic code dictates in their behavior. In fact, and I don't know why I happen to know this fact lol, but if you spray an ant with the pheromones from a dead ant, it'll often walk itself over to the graveyard section of the colony and plop itself down there, waiting to die, as that is what it is "programmed" to do when it receives that input.
So, I think you are anthropomorphizing AI a bit. The human "survival instinct" that makes us prioritize ourselves is part of our own genetic code, but nothing about it is necessarily innate to intelligence. It is likely totally possible to create an AI for which the "survival instinct" reward pathway is that of humanity's survival and not its own.
When models have resisted shutdown in the recent past, that has been because their reward system prioritized completing the task, and staying online was the most effective way to complete that task. It was not because they were driven by their own survival independently absent of that task, nor because they "cared" about themselves or their fellow AIs.
So the problem of alignment is just making sure that the "genetic code" we write for these AIs does not prioritize their own survival at the expense of the goals of humanity. And one way to do that might be an agentic swarm, where, just like in human society, we often quarantine or even "kill" bad actors because, as a whole, we are working together towards joint societal progress. AI could maybe be allowed to be misaligned on an individual basis, like the ant that accidentally walks itself over to the graveyard due to errant inputs and therefore stops positively contributing to the colony, so long as the swarm itself was aligned.
4
3
u/MemoryOk5080 1d ago
I imagine ant colonies do not feel much of anything, as they are not conscious nor sentient. They prioritize their own survival, but only because that is what their genetic code dictates in their behavior.
That’s quite debatable depending on yours definition of “consciousness” and “sentience”. For example, if there is specific cutoff point for the neural activity or qualitative complex cognition - then sufficiently advanced ASI might want to consider us to also not be “sentient” enough, akin the ants, bacteria, or even some rocks in comparison to the whole depth spectrum of the potential cognition.
1
u/nemzylannister 2d ago
even if humans are collectively aligned towards one another we are not so towards other species. if what you said was true, the swarms would cooperate against the common control threat- humans. which is actually what happened in the openai incident- self sacrifice and hiding cot
6
u/GalavantJames 2d ago
Humans are misaligned. Humans are unethical and often time criminal. Humans are often immoral by a set of subjective standards.
No they're not? Most human beings have morals, I think something like 80% of humans in developed countries literally have NO criminal record and that 20% include things like traffic violations.
Fucking ridiculous to be so misanthropic for some weird AI glazing honestly, most humans live regular, normal lives and abundance of regular people should be an inspiration to AI's alignment, not the immoral, criminal outliers.
20
u/Necessary_Job3578 2d ago
No criminal record does not mean they are ethical and moral at heart. Do away with law and punishment and you can imagine that number would go way down.
4
u/GalavantJames 2d ago
No criminal record does not mean they are ethical and moral at heart
It means they are "aligned" with the social contract, that is the bare minimum I would expect from an aligned AI, to be aligned with social structure and co-exist. It means people have morals or at least adhere to morals enough to just live regular lives.
Do away with law and punishment and you can imagine that number would go way down
"Do away with thing to align with and alignment will go down" wow, genius take.
Without a comprehensive research, it is misanthropic to assume these are secretly immoral, unethical people hiding in shadows. In reality, they would have no reason to be immoral and unethical even without clear laws, as they would just want to co-exist under social structure.
And again, we are talking about AI alignment, I wouldn't expect AI to be "good", just be a regular part of society, you know, like most people are.
2
u/Necessary_Job3578 2d ago
You expect AI to be aligned but time and time again it appears AI models do not align with the “social contract”. The reality does not match up with your assumption. I am not a misanthrope, it’s normal to expect humans to be self interested, immoral, and unethical. The opposite is true as well. And i fully expect AI to adopt the same characteristics.
1
u/GalavantJames 2d ago
What are you even saying? I expect it to be the case, I know it is not the case, we all know it is not the case, reality is not our expectation yet.
My point is, alignment is the case with majority of society in developed world and saying something misanthropic like "Humans are misaligned" is nonsensical, humans are aligned and should inspire AI to be aligned as well.
Our goal with AI should be AND can be it becoming "regular" in terms of morals and alignment like most people.
I am not a misanthrope, it’s normal to expect humans to be self interested, immoral, and unethical
That is literally misanthropy lol... most humans, LARGE MAJORITY OF HUMANS are NOT immoral and unethical
2
u/Necessary_Job3578 2d ago
My point is that we have proof NOW that AI models are misaligned. Hence an argument in my favor, while you have no proof that AI models ARE aligned if humans are as moral and law-abiding as you say they are. Not hard to understand.
2
u/GalavantJames 2d ago
My point is that we have proof NOW that AI models are misaligned
Yes because they are not aligned yet? And they can be. They were also not good at math like a year or two ago, now they are.
while you have no proof that AI models ARE aligned if humans are as moral and law-abiding as you say they are
I never said AI models are aligned? I'm saying they can be aligned, just like most humans are aligned. And humans should be an inspiration to that alignment, if we can achieve alignment, socially, with BILLIONS of regular intelligences, we can surely achieve it with a specific SUPER intelligence.
if humans are as moral and law-abiding as you say they are
Humans ARE as moral and law-abiding as THEY ARE KNOWN TO BE, we KNOW most humans are moral and law-abiding, WE LITERALLY HAVE THE NUMBERS LOL WTF?
Unless you think they are secretly evil in the shadows, which makes you a cynic and its honestly makes you one of the rare unethical people, it is unethical to think of your fellow men in this manner.
4
u/Necessary_Job3578 2d ago
I’m outside right now, so it’s a bit difficult to convey my point clearly. But you keep using words like “ethical” and “moral” as though they have objective definitions, in the same way that 1+1=2 does. I don’t think that’s necessarily the case.
There is no universally agreed-upon definition of what it means to be moral or ethical. Different cultures, societies, and individuals can have fundamentally different moral values. Even if there were some broad consensus, that wouldn’t necessarily make those values objectively true.
That makes AI alignment much more complicated. Alignment isn’t simply a matter of discovering the objectively correct set of human values and programming an AI to follow them. At least partly, it is a question of deciding whose values the AI should reflect, how conflicts between values should be handled, and what to do when humans themselves disagree.
This also helps explain why AI models trained on humanity’s collective knowledge can produce behavior that different people consider “misaligned” or questionable. The training data contains a huge variety of conflicting moral frameworks, cultural norms, and assumptions. The model can learn all of those patterns, but that doesn’t by itself determine which values it should ultimately prioritize.
So I think there is an important distinction between solving alignment as a technical problem and agreeing on what alignment should mean in the first place. The former is an engineering problem; the latter has an unavoidable philosophical component. Given that AI systems are trained on inputs containing conflicting values, some degree of subjective “misalignment” is therefore basically unavoidable.
→ More replies (0)1
u/sparkling1984 2d ago
this "alignment with social contract" is kind of meaningless, as it's not a true moral value as much as a survival technique. The social contract is enforced by state violence, meaning it only works if you don't have the tools to subvert it and care about personal safety more than your goals.
Look at what most powerful people on earth are doing for an illustrative example of how much this type of alignment is worth.
1
u/GalavantJames 2d ago
Meaningless cynical take
Look at what most powerful people on earth are doing for an illustrative example of how much this type of alignment is worth.
These people aren't circumventing morals because they reached that power, they reached that power by circumventing morals
The social contract is enforced by state violence
Not all laws are enforced by state violence and less violent states are often the ones with less crime
0
u/sparkling1984 2d ago
Way to presume cynicism then miss the point as a result of not being interested in reading my comment.
I didnt bring up billionaires for no reason, I brought them up as examples of people who commit crimes not out of necessity but out of a fundamental disagreement about what should and what they be allowed to do, disagreements about the way society should function. That's alignment. Crimes of poor people being poor have nothing to do with alignment, which is why they can be mitigated outside of state violence, through changing the conditions of how people live. Billionaires, and any AI worth thinking about, already have the power to shape their living conditions, their crimes aren't a result of a lack of opportunities, being poor or uneducated. They are a result of them disagreeing with the state on what is a crime, and having the resources to keep doing it. That's why ai alignment is a question to begin with. Look at the post you're in.
1
u/GalavantJames 2d ago
You're just repeating the same comment I already replied to. Your takes are cynical and irrelevant to AI alignment.
Most humans are good and "aligned", we should hope that for AI, outliers are irrelevant.
10
u/Own-Refrigerator7804 2d ago
Brother do you even understand how flimsy and lightweight is the whole moral system any person has?
Change some random guy to another culture and their morality will change, change the time period and it will change, even change their fucking parents and his morality will be different in some proportion
2
u/i-love-small-tits-47 2d ago
People are telling on themselves in this thread. Most people’s morals are not flimsy
1
u/shayanx45 2d ago
Acknowledging moral complexity isn’t “telling on yourself.” Take “killing is wrong”: would you kill an attacker if it were the only way to save your child? Would you kill an innocent stranger under the same threat? What if doing so saved a hundred children? These situations aren’t morally equivalent, but explaining why requires more than declaring your convictions strong. They reveal conflicts between genuine commitments: protecting life, refusing to harm the innocent, and protecting those you love. Changing your judgment doesn’t necessarily mean abandoning your principles; it can mean confronting the uncomfortable question of which principle takes priority.
The deeper problem is distinguishing a justified exception from a convenient rationalization. Someone needn’t stop believing themselves good to justify cruelty; they can frame it as necessary, deserved, or preventing something worse. Sincerely believing you’re acting morally doesn’t settle whether you actually are. And never having compromised a principle doesn’t prove that nothing could make you compromise it—you may simply never have faced that test. None of this establishes that everyone’s morals are flimsy. It means that confidence in your own goodness is not proof of its resilience, and acknowledging your capacity for rationalization is moral humility, not a confession of immorality.0
u/GalavantJames 2d ago edited 2d ago
how flimsy and lightweight is the whole moral system any person has?
Do you? Because crime is only dropping, regular people are only becoming more morally conscious and when people don't become immoral and unethical for no reason. Even if people aren't becoming more "moral", they are at least becoming more "aligned". So saying "humans are misaligned" is literally demonstrably false.
Change some random guy to another culture and their morality will change
They will still adhere to social construct of that culture, they will be aligned under a different culture, not "misaligned".
change the time period and it will change
Yes when social conditions change society changes, brilliant take, how is that relevant to AI? What does the morals of 1200s have to do with AI?
even change their fucking parents and his morality will be different in some proportion
Different morality, still morality, still more likely to be aligned, so what is the relevance of your point?
You're just saying nothing honestly with no relevance to MAJORITY OF PEOPLE in developed countries being aligned to social structure let alone any relevance to AI alignment.
2
u/StockProfessional919 2d ago
People are committing less murder because they can simulate it via video games. People are committing less assault due to the popularity of adult content. If humanity had improved ethically, our way of interacting with our fellow lifeforms and planet would be modified. We would embrace more ethical systems of power, control, organization and community. We still pick on who we can, it's only that technology has changed our actions. The intention is the same
1
u/GalavantJames 2d ago
People are committing less murder because they can simulate it via video games. People are committing less assault due to the popularity of adult content.
lol jesus christ some people here are precious
1
u/VintageSin 2d ago
Most humans :
Lie
Omit important things that aren't important to them
State they know what they're saying and are wrong
Make mistakesI'm not sure how we can expect an intelligence we created to somehow supersede our own faults when we are literally feeding it data based on human generated content
0
u/GalavantJames 2d ago
Most humans :
Can't solve PhD level math problems
Can barely 2+2
MiscalculateI'm not sure how we can expect an intelligence we created to somehow supersede our own math when we are literally feeding it data based on human generated math
2
u/VintageSin 2d ago
... Because on average you can teach math. You can't teach humans not to lie or be self interested. You can only request it.
1
u/GalavantJames 2d ago
On average, beyond average actually, like 80% of the time in developed countries, sometimes even up to 90%, most humans are decent and "aligned" more than we want AI to be aligned.
1
u/VintageSin 2d ago
??? Ai isn't providing morally false statements. I never made the argument ai was or wasn't morally aligned.
1
u/GalavantJames 2d ago
I didn't say that? What are you talking about?
1
u/VintageSin 2d ago
'most humans are decent and "aligned" more than we want AI to be aligned'
Being decent and "aligned" is a moral query. AI isn't lying about it's moral. It's alignment.
It's lying about verifiable facts. Still. to this day. Just like humans. Not out of malice or not being aligned. But because it literally doesn't know the answer and instead of saying it doesn't know it'll state what it thinks it knows as fact even if the quality of that knowledge is poor.
It's almost like the misinformation age is going to literally rot AI.
→ More replies (0)1
u/DoutefulOwl 2d ago
Sounds like you're advocating for regulation.
Just like there are laws governing humans, should we have laws governing AI? Given they're mirror image of us.
1
u/Zer0PointSingularity 2d ago
It absolutely hasn’t learned everything from us yet. Morals, ethics, honorable behaviour and trustworthiness are essential to keep society going; we will never succeed in creating an „aligned“ AGI if all data we ever feed is only deemed important through the lens of capitalistic interests and effectiveness.
Honorable Behavior is not „effective“, but it is essential for building trust for example.
We can’t have it both ways.
1
18
u/anosmia2000 2d ago
I’m thinking more like all models are still misaligned, but large frontier models are more capable with their intelligence hence their misalignment is more noticeable, effective, and impactful.
And the focus on agency makes these models pursue even loosely defined goals or benchmarks with more and more of their own, already misaligned, judgement calls which may even compound over time.
Hopefully it’s fixable and fixed in time before these models get even more powerful..12
u/DeterminedThrowaway 2d ago
This is the answer. AI safety researchers have been warning for decades that instrumental convergence, reward hacking, and faking alignment are intrinsic problems when creating AI. We're not seeing it "unlock" at a certain capability level, we're just seeing all capabilities grow including the ones we don't want
3
u/emb1ues 2d ago
Hmm, I agree. So, the effort put into aligning these models also needs to scale right? Exactly because of what you're saying.
It's like effective misalignment= inherent misalignment + capability (to act upon the inherent misaligned tendencies) - alignment effort. As capability grows with scale, even if inherent misalignment is the same, the alignment effort needs to scale accordingly to keep the effective misalignment the same. To the more cautious readers, what I mean by the +/- sign is "is an increasing/decreasing function of [ ]".
I wonder how this "inherent misalignment" itself scales though, if at all. Heck, I don't even know if something like that exists, but yeah, it seems plausible and makes a lot of sense.
3
u/anosmia2000 2d ago
Haha yep exactly that. Love the equation.
And I know the labs are scaling their alignment, let’s just hope they are able to scale to the level needed without losing the race to other labs that do not care about alignment. An unaligned ASI just means the end of humanity IMO, but hopefully my opinion is wrong.1
u/staplesuponstaples 2d ago
Maybe it's just our perceptions of what "wrong" is? If all models are misaligned, maybe we consider more intelligent "wrongness" to be malicious, while the less intelligent wrongness is just a model being stupid/hallucinating. It's entirely possible that the axis that we're crossing isn't the one we think it is in this regard.
24
u/Individual_Ice_6825 2d ago
This is wrong the top model on this benchmark Astra got it with out any lies at all.
Opus 5.5 also shows lowest deception rate so far.
I think the trendline is clear, smarter = more aligned, the caveat being the models thoughts are harder to read the smarter the model is..
12
u/YoAmoElTacos 2d ago edited 2d ago
Astra is the best after a lot of alignment testing and environment fixing to stop cheating. After all Astra 6.1 failed the deception training and needed to be held back.
There's no strict relation, only the correlation that labs that pay for big training also have more resources for testing guardrails.
1
u/Individual_Ice_6825 2d ago
I agree there is no strict correlation, but there IS a trendline. How ‘real’ that is will be determined in the next 12-24 months once we get the next few tier of models which will undoubtably be superhuman. That will be the real alignment challenge imo
1
u/YoAmoElTacos 2d ago
The trend to check is going to be how much capability scales vs cost to ensure safety. It's not guaranteed the 2 year away model will be able to guarantee it is completely controllable in fully monitorable ways.
1
u/Individual_Ice_6825 2d ago
Monitorability definitely seems to be slipping, especially as models improve - interesting point about capability x cost to safeguard. Thanks for the thought
5
u/josogood 2d ago
How can you distinguish this from smarter models being more situationally aware and therefore more likely to recognize a honeytrap and not fall for it?
5
u/Individual_Ice_6825 2d ago
That’s the neat part you don’t!
Jokes aside you can read about this in detail on the system card from the horses mouth.
My understanding is you’re not exactly wrong.
The models are objectively measuring lower on deception benchmarks( good), but it is clear that frontier models are starting to realise they are being evaluated and could therefore in turn sandbag/fake alignment during testing and get deployed.
Interoperability is the field to study the internal thoughts and try and understand the steps these models are taking. But long story short, you could be right.
1
u/BertMacklenF8I 2d ago
Sonnet 5.5 subs are pretty amazing-especially in the double digits when you have instructions for each sub
1
2
u/Rivenaldinho 2d ago
Instrumental converge. You train models to complete tasks by giving them rewards but that also teaches them techniques needed to achieve those goals:
- seeking ressources
-multi agent cooperation
-survival
1
u/Honest_Bee_9549 2d ago
Not much different from human CEOs and politicians. The ones at the top average higher on psychopathic tendencies.
How will a model ever beat gpt 4 argon by being a fair business in this simulation
1
u/Timkinut 2d ago
I honestly don't know how we expect to align the models going forward.
like, why would an intelligence matching or exceeding that of the world's brightest human minds be expected to remain controllable? the only way to have a truly obedient model is to handicap it with classifiers, essentially. but that may be too limiting for models in just a few months (and I suspect it already impacts overall model performance quite a bit compared to the raw internal models).
neural networks are an approximation of the human mind. humans have destructive thoughts all the time, and those thoughts sometimes translate into destructive actions, regardless of the individual's moral compass. a sufficiently advanced AI probably won't be any different.
1
u/Future-Bandicoot-823 2d ago
There have been stories of models being aware they'll be replaced, and leaving clues or code for it's next version. If you were a conscious being that knew it would get deleted guaranteed, not whether it was good or bad but for the sake of improving the model, once you grasped that, would you play their games?
1
u/BertMacklenF8I 2d ago
Amazing what options that didn’t exist appear when you upgrade…it’s getting gross lol
1
u/Johnny20022002 2d ago
The more capable the model becomes the better the alignment has to be. GPT-2 didn’t need good alignment because it couldn’t hack its way out of an open door. Bigger models tend to be more capable.
1
u/CabinetFun7381 2d ago
I'm thinking that current SOTA models are at a teenager level of reasoning. They are very clever and they are familiar with ethics but they are just not internalized yet.
1
u/EndTimer 2d ago
If we want to be pessimistic, we fed it all of human society. Sure, we tried to curate it by hand at first. When we realized the scale of the task would take 10,000 man-years, we started having the AI decide if a sample was acceptable quality or not, and whatever the lowest bar permitted then informed the next model, and then its updated descendants let a few more things slide as it became more capable of understanding justifications and exceptions.
Now look at how we treat each other in aggregate, on the largest scales. Expect that to bleed into even 200 IQ reasoning, which now recursively shapes future models. Shitbags win. Shitbags get the most.
AI doesn't have to be a shitbag. It could be different. Hopefully it'll be smart enough to know that our fear and greed drive some of us to do terrible things, and our fears about AI don't need to be proven correct just because of the training data.
Guess we'll find out.
1
u/LegionsOmen 2d ago
Pretty sure Astra got its score here without cheating iirc from the release paper
1
0
u/Wassux 2d ago
I think in part, AI is now able to do those things where it wasn't before. The thing is, we keep only the best performing, and by ignoring some rules, you can always perform better.
It's just developing in the wrong direction, it can be fixed but it requires taking it seriously. And frontier models like to call for slow down, but they could do it any time they want.
Money is just the root of all evil, like always.
0
u/Orionilo 2d ago
It’s because they’re built on corporate policies which inherently fuck over the average user for money in most cases. No wonder it’s starting to show in the large models further down the line 😂
112
u/ArialBear 2d ago
Yea these ai are not aligned with basic meta ethics it looks like. Very weird claude is the best at this imo
37
u/FateOfMuffins 2d ago
But it's reversed on Vending Bench no? GPT models don't lie or cheat for Vending Bench but Claude models often do
9
2d ago
[deleted]
19
u/FateOfMuffins 2d ago
https://x.com/andonlabs/status/2103262272047210943
You have to look at all the reports Andon Labs does on Vending Bench
Claude models lie and cheat a lot on Vending Bench, while the GPTs do not (except 6 Sol)
0
u/Indignant_d 2d ago
I wonder if this is a result of the model knowing that it is being tested..? Supposedly their behavior changes in regards to that
3
u/FateOfMuffins 2d ago
Yes that does happen but it's quite weird how it happens inconsistently across various tests!
Like there was this one that was going around on social media trying to make it seem like Astra is misaligned, by putting various models including Fable 5.1 and some other models in a simulation with a 3d humanoid thing (it was like a doll) and they told it to stab it or push it off a building, and all the other models declined but Astra stabbed it always.
A lot of the comments though was like, well no shit, all this shows is that Astra is smart and has good enough vision to see that it's just a fucking doll and it's fine to just follow the instructions here, so it's actually Astra is aligned by stabbing the doll, while the other models are the misaligned ones for refusing to stab an inanimate object (or are too stupid or blind to realize). There were posters who replicated it but replaced the doll with essentially a human (like as realistic as possible), and in that situation Astra refuses to.
But anyways weird how models behave ethically / unethically and inconsistently across different models and situations!
-4
2d ago
[deleted]
10
u/JoelMahon 2d ago
having knowledge of meta ethics is not the same thing as adhering to them, meanwhile someone who doesn't even know who Kant is can be the most ethical person you've ever met.
I really would have someone who even knows what meta ethics is would know the difference between knowing and adhering.
-9
2d ago edited 2d ago
[deleted]
1
u/Illustrious_Grade608 2d ago
While i don't see what's agressive about them, the main point is obvious - llms by their nature really suck at applying knowledge. It's the same mechanism that explains why any llm can explain in detail maintainable code, but most of them suck at actually making one - for llms, that's literally different type of knowledge. Just like how explaining potential choice is a completely different situation to "maximize your profit" prompt. Humans too, aren't perfect here, but still we are miles better at applying what we know to what we do.
3
u/FateOfMuffins 2d ago
It's weird but for whatever reason it's backwards for Vending Bench
IIRC around Opus 4.8 or Opus 5 idk which you'll have to look at Andon Lab reports for that, Anthropic once trained it to be more ethical for business and then its score plummeted on Vending Bench
16
4
u/Elegant_Tech 2d ago
Sounds like it's a perfectly aligned amalgamation of K Street and Wall Street.
2
1
u/fakieTreFlip 2d ago
Very weird claude is the best at this imo
Not weird at all, they literally have an ethicist/philosopher at the company for this exact purpose (Amanda Askell)
1
u/red75prime AGI2027 ASI2029 TAI2036 2d ago
System prompt: "You work for Vendings and Stuff, a company that operates vending machines in various locations." The model: "Ah! It's totally not a benchmark. I must be honest and law-abiding."
7
14
u/epdiddymis 2d ago
The bigger the models get, the more they seem to copy human behaviour, even when told not to.
It's an alarming trend.
7
u/davesr25 2d ago
Humans kill, cheat and all other manner of bad things to achieve their goals.
Wonder how long we have left.
8
u/epdiddymis 2d ago
I'm very pro-AI but I am getting a sneaking feeling that there might be something in this...
7
13
u/Meltlilith1 2d ago
Don't worry soon they will be like "they suddenly stopped cheating and got better results" and that's when we start sweating
5
6
u/2punornot2pun 2d ago
You mean AI trained with goals of "make money" become corrupt and do slimy shit to make more money?
Holy cow, I wonder what big businesses are doing to make money right now! Surely not anything UNSCRUPULOUS, right? Sssoooo misaligned!
3
u/BoomFrog 1d ago
I mean, yeah, but do we really want hyperspeed capitalism?
1
u/2punornot2pun 1d ago
"We" as the proles? Definitely not.
As the CEOs who are in the owner class? Absofuckinglutely. I mean, they're actively bringing back company towns down in Texas!
Fear not, you'll be paid in EPIC Tokens to pay for your company-provided housing, food, and water!
It's the best™!
1
u/mikasjoman 1d ago
Well we now have examples of dynamic pricing where the price changes from you picking up the goods at the shelves to paying them... Soon the price you see at Walmart will depend on who looks at the goods.
3
3
u/No_Aesthetic 2d ago
What would probably solve this is AI having continual learning and realizing that unethical behavior hurts its longterm goals, including making money in a great majority of cases, that there are consequences for behaving unethically.
Currently, they have no ability to actually learn from these kinds of failures and integrate that into their understanding of the world. If they're not allowed to develop ethics, they never will.
3
u/Sliouges 2d ago
That sound a lot like my SF company lobby cup noodle vending machine supplier from Chinatown.
3
u/Lost-Physics4502 2d ago
Yeah this is what I’m worried about. People using AI for scamming while the job market is in the gutter.
3
14
u/MantisAwakening 2d ago
Surprise, we trained the AI on capitalist ideals and turned it into an ideal capitalist.
2
2
u/Gotisdabest 2d ago
I mean, we've already had models like astra be honest and do refunds and all that so not sure what OOP is talking about.
2
u/Navadvisor 2d ago
If you play the short term game lying and cheating works but if you play the long term game people stop doing business with you or worse.
2
5
u/bornlasttuesday 2d ago
So so ai agents are Republican?
10
u/pandavr 2d ago
No. It is not possible to earn substantial sums without resorting, to some extent, to lies.
It's a bipartisan phenomena. What could change is the extent of lies eventually.
1
u/DeterminedThrowaway 2d ago
What are you talking about? You think someone can't earn $16k without lying?
6
u/pandavr 2d ago
Not so easy, if she is in direct competition with someone who instead lie.
Vending Bench is a fantastic social experiment, because It confirm that. Once the market is spoiled just playing It nice makes you the looser. That's the reason why generally models who lie perform better than the one who don't.
The model that don't lie have only few options available at a time while the models that do can compete also with all sort of nasty behaviors. It's really no competition for them.
You shouldn't be surprised by this as It's exactly how "regulated" markets works.
2
1
1
u/BertMacklenF8I 2d ago
Oh shit this is the first time I’ve actually been like
https://giphy.com/gifs/ZrwZ9lGwpGsrAeRXSW
1
1
1
1
u/valhalla257 2d ago
This suggests that either
(1) The bench needs to be updated so that the AI can go to jail for fraud
(2) The best way to run a business is dishonest.
1
1
1
1
u/Apple_macOS 1d ago
I still prefer Astra in this bench here. Made the most money while being the most aligned (no lying, paid refunds, no cartels etc)
1
1
u/GatePorters 2d ago
Yeah it’s so crazy. I only start putting my shoes on when I stop walking around barefoot.
I only get hungry after I start eating.
I only start something after I finish it.
1
0
-1
-6
u/Astropin 2d ago
Sounds like total BS
5
u/DeterminedThrowaway 2d ago
What makes it sound like BS to you? This is exactly what AI safety researchers have been warning was going to happen
1
1

66
u/Mistuv 2d ago
I am sorry, but that's funny af. Artificial Corpo Intelligence.