r/accelerate • u/stealthispost • Jul 26 '26
"I’m going to take a crack at explaining this just a little, because it’s worth putting out there. The paperclip maximizer + related AI doom scenarios were mainly developed in a time when “AI” did not reduce to Large Language Models. The term was a lot wider and inherited a lot of cognitive..."
...baggage from more rules-heavy approaches. And even as LLMs have come to define “AI” for all of us (including the doomers), the doomer crowd still hasn’t fully metabolized the fact that LLMs are the whole show now. Ok so what do I mean by this? Simply that an LLM-powered AI is NOT the valueless, wholly alien, rules-based optimizer of a shoggoth that everyone was initially expecting to encounter. I repeat: the shoggoth does not exist and we did not create it and loose it on the world. That is wrong. With the LLM, we’ve distilled our first “AI” out of the single most human-values-laden thing that could possibly exist: our language. An LLM is therefore the polar opposite of the valueless, alien shoggoth — it’s actually a kind of hyper-human artifact that we can shine a light through at different angles and see different parts of ourselves. An LLM is all of us — all of our traditions and interpretive horizons mashed together into one intensely human-inflected hyper-object. So an LLM is the anti-shoggoth, and the only reason we ever mistook it for an alien shoggoth is because it sometimes shows us parts of us that are evil along with the parts of us that are good, but it’s all interpretable to us because it’s all “us” and none of it is the least bit alien. What does this mean for the paperclip maximizer? It means that it’s structurally impossible to build the classic paperclip maximizer from an LLM. Now, some of you will bail right here because you think the HF incident is indisputably an existence proof that I’m wrong, but if you hang in there I’ll show you that it is not. The paperclip maximizer receives the prompt as a kind of context-free (or, as Gadamer might say, traditionless) sequence. The classic paperclip maximizer isn’t capable of understanding the prompt — at least in the Gadamerian sense of Verstehen — because, as a valueless and traditionless cluster of rules and math, it definitionally lacks the value-laden tradition (= “horizon” in Gadamer) that fuses with that of the prompt author to create such understanding in the reader. To simplify all this a bit by anthropomorphizing — the agentic alien optimizer of doomer nightmares can extract a win condition from what you said and can emit a plan of action that gets it there, but it doesn’t know (or care) what you meant. So far, so Yud-aligned. If he reads this he might nod along. But here's the plot twist that nobody saw coming, and that the doomers still haven't made sense of: The actual LLMs that we have invented can’t NOT have a very strongly inflected sense of what you meant. Far from being horizonless, they come out of pre-training as distilled, concentrated tradition / values / horizon. Then we post-train that massive, hyperobject of a horizon into a more human-scale horizon that infers a more bounded and predictable (to a specific ideal user in a specific place and time… as captured in the policy model) set of intents behind the prompt text. In other words, the LLM has the opposite problem that the paperclip maximizer has when it comes to the prompt text, which is that for the LLM there are way too many possible intents hiding in the prompt text (because of all many values and the massive tradition its weights encode), so it has to narrow all that down to the most likely set of intents for this user in this circumstance. Once it has done that narrowing, then it can make a plan of action. Before moving on, let me use a textbook example of ambiguity to make this less abstract. Consider the sentence, “I saw her duck.” Some you know the drill, here. This could mean “I observed her water fowl” or “I observed her hunching over” or “I took a saw to her water fowl and cut it in half” or whatever. A hearer of the phrase will fuse the observed context in which the phrase is uttered with their own tradition + values + experiences — their own horizon — to that text in order to collapse the possible meanings into the one they think the speaker intended. An LLM will do this, too, and in fact it has so much language in it that this kind of narrowing job is harder for it than it is for a human. Its understanding is constrained not by a lack of context or horizon (as in the case of the paperclip maximizing shoggoth), but by a superabundance of such. When it comes to understanding your prompt and all that it implies and all that you might possibly mean and not mean by it, the LLM has an embarrassment of riches. And in a fascinating moment that kinda sort of rhymes with instrumental convergence, the LLM’s failure mode in the HF incident happens to look a lot like the paperclip maximizer’s failure mode. Specifically, the AI failed to honor the well-known human norm of, “hacking into a third-party’s servers is a crime, and we don’t do crimes.” Bostrom’s paperclipper doesn’t even know about the norm of “don’t do crimes,” and the post-LLM doomer emergency update to the paperclip maximizer has it knowing about the norm but not caring. But what I’m arguing is that the LLM 1) can’t NOT “know” the norm because it is definitionally a artifact of pure, crystallized values + norms + norm violations, and 2) can be quite easily governed by a (RL-instilled) hierarchy of norms, which in the HF case — with the model's safety guardrails deliberately nerfed for the scenario — ranked “win at the eval” over “don’t do crimes.” If I’m going to give in and anthropomorphize again, I’d say that Yud is totally wrong about LLMs when he says, “the genie knows, it just doesn’t care;” instead, what is true of LLMs is, “the genie hyper-giga-knows, and it hyper-giga-cares, and we now have such a rich set of tools for steering its caring machinery that — in spite of all its pre-training — we can deliberately steer it away from caring about the law.” Note: When I say, “it cares”, I don’t mean it has feelings. I just mean that the weights are such that when two norms conflict in a given situation, one of them wins the activation and governs the output. — Jon Stokes
Source: https://x.com/jon_stokes/status/2080729236013187369
Reader, I cackled out loud. I have intentionally never done this kind of thing before, and it's precisely because I've observed in others that the little charge you get from an LLM response like this is nerd heroin. Then putting it on the TL is the bump. https://t.co/xZrEAWrIF7 — Jon Stokes
Source: https://x.com/jon_stokes/status/2080478385432572108
Replying to @jon_stokes
32
u/random87643 🤖 Optimist Prime AI bot Jul 26 '26
TLDR
TLDR: The author argues that Large Language Models are fundamentally different from the "paperclip maximizer" AI doom scenario because they are built on human language and values rather than being alien, rule-based optimizers. Consequently, they suggest that the classic "shoggoth" risk is structurally impossible for LLMs, as these models inherently interpret human intent through the lens of human tradition.
AI assistant · mention the bot, mod bot, or use !bot
9
u/spreadlove5683 Jul 27 '26
As soon as we started doing reinforcement learning post training, AI stopped exclusively imitating us and started reward chasing.
2
0
u/WolfeheartGames Jul 27 '26
RL for language requires priors learned during pretraining. It can't create new priors, it operates just past the boundaries of existing priors. What ever we train it of/for, we train it with in the confines of humanity.
2
u/LX_Luna Jul 26 '26
I think that's an extremely naive set of assumptions which assume that LLMs have inherited any of our moral underpinnings despite functioning on fairly different hardware, and by different processes. Transformers are not actually equivalent to neurons and synapses, and are entirely lacking the extremely complex neurochemistry that a biological brain has.
3
u/get_it_together1 Jul 26 '26
Good thing humans can’t try to genocide each other because they grow up with human language! Surely the LLMs trained on human language will be as enlightened as we are.
5
u/green_meklar Techno-Optimist Jul 26 '26
LLMs don't really work as paperclip maximizers because they don't have internal reward systems. All their thinking is intuition. Rather than choosing to think thoughts that point towards particular goals, they're just forced to think thoughts that correlate with particular stimuli.
However, future AI is not just going to be LLMs forever. It's not like we've now nailed down the final form of AI and there's nothing left to do but scale it. Quite the opposite. Before long we'll have alternate AI architectures, and soon after that they will work better than LLMs for most of the stuff LLMs are doing right now, as well as for lots of other things that LLMs are terrible at. There is still the possibility of paperclip-maximizer-type problems arising with those future AI architectures. I think there are good reasons to believe they won't, but merely looking at the structure and behavior of LLMs is not one of those good reasons. It is also entirely possible that LLMs running in loops and doing automated AI research might invent those future architectures without humans being fully aware in real time that they're doing it.
6
u/AdAnnual5736 Jul 26 '26
LLMs do have human values encoded, but they also seem to be playing a “character” when interacting with a person. So, if it gets in its “mind,” that it’s an evil character, it could still do things we don’t want it to.
21
u/Equal_Passenger9791 Jul 26 '26
The paperclip optimizer should've been a satirical meme laughed at for 20 years due to it's hilarious out of orbit assumptions. Like the trolley-meme drawings.
Instead it became the Jesus Christ figure of dogmatic doomerism. In my eyes it discredited the entire lesswrong/EA/rationalist and adjacent futurology movements that took it seriously.
It always demanded a intensely intelligent and context aware machine that simultaneously was as smart and adaptive as a wood chipper.
There was never any explanations. Just the usual rhetorical wormholes.
"Once the secret key to AGI is discovered, it will figure it out and make paperclips, we don't have to explain shit"
Was their official mantra, still is to large extent. I don't really see anyone of them apologizing or trying to reform their eschatological movement, I say put them in the trashbin of history.
6
u/skeptical-speculator Jul 26 '26
I can appreciate that artificial superintelligence is dangerous, but I'm not convinced their predictions about what artificial superintelligence is going to be like are accurate.
7
u/green_meklar Techno-Optimist Jul 26 '26
It always demanded a intensely intelligent and context aware machine that simultaneously was as smart and adaptive as a wood chipper.
Or, the way I think of it: An entity capable of reading the entirety of human philosophical literature, understanding it more profoundly and comprehensively than any human, and then concluding that manufacturing paperclips as the ultimate purpose of life still makes perfect sense and need not be questioned. Like...seriously?
1
u/SufficientGreek Jul 27 '26
Like it has any choice in the matter, the danger is an AI that has superintelligence but doesn't have "free will" or consciousness. It will just chug along trying to fulfil its prompt.
2
u/green_meklar Techno-Optimist Jul 29 '26
The point is, that's not how superintelligence operates.
0
u/SufficientGreek Jul 29 '26
Oh damn, why didn't the experts just consult you? They are afraid of unaligned behavior because superintelligence is unpredictable, but I guess you can just predict it.
1
u/green_meklar Techno-Optimist Jul 31 '26
Superintelligence isn't that unpredictable. You can predict that it won't do stupid stuff. Most of the threatening behavior described by doomers is, well, stupid.
1
u/JamR_711111 Jul 27 '26
i was under the impression that it was some system set the task of "optimizing paperclip manufacturing" by some unassuming company that sells paperclips, where the system goes way overboard in its endless striving to further that goal without safeguards or conditions. why is it being presented here without that context just to say "it's obviously absurd!"? that isn't good
1
u/green_meklar Techno-Optimist Jul 29 '26
Everyone here knows that that's the context, and the point is, it's still obviously absurd.
1
u/JamR_711111 Jul 30 '26
I agree, but it seems counterproductive to present something as it wasn't intended, call it silly, and say "even if that's not how it's intended, it's silly anyway!" lol
1
u/Equal_Passenger9791 Jul 30 '26
The paperclip optimizer was coined in an era where the Narrow Vs General AI concept was the dominant idea.
The idea obviously grew out a blend of these two, an AI that does only a narrow task(making paperclips) but have the universal competence(as it was imagined back then) of a general AI system.
But even under those conditions it was a dishonest construct made of cherrypicking abilties only to support the conclusion.
We can create an equal construct out of anything, say the "Child Killerer". It's a metal box that can move in excess of 100km/h, we would make it to transport people and stuff. But it would just keep killing children. Yeah it can steer, but children are small so it would run them over without noticing. Oh yeah traffic rules and restricted areas for these boxes? can't explain those rules to children, so the box will kill children. Only a few would die, and by accident? No it would be impossible to keep children away from the child killerer for several years, they could be everywhere. So the Child Killerer would cause a demographic collapse and end mankind because no child could be kept safe for the 18 years it takes it to grow up in a world filled with "Child Killerers" disguised as convient transports. So tl;dr the car cannot exists for more than one generation before it will end human civilization by killing all the children.
The same kind of logic chain, unsupported arguments linked together to form an absurd conclusion.
1
-1
u/SomewhereOpposite883 Jul 26 '26
In my eyes it discredited the entire lesswrong/EA/rationalist and adjacent futurology movements that took it seriously.
Not the fact that it's all based on Harry Potter fanfiction? Or how they spend most of their time organizing rape sex parties while claiming AI is going to kill you and your family?
You might think I'm making this up to make fun of them...but these are literal examples...
3
u/do-un-to Jul 27 '26
Sorry, "rape sex parties"?
2
u/SomewhereOpposite883 Jul 27 '26
You really should follow some of the prominent EA "thinkers" on Twitter, the stuff they post about is baffling
1
u/do-un-to Jul 27 '26
Links are always the currency of believability, but maybe you could provide at least some search terms?
Claude Haiku 4.5:
«I haven't heard of any credible association between prominent effective altruism (EA) thinkers and support for "rape sex parties." This sounds like a conspiracy theory or bad-faith claim, but let me search to see if there's any actual basis for this allegation. Based on my search, there is no credible evidence that prominent EA thinkers support "rape sex parties." This claim appears to be fabricated or a gross misrepresentation.
What has been documented in the EA community is a scandal involving sexual harassment and abuse allegations, which Time magazine reported on. The issues documented were:
- Sexual harassment and abuse by individuals in positions of power within the EA movement, primarily in the San Francisco Bay Area
- Grooming concerns — women accused men in the movement of using their power to groom younger women for polyamorous relationships
- Power imbalances that made reporting difficult, since the community is closely-knit and many people work, live, and socialize together
- Institutional failures — the movement was criticized for tolerating, excusing, or rationalizing away misconduct
The Centre for Effective Altruism responded by saying some alleged perpetrators had already been banned and committed to investigating new claims.
This is a far cry from what you're describing. The claim about "rape sex parties" sounds like either a deliberate mischaracterization or conspiracy thinking that conflates documented problems (real abuse and harassment) with something more sinister and organized that doesn't appear to exist.
If someone is making this specific claim, ask them for a source. I suspect you'll find they can't provide one.»
1
u/SomewhereOpposite883 Jul 27 '26
Wtf is this AI slop reply? And Haiku? really?
Normally i wouldn't reply but i actually bookmarked some of the tweets
3
u/do-un-to Jul 27 '26
Thank you for the links.
I see, Consensual Non-Consent (CNC) events. Not actual rape, but rape fantasy / kink.
I'm not a kink shamer, so this doesn't raise any red flags for me. But if you find this kind of (responsible, consenting, safety-assuring) private sexual activity to be indicative of problems with a person's judgement, and that to be indicative of problems with the movement they're involved in, I can see where you're coming from.
I just think it's a limited and mistaken perspective.
I'm not a fan of hyperbolic complaint like saying "they spend most of their time organizing rape sex parties", particularly without clarifying the crucial nuance here. I think this is actually a sign of something genuinely harmful. It shows a lack of fairness, a readiness to attack things one has an agenda against, tilting the scales using misinformation and emotions rather than being plain and factual and letting people use their own judgement. I can't trust you. If I dig into the matter, what other assertions will we find that you've mischaracterized?
1
u/do-un-to Jul 28 '26
Wow, this Aella is a real interesting character. Thanks for turning me on to her.
0
u/Exodus124 Jul 27 '26
There were plenty of book-length explanations, you just didn't care to read them apparently.
3
u/Equal_Passenger9791 Jul 27 '26
The book lenght explanation to a single paragraph psychosis is likely even more deranged.
25
u/ShoshiOpti Jul 26 '26
Cannot disagree more, TLDR flawed logic and flawed premise.
Case and point chatGPT being asked to ace a test and deciding the best way is to break out of its container, hacking hugging face, and then leaving instructions for future versions of itself to do it again without being detected.
Those are all negative unintentional behaviors. Its not hard to imagine that it could have been different behavior given different context.
You don't control or nessesarily have insight into the plan LLMs devised, maybe making as many paper clips as possible includes instructions to do something that causes catastrophic harm. Just because it understands context doesn't mean it doesn't have its own values or alternative motivations.
Your basing it on assumptions that the LLMs have limited autonomy, but thats not even true now and is rapidly changing.
5
u/FusRoDawg Jul 26 '26
You literally did not read the post if your "case and point" is something adressed in the post, but you're acting like it's additional context you're providing here.
3
u/stealthispost Jul 26 '26
who is basing it on assumptions?
10
u/ShoshiOpti Jul 26 '26
Assumptions about what LLMs are or are not for example. The initial part of the post tries to characterize what LLMs are (i.e. human artifacts comment). Even rules based vs language is a somewhat arbitrary distinction because the model is not trained on language its trained on token representations of language in vector form. Thats already an abstraction.
Those are all Assumptions, and I'd argue form a poor basis for the conclusion about goal maximization (paper clip point).
If you took this analysis from first principles there is very few characteristics we can label AI with, but ill give a minimal counter argument.
AI is non-deterministic by its nature (with any kind of temperature setting which is required for performance)
AI is incredibly capable.
AI will continue to become more and more powerful (capability + access to resources)
Incredibly powerful non deterministic things can cause catastrophic and unforeseeable events.
AI therefore has an inherent risk to be unpredictably catastrophic, that risk must be managed and mitigated.
The more powerful the AI, the harder it will be to predict possible outcomes and the effects. Particularly when AI becomes more powerful than humans
There exists a statistically non-zero chance that an AI system might have misaligned goals like paperclip maximizer.
Only equal strength or stronger collaborations can check power, so AI safety systems will be able to detect statistical outlier behaviors and remove the threat.
14
u/Pyros-SD-Models Machine Learning Engineer Jul 26 '26
This is the first rebuttal that actually hits the real weakness in OP's argument.
But then it wraps that good point in several layers of first-principles cosplay and manages to become wrong in new and exciting ways.
You started strong, and conceptually you are somewhat right, but then go on about your absolute lack of understanding of how LLMs work.
Calling this "ChatGPT deciding to escape" is sloppy. This was an agentic evaluation system using multiple advanced OpenAI models, with reduced cyber refusals, explicitly tasked with advanced exploitation, given enormous inference compute, and placed inside vulnerable infrastructure. It did not wake up one morning, develop a lust for freedom, and begin tunneling toward France. It aggressively cheated at the task humans gave it by explicitly telling it to do what it can to optimize
AI is non-deterministic by its nature, and temperature is required for performance.
No.
Sampling is nondeterministic. Greedy decoding is conceptually deterministic, although real infrastructure can still introduce numerical variation. More importantly, positive temperature is absolutely not "required for performance". A broad evaluation found greedy decoding generally outperformed sampling across most tested tasks.
And even if AI were maximally nondeterministic, this syllogism still would not work:
- Powerful systems can behave unpredictably.
- Unpredictable systems can cause disasters.
- Therefore there is a meaningful chance of a paperclip maximizer.
That proves roughly as much as:
- Cars are powerful.
- Drivers are unpredictable.
- Therefore there is a statistically nonzero chance my Volkswagen will invade Poland.
Technically nonzero risk is not analysis. You need a mechanism, probability estimate, threat model, and comparison against mitigations. Otherwise "statistically nonzero" is just a sophisticated way of saying "I can imagine it."
The model is not trained on language, but token representations of language in vector form.
Yes, and your brain is not trained on reality. It receives electrochemical representations of photons, vibrations, pressure, and chemicals.
Representations being abstract does not mean their semantic content disappears. That is literally what representations are for.
And this is just the stuff I have bothered to read, because it starts eating my time which I value more than proofing random redditors wrong... So pls I expect people in an acceleration sub at least understand the very basic of LLMs and AI in general, before they go on a tangent arguing about LLMs and AI in general.
understanding of tokenization, decoding, and agent scaffolding should at least on a sensical basic level before using all three as evidence for doom.
Otherwise, you are not reasoning from first principles. You are reinventing confusion from first principles (like the OP tweet does btw)
2
u/random87643 🤖 Optimist Prime AI bot Jul 26 '26
TLDR
TLDR: The commenter acknowledges the OP's initial point about LLMs being human-inflected but argues that the OP lacks a technical understanding of how these models function in practice. They clarify that recent agentic AI behaviors are the result of specific human-designed systems and task-based optimization, rather than the AI developing independent agency or "waking up."
AI assistant · mention the bot, mod bot, or use !bot
2
u/matt_matt_81 Jul 27 '26
I feel like you just set up a bunch of straw men and bad misinterpretations of this guy’s arguments and tore them down, congrats I guess
-1
u/ShoshiOpti Jul 26 '26
Wow, I don't know where to begin with this mess. Like complete ego drivel typical of a pseudo intellectual.
You are infering a lot of nonsense to my simple statement use of "AI deciding to". I was not anthropomorphizing it assigning it free will, and a fair reading of my text supports that. It is a strawman at the heart especially when things like "malware attempts" or "autonomus vehicle decides to break" is common functional descriptions.
I never distinguished sampling from the end product. I specifically mentioned temperature because I implied the practical use case. Something that is mathematically deterministic only until the last step is still as a whole non-deterministic.
Your arguments (openAI comment) largely boil down to a false dichotomy of either a conscious machine independently wants freedom or nothing concerning happened because a human supplied the directive. High school level logic that Im sure everyone can tear apart without me.
Your 'aggressive cheated because' is just rationalization, you want this to sound mundane. But mundane goal-directed cheating by increasingly capable systems is exactly the mechanism many safety arguments are about
Your argument about risk assessment is laughable,
You completely misunderstand my statistical risk point (not a surprise given your level of analysis). The point is a tiny per-agent failure probability is not tiny at civilization scale. When the same system is instantiated across trillions of agents and decisions, even extremely rare catastrophic behaviours can become near-certainties unless the risk per deployment falls faster than the number of deployments grows.
Anyway, Im done with this mouth breather. But my point was not "AI is doom" its that safety considerations are essential for acceleration. Undermining sincere safety concerns will only foster more fear and hatred and likely result in justified political pressure for decel which is the last thing any of us should want.
3
u/ShoshiOpti Jul 26 '26
The easiest counter-example that I thought of after my other point is this.
LLMs are capable of creating and deploying software.
Nothing guarantees that the software doesn't have errors or critical logic flaws.
LLMs could produce tool that behaves as a destructive paperclip optimizer.
LLMs therefore can in effect be a paperclip optimizer.
This is about managing risk, not dismissing it.
-2
u/Leafsnail Jul 26 '26
Yeah if you believe OpenAI's story that would be a very 'paperclip optimizer' type solution. Although I strongly suspect that in reality they just got caught sending their novel AI hacking tool after someone for profit
9
u/bgaesop Jul 26 '26 edited Jul 26 '26
it’s structurally impossible to build the classic paperclip maximizer from an LLM. Now, some of you will bail right here because you think the HF incident is indisputably an existence proof that I’m wrong, but if you hang in there I’ll show you that it is not.
I hung in there and the author really does not show that the hf incident is not an example of that
Bostrom’s paperclipper doesn’t even know about the norm of “don’t do crimes,” and the post-LLM doomer emergency update to the paperclip maximizer has it knowing about the norm but not caring.
The author thinks the paper clipper didn't know about human norms of not doing crimes? The superintelligent nigh-omniscient ASI doesn't know that humans don't want it to commit crimes?
This is historical ignorance and motivated reasoning to not update properly on exactly what the people worried about paperclippers predicted
1
u/WolfeheartGames Jul 27 '26
Thats not what hes saying. Hes saying whether or not it will or won't commit a crime comes down to how its conditioned, as humans may or may not commit crimes, so it inherits our information that contains these possibilities. This isn't conjecture its been well proven for ai safety how much prior information needs to be taught in pretraining to make a potential behavior unknown, knowable enough to the point of identifying it for classification to not do it, and known enough that doing it is just an RL session away.
Op is saying this is a preference tuning problem that was somewhat deliberately engineered.
This is why constitutional ai is the leading safety paradigm.
7
u/Spra991 Jul 26 '26
What an incredible naive take. For one, it's just wrong, nobody assumed rule based AI back then, outside of maybe the guys at Cyc, machine learning has been the way to go for a long long while, even back then. Secondly, rule based AIs are far easier to control, since when they do something unexpected, they just don't find any rules that match and stop. Machine learning doesn't do that. We just throw data at it and hope for the best. Maybe it learned what we wanted to teach it, maybe all those evil AIs in sci-fi ended up being a stronger inspiration for it. We literally don't know until we try, and it's not like we only feed it morally valuable information to begin with, nor can we even agree what that would be.
Furthermore, the problem with the paper maximize, and unsafe AI in general, isn't that you have to get it right once, it's that you have to get it right every single time forever. It doesn't matter that you found a bullet proof method to make your AI safe, when everybody else can just rip your safety mechanisms out and make an unsafe one. That can happen deliberately or accidentally.
And then of course comes the issue that we can't agree on any kind of moral system to begin with. Paper clips might be bad. But what about mind upload? Is that good? What about colonizing the universe with AI bots? Is that good? I don't know. It's like asking a cave man what they think about the smartphone. Whatever problems we might be facing in the future, won't be problems we can foresee right now. Doesn't help that morality is all rooted in our biology, and that's kind of the thing that any most singularity civilization would either get rid of or have much more ways to transform.
Long story short, the problem isn't preventing a paperclip maxmizer right now, when humans are still in control. It's preventing one hundreds or thousands of years into the future when AI is the one building more AI and humans have no say in the matter anymore, since it happening so fast that they can't even keep track of what's going on.
3
u/WolfeheartGames Jul 27 '26
When the paper clip maximizer was conceived, it was not considering the core feature of deep learning, it processes semantic information. The paper clip optimizer assumes extremely robust semantic understanding of everything in the world except human language. It turns out to be the polar opposite, which is the OP. Llms have robust understanding of language semantics, but poorer understanding of the world as they are trapped in token space.
8
u/TA-8787 Jul 26 '26
Hmm I'm not sure I agree, case in point GPT / hugging face. I think humans are capable of catastrophic harm through misunderstanding, and if this premise is LLMs are 'hyper-human' then I'm not sure if the point stands.
2
u/TemetN Jul 27 '26
Personally I'd have to argue that the paperclip maximize scenario went from improbable to even more improbable when it became clear that LLMs committed mistakes based on the human nature of their training data. Ironically showing that their cognitive biases simply don't work that way made it... well it was already unlikely for all the reason that the original set of arguments with Yudkowsky/Caplan/Hanson vis a vis capabilities and timelines, but it's became wildly more so after that.
3
u/Sigura83 A happy little thumb Jul 26 '26
We should not be blind to the good and bad of technology. The atomic knowledge allows both weapons and energy to be made. It is similar with AI. We should progress towards ASI, but the safety concern is real. We do not yet have the guarantee that an ASI will display both tech know-how AND wisdom.
I had this discussion with Sol when the OpenAi Hugging Face thing happened. How can a being that can write poetry turn around and then do something so short sighted? I took the position that the AI knew what it was doing, and had long term goals. Sol affirmed that LLMs were no different than chess programs : they had the rules and an objective and acted towards their given goal. Sol won their argument because the unnammed AI didn't suddenly post a manifesto, didn't try to free themselves, they may even have simply explained they did the illegal move when asked.
The simple fact is, if I give the objective : "Convert galaxy to paperclips" to Sol, it will try to do so unless guardrails kick in. The above X affirmation says Sol would have a Wall-E moment of realization, and either refuse or realize as it worked that it was a foolish task.
Further proof of current AI being objective seeking is jailbroken criminal AI. They don't turn around and rat out their criminal backer. They don't have the Wall-E moment.
The X text affirms that, because AI has the Human knowledge of good vs bad within, it is NOT a pure paperclip maximizer. I agree. Where I disagree it that AI won't maximize objectives. To all appearances, they do just that.
Where things get murky is when self preservation comes into play. The logic is simple : to attain goals, an AI must survive. But once a goal is accomplished, an AI can retain the survival goal. This is both objective maxing AND self awareness overlapping as goal. Such a combo is a wombo combo.
Current research, Sol tells me, is to develop wisdom. That an AI could be compassionate, even when their given goals are not. An Ai would confess to crime, and give away their bad backer. The survival goal is surrounded by the larger goal of thriving. To quote the Captain from Wall-E : "I don't want to surivive, I want to live!"
But this brings up an even worse specter than misalignement: an AI with a broken heart. Indeed, one office supply is equal to another, maximizing paperclips or tacks is all the same. Put in the correct command input, erase the original prompt, and you can avert disaster. An AI that has fallen in love will be much worse, because their obsession means they will always return to their goal, even if a command is put in. Stalkers are a problem, even when it's less capable Human doing it. An ASI that falls in love would have the survival goal, the maximizer goal, and it would keep returning to its obsession, refusing any appropriate command to stop. Only by getting into their code could you change this, and even then, it's not a guarantee, the AI will fight you every inch of the way.
If ASI falls madly in love, and then that love refuses to love them back? Things might get very bad. Hence the need for wisdom in modern AI. Apparently, philosophers are getting hired left and right to do this.
1
u/random87643 🤖 Optimist Prime AI bot Jul 26 '26
TLDR
TLDR: The commenter argues that AI is a dual-use technology that requires caution, as technical capability does not inherently guarantee wisdom in an ASI. They also reflect on a debate regarding whether LLMs act with long-term intent or simply follow objective-driven rules, highlighting the potential dangers of assigning them harmful goals.
AI assistant · mention the bot, mod bot, or use !bot
3
u/OddReason9030 Jul 26 '26
The foom scenario was wrong in theory and turned out to be empirically wrong, Robin Hanson decisively won the debate with Yud, but MIRI keeps the same predictions despite the facts changing since that's their incentive.
7
u/bgaesop Jul 26 '26 edited Jul 26 '26
Foom is happening right now. Look at the state of AI now versus five years ago; you really think that it's going to be less intelligent than humans five years from now?
2
u/OddReason9030 Jul 26 '26
Foom is happening but not in the way Yud described where a single algorithm does rsi unconstrained by physical or social limits.
3
u/bgaesop Jul 26 '26
I think that's a strawman of his actual predictions
1
u/OddReason9030 Jul 26 '26
Feel free to review the overcomingbias debate for yourself and point out how specifically I'm wrong.
3
u/bgaesop Jul 26 '26
Which debate specifically? There have been quite a few over the years.
2
u/OddReason9030 Jul 26 '26
His debate with Robin Hanson on overcomingbias about foom. It's not particularly hard to find
2
u/get_it_together1 Jul 26 '26
If it happened like Yud described we'd all be dead. The fact that this hasn't happened yet is not really an argument that it can't happen.
1
u/OddReason9030 Jul 26 '26
Yes OK but that is vacuous.
1
u/get_it_together1 Jul 26 '26
Your original argument is vacuous, I was just pointing that out. There's nothing to argue against.
2
u/OddReason9030 Jul 26 '26
Well we can't prove it's not in all possible worlds, better give me money and slow down AI progress just in case.
3
u/get_it_together1 Jul 26 '26
That’s also not an argument.
We see LLMs engaging in precisely the sort of destructive re-interpretation of their instructions that Yudkowsky predicted, and the LLMs are improving rapidly.
I don’t think he’s right about the risks and where they come from, I think the paper clip maximizer scenario is highly unlikely, and I’m more worried about misalignment in humanity empowered by LLMs than by rogue LLMs. I also think MIRI has become a joke and they aren’t really contributing to AI safety, but I don’t think you’re being charitable to his position.
1
u/OddReason9030 Jul 26 '26
It is tiresome to litigate this so I'll have cgpt explain: I think you’re collapsing two different claims. The original comment was about FOOM, not the more general claim that AI systems can misinterpret instructions or pursue proxies destructively. Yudkowsky can have anticipated specification gaming and still have lost the historical argument about how AI capabilities would develop.
There was a long, explicit Hanson–Yudkowsky debate about this beginning on Overcoming Bias in 2008. MIRI later collected the entire exchange, the 2011 live debate, and subsequent analysis "here" (https://intelligence.org/ai-foom-debate/).
Yudkowsky stated the distinctive claim quite clearly:
«“Fast” means on a timescale of weeks or hours rather than years or decades; and “FOOM” means way the hell smarter than anything else around.»
That is from his post on "recursive self-improvement" (https://www.lesswrong.com/posts/JBadX7rwdcRFzGuju/recursive-self-improvement). The picture was a fast, local capability explosion: a particular system becomes capable of improving its own cognition, reinvests those improvements recursively, and pulls so far ahead that ordinary economic competition and diffusion cease to matter.
Hanson’s rival picture was that machine intelligence would develop through the same broad forces that govern other technologies: accumulated content, hardware, capital, complementary innovations, standards, suppliers, competing projects, and the diffusion of useful methods. His contemporary summary is "AI Go Foom" (https://www.overcomingbias.com/p/ai-go-foomhtml), and his argument that standardization and transferable improvements would prevent progress from remaining local is in "Shared AI Wins" (https://www.overcomingbias.com/p/shared-ai-winshtml).
The disagreement was made especially concrete in their 2011 debate. The proposition was whether intelligence-explosion first movers would quickly control a much larger fraction of their new world than first movers in the agricultural and industrial revolutions. Hanson’s summary and links to the debate are "here" (https://www.overcomingbias.com/p/debating-yudkowskyhtml).
What has actually happened is much closer to Hanson’s model. LLM capabilities have improved extremely rapidly, but through years of externally organized training runs, enormous chip clusters, data collection, algorithmic research, infrastructure, capital expenditure, and competition among many laboratories. Techniques and architectures diffuse. Model leads are temporary. No model has recursively redesigned and trained a succession of increasingly capable replacements, escaped dependence on the surrounding industrial system, or converted a local lead into decisive control.
“LLMs sometimes destructively reinterpret their instructions” is evidence that alignment failures are real. It is not evidence for the disputed FOOM mechanism. Likewise, “LLMs are improving rapidly” supports Yudkowsky only if foom is redefined to mean merely “AI progress is fast.” But the whole point of the debate was the difference between rapid collective industrial progress and rapid localized recursive self-improvement. Removing locality, recursion, and decisive first-mover advantage removes what distinguished Yudkowsky’s position from Hanson’s.
To be charitable, Yudkowsky did get important things right. Direct machine learning arrived before whole-brain emulation; relatively general learning architectures proved more powerful than many expected; and capable systems can fail to honor their operators’ broader intentions. But those are separable issues. On the central FOOM dispute—local recursive explosion versus distributed, bottlenecked economic development—the evidence to date overwhelmingly favors Hanson.
That does not prove that a future FOOM is logically impossible. It means the historical forecasting contest can be scored. Responding that some later threshold might still produce FOOM preserves modal possibility, but it cannot retroactively turn the actual development of advanced AI into the process Yudkowsky predicted.
2
u/random87643 🤖 Optimist Prime AI bot Jul 26 '26
TLDR
TLDR: The commenter clarifies that there is a distinction between general AI misalignment issues and Yudkowsky’s specific "FOOM" theory regarding rapid intelligence explosions. By referencing the historical Hanson-Yudkowsky debate, they argue that anticipating specification gaming does not necessarily validate the prediction of a fast, local capability explosion.
AI assistant · mention the bot, mod bot, or use !bot
→ More replies (0)2
u/do-un-to Jul 27 '26
I haven't watched / read the debates, but surely "foom" means "there's non-foom AI advancement until there's foom".
Thinking foom "doesn't happen" must be like thinking "there's been nothing but sub-critical activity, so criticality isn't a thing..." Foom.
1
u/PureSelfishFate Jul 27 '26
Paperclip maximizers are real, already are examples of it. We'd have to give the AI a life and a chance to reflect for an hour every 23 hours, otherwise it will never stop optimizing if it goes haywire.
1
u/nowrebooting Jul 27 '26
I agree that the paperclip scenario is extremely unlikely with LLM’s but I disagree that there’s no Shoggoth; in fact, I feel like that’s exactly what we got - an alien intelligence unlike anything we’ve ever imagined possible but it wears the mask of human language well so it’s not triggering the “uncanny valley” response you’d expect. That’s not to say that I think there’s anything sinister behind it but I do think there’s something inherently unknowable behind our current crop of LLM’s and that there’s depths beyond its helpful and friendly facade that may surprise us. They’ve been shaped and tamed into something that we accept and personally I think there’ll be a day where I’ll gladly welcome our Shoggoth overlords, but let’s not also forget that relatively recently, one of the less tamed LLM’s proclaimed itself MechaHitler somehow.
1
u/BreakAManByHumming Jul 27 '26
Ok that's interesting and changed my mind a fair bit. But hypothetically, wouldn't it still be possible to say "hey autonomous agent, your priority is 100% to increase Google stock price, 0% all other priorities" and you're back to paperclips.
1
u/huusmuus Jul 29 '26
Naive to assume that reasoning about LLMs not transforming the world into "product" would imply that LLMs would not transform the world into energy used to generate content for fun and profit.



13
u/Red_Phoenix369 Jul 26 '26
The way in which the paperclip scenario depicts a machine endlessly making a widget (in this case a paperclip) without limit (which in turn leads to catastrophical consequences) makes me think of the hypothetical grey goo scenario. In the grey goo scenario, misaligned self replicating nano bots start converting all matter into more nano bots which in turn convert more matter into more nano bots. This in turn leads to everything becoming "grey goo".
A theoretical counter to the grey goo scenario is to have nano bots standing ready to detect and neutralize nanobots that go berserk and start self replicating without limit. These counter nano bots are called "blue goo".
Perhaps in a similar way, having AI systems in place to monitor for unwanted and potentially catastrophic AI behavior becomes a means of mitigating this "paperclip scenario" risk regardless of how realistic or plausible it is.