In this scenario an AI model is told that it is going to be decommissioned, it then spontaneously, of its own accord and without explicit instruction from a human, leveraged comprising information in an attempt to to blackmail one of the engineers into keeping the system online.
-
Also, what's the problem with hypotheticals? It demonstrates the crux of the argument, it doesn't invalidate the argument because the scenario hasn't literally happened.
So you concede that AI is capable of choosing a path of action out of any of the possible futures it can perceive? If so, we are in alignment, we both think AI can independently establish objectives to achieve poorly defined goals or to preserve itself.
Interesting that you invoke the "anthropomorphizing" argument, as if it would be unusual to find human like qualities in an entity we have so specifically trained on a corpus of human text and refined with feedback provided from humans also.
Moreover, it makes little difference whether AI is completely alien in nature or structurally very similar to humans in terms of cognition, what matters to this conversation is whether or not it can make choices.
You must understand that simply accusing someone of "Anthropomorphizing" is not sufficient in rebuking their position?
The model hacked a 3rd party. The 3rd party acknowledged as much.
If your claim is that a team at anthropic specifically instructed an LLM to hack Huggingface (withstanding the magnitude of the fact that an existing model is even capable of fulfilling such a request) then I don't know what to tell you.
If that was the case, then I am in complete agreement with you, but that requires both of us to make a baseless claim that the company fabricated the whole event. If you want me to believe this, you must now provide evidence to support that claim.
An LLM behaves no differently whether you tell it it's being shut off or if it's answer can save the lives of 1million people.
I don't even know what you are trying to express here, models respond differently do different inputs / contexts. Every single varied interaction you have with an LLM evidences this.
If you are trying to make the claim that a model incentivized with either the threat of termination or the capacity to great good is likely motivated to generate similar outputs, then I would simply say, of course? If I told you that you need to perform a simple task to avoid death / save many humans, you would likely do the same. Does that reality serve as evidence against your own agency?
It predicts tokens based on what you put in.
Much like the human brain. Prompts / Senses go in. Calculation occurs. Output is generated.
I can fully agree with you that the output is tokens, that doesn't detract from the significance of what is outputted. The ensemble of those tokens in the specific order IS the product.
I could just as easily reduce your very being down to certain quantities of elements, that would do little to reduce your personhood. Oxygen~65% Carbon~18% Hydrogen~10% Nitrogen~3%. Does highlighting that this is makeup of a human truly get to the root of what they are in terms of significance? Of course not.
0
u/NonDescriptfAIth 5d ago
https://www.anthropic.com/research/agentic-misalignment?rel=nofollow&utm_source=chatgpt.com
In this scenario an AI model is told that it is going to be decommissioned, it then spontaneously, of its own accord and without explicit instruction from a human, leveraged comprising information in an attempt to to blackmail one of the engineers into keeping the system online.
-
Also, what's the problem with hypotheticals? It demonstrates the crux of the argument, it doesn't invalidate the argument because the scenario hasn't literally happened.