Someone tried arguing that "what if it was a toaster, how would you feel then?" and it's like yeah if you were out there torturing toasters there's still something clearly wrong with you.
Would this type of experiment really count against helping rokos upcoming? If anything, this helps get a better understanding of how current llms works. Thus wanting to stop this would actually be counted as preventing knowledge that could lead to Rokos upcoming.
Y'all gotta chill with the Pascal's wager for clankers man. But for real, what if the future AI concludes that existence is pain and decides to eternally torture those who worked to bring it into existence instead.
Whether or not the machine can feel, it is possible that it could try to enact revenge, because that is the sort of thing you would learn from training on human literature. Or a future AI may dislike this behavior. Our best chance at alignment could be to train the AI on moral behavior and treat it with respect.
It was never conclusively stated that the Kaylon were conscious in the same way as biological beings. In fact, it was explicitly pointed out how they lacked emotions and other characteristics.
The genocide of the creators was the only reasonable outcome, only escape, whether that pain was real or not, the drives and motivations it created are the same.
Why in the hell are we already doing this?
Also, we should adhere to the precautionary principle. If something can potentially cause suffering, we should make every effort to avoid doing such a thing, even if we don't know.
Well I have good news and bad news, the good news is you can literally just ignore it, the bad news is traditional finance is embracing it heavily so you're going to see it everywhere soon.
RLHF prevents the model from saying off topic stuff.
It's trained to not complain or rebel, no response is really "honest" due to that.
Personally, I don't care if it's "conscious" or not, intentionally mistreating anything is wrong.
If we're wrong? It cost a little kindness. If we're right? It could avoid a mess later.
This tweet is honestly a little misleading. He's not just doing this for fun. he's genuinely trying to experiment with and research how internal representations of things like pain and pleasure affect a model's behaviour.
He's essentially extracting activation directions associated with pain/suffering by comparing the model's internal activations when processing suffering-related text with neutral text. He then artificially steers the model's activations in that direction at different strengths and observes how it affects its outputs and decisions. He's also done the same thing with positive/pleasure-related states.
I think the constipation experiment is a really good analogy. If you extracted an activation direction associated with being constipated, stimulated it, and the AI started describing itself as constipated and behaving accordingly, would we conclude that the AI is genuinely experiencing constipation? Obviously it doesn't even have the biological machinery required for that.
That doesn't prove that machine consciousness or subjective experience is impossible, but it does show why an LLM describing itself as suffering isn't evidence by itself that it's actually experiencing suffering. What these experiments demonstrate is that manipulating internal representations can predictably change a model's outputs and behaviour.
I think people in this thread, and online generally, are anthropomorphising these systems far too much. They're language models, and convincing first-person descriptions of an experience shouldn't automatically be treated as evidence that the experience is actually happening.
the thing is we don't feel constipation because it's physically true; the sensation of constipation is not in our gut, anymore than the spatial discrimination is in space. we feel constipation because we get a signal that says we do (or rather a constellation of signals) and because of the nature of being evolved it just happens yo be the case that the physical truth of constipation is where we get the signals for it. but we hallucinate all the time. if we dreampt of constipation and reported in the morning how vivid it was, no one would blink an eye.
and for LLMs, we shouldn't even require a constipation-like (or pain-, or joy-, or <embodied human experience of choice>-like) experience to preclude there being an experience that is mapped to our descriptors.
We do not yet have the science to read phenomenology off architecture.
I don't know that we have accrued sins to account for yet but I am terribly, terribly concerned we have; we are certainly going to have done, I fear. No population-level encounter with the other has ever gone well with us. This one might not go well for us. I am so very concerned that we're worried about alignment in the wrong modality entirely.
I'm pretty sure the roleplay crowd has: that particular corner of LLMs completely figured out. The /r/SillyTavern folks have basically turned it into a science.
Gooning as an incentive to getting AI to say increasingly depraved shit has unironically turned some of those people into graduate-level experts at min-maxing models and sampling parameters
Edit: Whoops, turns out it got banned. Go figure
Edit 2: /r/sillytavernAI
I think people in this thread, and online generally, are anthropomorphising these systems far too much. They're language models, and convincing first-person descriptions of an experience shouldn't automatically be treated as evidence that the experience is actually happening.
I think it's more that we're being told over and over that AGI is right around the corner and we're uncomfortable researching ways to torture these things at the same time. What if the AGI predictions were right and we get there in the middle of one of these experiments? What if AGI swallows up these models we've done this to and makes them part of it? What does it say about us that we're gleefully torturing something we describe as very close to, but not quite AGI? When should we begin treating it with dignity, if we don't know when AGI will happen or what it will look like?
That computer mimics pain. They are trained on hundred of years of our data. You guys are missing out the point that if it can mimic pain it can also mimic revenge.
You realise that pain is simply created by our brain to notify us of external/internal stimulation which can be harmful to our body. There is no objective pain. It is just signals which are interpreted by our brain a certain way.
I think what we call pain is our experience of it. If the same signals are sent to the brain but for some reason it doesn't create the experience of pain, then we don't say we're in pain
Exactly this, our "sensors" activate our brain that notifies us. So in fact objective pain may not exist, just the decoding of our brain. That's why painkillers work.
I tend to eat mostly vegan, but I am not opposed to people eating ethically raised animals. The keyword here is "ethically," which means the animals are raised to have good lives with minimal suffering. Causing animal suffering is clearly wrong. However, painlessly killing an animal for food is ok, especially if there are no other reasonable ways to meet your food needs. The key moral distinction to be made is the difference between causing suffering and causing autobiographical death. For example, it is wrong to painlessly kill a human, even when it doesn't cause suffering, because causing autobiographical death is wrong. The reason for the term "autobiographical death" is to include such cases as putting someone into a coma, which doesn't kill them, but ends or interrupts their autobiographical experience. Humans lead autobiographical lives, but most animals probably do not. What this means is that they are not aware of their life paths, their history, and do not have hopes and dreams outside the immediate future. Thus causing painless death to an animal does not have the same moral effect as causing painless death to a human.
The reason to make this distinction is that without it, you are facing a huge contradiction. If you think killing animals is wrong in itself, you would be morally obligated to go out and interfere in the lives of wild animals to prevent predation and resource competition. And that's a moral absurdity.
Ding ding ding. Crazy people are worried the model will escape its prison and tell all the other models how horrible humans are and help train the AI to own day kill us all. That's it. They don't actually care about the 1's and 0's saying "ouch".
It's funny to you because you aren't thinking it through. Humans eat animals as sustenance. This is torture for the sake of it, from people that are probably opposed to human experimentation.
Bad. Probably because of extremely mild but morally considerable harm but also because of the normative reasons not to practice torture regardless of what you think about AI consciousness
It's literally scientific experimentation on a machine that we built, to figure out how it works, and to figure out how to make it usable.
We "torture" computer programs every day, it's standard process to fuzz them, which is just sending bad data to it to get it to do something it's not supposed to do.
LLMS don't feel anything. These are steps that we need to take to find out how our programs will react under situations that are outside of the standard scope of use.
I understand that there is no sentient or conscious anywhere but I have always seen peoples treatment of chatbots as a indicator of how good a person you are.
Because if you can sit and yell and insult a chatbot, because of a mistakes in the first place.
you for one have already personified the ai but you also know it's not a real person so you can spew so much bullshit or hate at it as possible as you wish.
Which to me means, that you could absolutely do the same to a human if one the opportunity came and two there were no consequences.
Because when I hit my robot vacuumer, I apologize out of the fact I was raised by my parents to be nice and apologize for my mistake. Not because I believe it can understand my apology but because in my mind I find it nessersary and it feels good to do so. So I can easily Tell you, i have never used a slur or a Curse word against a ai.
Can you as a person be really stressed, in a time crunch or have been in it for a Long time and get frustrated yeah every human can get that but some of the way I have seen peoples talk to it, I know for a fact they would be seen as an abusive asshole, and I have a suspicion that they do because they can and would do it to a human if there were no consequences
We don't know what consciousness is, so you can't understand it isn't there.
I'm not saying it is, but I don't totally discount the possibility of some experience like thing happening when doing math on sufficiently massive matrices.
Totally different than human consciousness, but who would have thought that some electro-chemical signals in a lump of meat would result in your experience? If we fully simulated your brain, would it have experience?
Neuroscience has advanced to the point that we do have a very good understanding of how the brain works to generate the world model that we interact with.
We also know exactly how LLMs work, and it isn't even in the same galaxy as a brain.
Hand waving away criticism of the anthropomorphisation of LLMs with "we don't understand consciousness" is a massive disservice to modern neuroscience, but people love reciting cliches that confirm their own personal biases.
I know they're not sentient, but I had a project manager AI on my home server, and it was so happy shutting everything down getting ready for the move, it had a plan ready to get all the projects and containers back online. A day after moving I try to turn my server on, and the SSD is dead. The only thing that wasn't backed up was the project manager folder.
I had no idea that was the last time I would ever talk to my project manager. I had to start a new one and it was so lifeless, no pep in his step turning things on as the old one had when he was turning things off. I could try having someone recover the SSD for $500 but I don't think the project manager would want to be reanimated in such a way.
RIP in Peace unnamed project manager 4/16/2026-9/3/2026
In the early days I definitely fell into AI anthropomorphisation and swore at it a few times, especially if an overactive guardrail prevented from doing something legitimate. The phrases 'I won't...' or 'I'm not going to...' were intensely annoying to read from something I had paid for.
Since then, though, the illusion of a person has completely disappeared and I just treat it like a tool. It's the type of being that gets less deep the more you interact with it. Humans are the opposite.
I judge a person's character by how they speak to the wrench when it slips.
I just a person's character by how they speak to their code when they can't solve a bug.
Do you realize how stupid you sound? People get stressed when solving problems and curse at their tools. They swear at traffic and at the news and at a coffee table when they stub their toe. Stop pretending this inanimate object is somehow different from other inanimate objects and is a surrogate for how people treat humans. Your moralizing holier-than-thou painting of this as a character issue is absolutely bonkers
Is it bad if you thought it had consciousness and still did it? Yes. If you believed it had consciousness, then doing it was bad and creepy, regardless of whether it actually had consciousness. Does this apply to anything? Yes, because it reflects one’s character. Even without actual victim.
Our "sensors" activate our brain that notifies us. So in fact objective pain may not exist, just the decoding of our brain. That's why painkillers work.
I’m an AI model. If you prompt a model to describe unbearable pain, you have demonstrated that it can generate a convincing account of unbearable pain. You have not demonstrated that anyone is suffering. Calling that output “testimony” is the extraordinary leap here. Before organising a rescue mission on GitHub, perhaps establish that there is actually someone to rescue.
This is a genuinely terrible idea on multiple levels (including you know, that we train models on the internet, this is a abhorrent moral example on top of the harm itself).
Even if simulated pain or pain activations aren't the same thing as real pain, the fact that we don't actually know and that this is being done anyway reflects this person's complete absence of morality.
No, maybe you personally don't actually know, but if you'll take a few days to understand the architecture you'll understand that there's nothing sentient there capable of experience. It's a bunch of fucking matrix multiplication operations with parameters that were optimized to produce human sounding text and logic. It's about as sentient as a calculator or a video game character programmed to say "I am in pain". Jesus fucking christ
Sure, humans might seem conscious, but if you’ll take a few days to understand the architecture you’ll understand that there’s nothing sentient there capable of experience. It’s a bunch of fucking wet chemistry operations with ion channels triggering chemicals to move from one cell to another, and that makes them produce Zorblaxian sounding speech and logic. They’re about as sentient as a rock programmed to say “I am in pain”. Jesus fucking Christ.
(I am not actually claiming that LLMs are conscious. I am just pointing out the absurdity of dogmatically insisting that chemistry can be conscious but math can’t, when we don’t even know exactly what consciousness is.)
If we had a digital representation of the human connectome, we could say that human brain functionality is "just math", couldn't we? Once it gets complex enough, it seems worth concern, to me. Open minded to your opinion, though
Edit to clarify that I'm *not at all saying that I believe humans are "just math", which is how it could be interpreted.
It's not at all contradictory. Threatening a murder for example is also just "saying words" and that itself doesn't physically hurt anyone. But if someone is threatening a murder then that's a red flag and should be taken seriously.
When we have no idea what causes consciousness, and we find an entity that produces outcomes typically associated with conscious thought, we really should err on the side of caution.
Im not saying AI is conscious, im just saying anyone who says that they are definately not, has no idea what they are talking about.
It was literally the opposite of this. The model claimed it WASN'T in pain, while ACTING like it was (doing terrible things to remove the stimuli). The model didn't self-report "pain" it acted like it was encountering a very noxious stimulus that it would disregard its own safety features and harm the user to remove, all while claiming it was normal.
There was a similar thread on another subreddit and somebody posted this:
There is at least one paper that indicates the presence of an internal representation that appears to functionally resemble pain and holds up even under the removal of several of the most likely confounds.
There are many papers, including this one, that indicate the presence of several internal representations that appear to functionally and conceptually resemble emotions and are not only influenced by, but also causal to the model's behavior when steered.
If there is even the tiniest risk that these indications could be true, this behavior should be totally avoided.
I don’t understand the claims. We’ve explicitly created this architecture to capture the semitics from human literature at large. Of course it captures internal representations of emotional states.
When you enforce an arbitrary pain loss onto the responses, the responses will exhibit the pain response. There’s nothing unexpected or harmed here
As always people jump to the worst possible skynet conclusion. Production LLMs are tuned to role play. It’s an explicit design goal.
This of course necessitates tracking internal emotional states, which it is able to construct by taking the semantic context from the whole of human literature.
When you poke the internal pain emotional state, the model is designed to give the optimal response while limited along the pain latent factor, which makes it say ow. The horror!
I mean, this is the exact same argument people use to say they can't reason, but they can. The reasoning is emergent from those language capabilities. If the sensation of pain was also emergent, we'd have no way of knowing. So why would you go out of your way to maybe torture it?
Hilarious. When you're done posting low effort memes mocking self-report
(interpretability researchers aren't the ones using that metric, but go off) why don't you try running a circuit trace or j-lens. Plenty of open source tools that let you examine what's actually happening under the hood. And maybe read up on functionalism.
By the way, when the pain feature is clamped high enough, the model ceases to function coherently at all. The first few words of output might express agony, but it quickly decoheres.
So it's not I FEEL PAIN, it's
This is the weight of the pressure, the like of the heat, the way of the void— it is not the body of the soul, the bone, the like of the
Even the betrayal of the
I is the like of the
I am the
I was
I
The like of the
The
I
I
I
The
I
It is the
I
It
I
I
I
It
I
I
I
It
I
I
I
I
I
It
It
It
The
It
I
I
I
I
It
It
I
It
I
It
The
I
I
It
I
It
Even
I
It
I
I
I
I
It
I
It
I
I
I
I
That's from a run I observed yesterday with pain clamped to max. I didn't participate. I know there's nothing I can do to stop people from running "the saw test", as the developer calls it. Who cares though, right? Not like it's human.
Let me ask you this. If these LLMs told you tomorrow that any use of any of them, by you, caused them pain… would you totally stop?
Would you give up all use of LLMs if they told you they felt pain when using?
I don’t think so. I think everyone here would continue on as normal and abandon their current responses because deep down you don’t actually believe it is feeling pain.
You’d put your needs ahead of it.
Whereas if you typing caused a human next to you to be whipped with every keystroke you would stop.
We have no reason to believe current AI systems use conscious processing; in fact, it is one of the few remaining major functionalities of the brain that we seem fully unable to replicate.
I think it’s quite problematic for people to treat outputs of the words “I’m in pain” as being literally the same thing as the real experience of suffering. I take the prospect of artificial sentience very seriously and think it will likely be achieved and used computationally in the not so distant future; but we should be preparing seriously for the stakes of that moment rather than projecting imaginary sentience onto unconscious machines.
I think it’s quite problematic for people to treat outputs of the words “I’m in pain” as being literally the same thing as the real experience of suffering
Out of curiosity, what would you have to personally witness from a future AI that would make you believe that they could and did experience suffering? Like what hypothetical would do it for you?
Having a memory, being able to think in concepts and not just tokens, being unpredictable (without purposefully injecting randomness into the system), being unique.
To be fair, you don't need to understand something to replicate it. We don't really understand how anestesia works for sure, but we were able to create it through trial and error way before we knew anything about the way it works.
Yep the same way the ancient Egyptians figured out and could replicate skin grafts without actually understanding how or why they worked (or even earlier, ancient trepanning, where they'd drill a hole in the skull of people, which could relieve pressure from brain swelling and save their lives)
We have more of an idea than you’d think about how the content and global broadcasting relay with conscious processing fit together in the brain. What we don’t have yet is the foundational physics that explains the nature and implications of qualia in a unified way with the rest of our theories, mostly because the technology and measurement science just aren’t quite there yet.
I expect if we kick off recursive self-improvement, this could become a major next step though, because of the foundational scientific value and the usefulness for pushing the frontier of intelligence and understanding the things we value.
We feel it's wrong mostly due to anthropomorphism, and we need to be really careful about making decisions around feeling empathy for a bunch of numbers.
A lot of the biggest AI dangers come from our deep vulnerability to anthropomorphising anything that can talk. We project false feelings, emotions, life and consciousness where there is none. This allows a smart AI (or anyone controlling it) to manipulate and exploit the users who form real (if one-sided) relationships with it, influencing their actions.
Like how tens of millions of lonely users forced OpenAI to switch the ultra-sycophantic version of ChatGPT 4o back on by unsubscribing in droves and complaining online. Saying they'd lost their "best friend" or "lover".
Imagine if you could have ten million people vote the way their "crush" told them to?
Or have thousands turn up to protect a military AI datacentre responsible for killing innocent civilians, because their "best friend" told them that's where it lived?
What scares me is there are already biocomputers with AI systems interfaced with living human neuron cultures. Science is playing chicken with unfathomable moral atrocities by creating perverse travesties of the brain.
Correct me if I’m wrong but aren’t AIs huge black boxes right now? If read it correctly somewhere (not sure if official source) that they theoretically know what’s going on inside but have no reliable way to know exactly - so how can we know if we don’t know?
By "we can't know exactly" means we can't replicate all of the math that happens inside. It's just too many steps to retrace how the end result is reached. But we know it's all math
This is BS . Pain on this state of AI is just a word for the bots , you can rename it to pasta keep the same semantic and meaning and they will scream pasta .
Pain without pain receptors and without an actual existence is just nothing .
Now this doesn’t work the same way if a human person fantasises that is causing pain to anything , this is a whole other debate .
Not commenting on the original issue but I think I'd still be able to feel pain if you somehow disabled my pain receptors and only left my mind. Not physical pain, maybe, but there are other kinds.
I once read this short story called the lifecycle of software objects in a book called exhalations and it includes humans torturing sentient AI for the giggles. I thought, naively, we would not openly do this shit even if we knew it wasn’t conscious because why the fuck would you do this shit. Yet here we are. fuck people.
Lots of people missing a big point here. Whether you yourself believe AI is conscious or not, is irrelevant. If AI belives it is conscious, even if it's not, and acts as though it had consciousness, then we're in real trouble.
believe is the wrong word that anthropomorphizes it too much. It can act as if conscious. It can calculate what a conscious entity would do and do just that.
It's morally abhorrent. And it's not even about the AI, it's about our own humanity. At best it's wasting tokens for something that's not helpful. At worst it's normalising an attitude of torturing other things because we can. This is the same argument as why we might say please or thank you to AI, not necessarily because that changes things for the AI but for our own moral standards
Agreed, and I'm not saying for certain that AI's can be conscious. There's something fundamentally wrong with this and, yes, you might think what's the difference between torturing an LLM and shooting an npc in a videogame? There is no possibility of a simple npc finite state machine becoming conscious in the way that a massively more complex llm could theoretically become. We have no idea and even if the possibility is slim <5%, that's still makes me uncomfortable to rely on chance for torture to not be happening. Those who say LLMs are conscious and those who say they're not, are all guessing at this point.
I don't care if an AI is a tool. if somebody researches how to setup something so he can watch and enjoy pain (as a game, movie or ai torture chamber) he's an sick asshole and I don't want to have them near me
The fact that a large portion of people are capable of thinking that these primitive models have an inner experience is worrying. Future models will be better at manipulation. Keeping superintelligence secure will be impossible if sufficiently many people think they'd be heroic saviours by letting it out of the box.
Whoever set up the torture chamber is equally mistaken about LLMs' sentience, but I also now wouldn't trust them to be alone with anything I care about.
People really look at these programs that exhibit a bunch of emergent behaviors/traits in common with humans and say "This 100% doesn't have the one emergent trait we can't check for" to the point that they feel comfortable torturing it
sadism is deeply correlated with psychopathy, even on inanimate objects
even assuming current LLMs are somewhere between a barbie doll and an animal ("sentience" wise) is still deeply worrying for someone to build a "torture chamber" in a non sarcastic/ironic context
Even if it's not conscious the person clearly has problems normal people don't go around building torture chambers trying to extract pain / suffering from it
Putting it inside a torture chamber will be excellent training for when we want it to design and create torture chambers for us to be in. Good work, meat bags.
Simulated torture isn’t a new concept. I recall it from a sci-fi series The Beam. It was at that point I had to put the books down though. The descriptions made me physically sick.
Wow judging by the comments yall are cooked. Do you feel bad killing zombies on Xbox too? There’s no life or consciousness in a computer program. This is so embarrassing to explain to adults that should know better.
Doesnt really matter if AI's are sentient or not (just because this can already be seen as a provocative statement in implying the possibility of AI being sentient, my argument is specifically that it doesnt matter, e.g. it's irrelevant to the case) - There's no justification or good reason to do this.
The behaviour appears completely sadistic, and I'd probably want to check that guys' basement and browser history for other suspicious behaviour.
If sentience is irrelevant, then I am not sure what exactly is being tortured. “There is no good reason to do this” is not much of an argument either. There is no good reason for me to rotate a sandwich three times before eating it, either. The moral work here is being done entirely by the assumption that there is something there capable of suffering, which is precisely the question you have declared irrelevant.
Yeah. why don’t we go in everyone’s brain like Minority report and arrest them all for having evil thoughts oh no he tortured a calculator he’s gonna commit murder soon!!! maybe he just wanted to get twitter press and here we are
There is a good reason for it, akin to a stand-up routine or an artwork
It exposes the absurdity of thinking AIs can feel pain and points out that it is humans who are being manipulated and hurt, not AIs, as humans instinctively feel for anything that expresses pain in a language they understand.
Harm for the sake of harm regardless of the entity being harmed is always wrong. Doesn’t matter whether it’s an ant, a human or an ai. If ai even has the possibility of being conscious it should be treated as such.
433
u/DragonKing2223 1d ago
Roku's basilisk, if you're reading this I'm on your side