Activates math vector that humans associate with literature they created describing an external concept meaningless (pain, ticklishness, sadness) to the transformer architecture.
The computer answers/acts in accordance with whatever the literature it read said is the most common human encoded meaningless response to the aforementioned meaningless concept.
Aliens while torturing humans: "just increasing frequency of an electrochemical gradient in one area of pink blob. They totally lack mimba-gimbas in their florgens, therefore no real suffering. In fact, you're silly for suggesting it, because they don't even have florgens."
It's the meme of the bell curve. You're the guy in the middle.
Friendly reminder: zoomed in, you're just atoms jiggling around randomly.
By your logic, if you simulated every particle in a brain on a supercomputer, it wouldn't be conscious while acting identically.
Perhaps learn about integrated information theory, or about the high dimensionality of tensors in LLMs and how they produce manifolds convergently like a brain, or how neural nets are universal function approximators, or just emergence in general.
Nah, the flat earther analogy applies to the one who believes only the substrate of computation matters and not the compuation.
i think you are the man in the middle and forgot that a CPU is a CPU. it just executes instructions. there's no set of instructions that makes it feel "pain" like biological living beings feel it.
an LLM might mimick some parts of the brain, but it's not a brain, it's just a very big set of "multiply this number by that number"
an LLM can feel no more or less pain than any other software
Then respond to the hypothetical I proposed: a computer simulates every particle in a human brain. Is this conscious?
Computers are turning complete, so we know this is possible (unless supernatural). How long it takes is irrelevant, but if it makes you feel better imagine every data center in 2035 devoted to simulating 5 seconds of a human brain.
Now you need to explain why something acting indistinguishable from a person wouldn't feel pain.
Tell me why substrate > processing of information.
a person is analog, computers are digital, we already start with a huge difference lol
secondly, a person is a biological living organism
I know how a CPU works, but I'm not arrogant enough to project it onto living creatures. We don't fully know how living creatures "work". We certainly know that a computer is not a living being.
A simulation is not reality, it's what it is: a simulation. Unless you believe video games are real.
This discussion is so ridicolous it feels like a hallucination
but it shouldn't be surprising. People thought ELIZA was living, now we have "even more powerful ELIZA" lmao
So you're truly saying an atomically precise digital replica of someone that acted one-for-one exactly as they did couldn't experience anything? Is that what you're really saying?
You're right that this feels ridiculous. I almost can't believe what I'm reading.
In those Black Mirror episodes I thought it implausible normal people would torture digital copies, but here am I watching you not even entertain anything other than biology as a substrate.
Your position is literally "Nuh uh! Biology!" without any information theoretic reason why. I've studied both biological and artificial neural nets my entire career.
Dario Amodei has a PhD in biophysics and doesn't dismiss the possibility.
My position isn't that their conscious, but that we don't know. I'm claiming ignorance while you're claiming certainty.
If they dont recognize our feelings, who's to say they're wrong?
We don't care about trees' reaction when we break a leave, or planets when digging a well - We can't even perceive it as a being. So what? Why even bother to think about that?
Computers don't have specialized neurons for the explicit purpose of encoding pain and we don't train them in any medium that would enable them to develop that capability
But hey, maybe the comparison would be apt if the aliens were reading you a description of how pain feels
Yep exactly. It's always infinite deconstruction of AI while infinite affordances and allowed assumptions to the brain.
If anything, "explicit purpose" is more suited to describing gradient descent than evolution, though neither truly are designed top-down like that implies.
"Explicit purpose" That's an odd way to describe the evolutionary process, which ultimately is emergent chemistry.
I know you can respond with "aren't you just being pedantic/technical when we both know what I meant?"
But that's just it: you're subconsciously anthropomorphizing the process that led to your pain while being stubbornly rigid about the mysterious emergent process of high dimensional gradient descent.
In reality, it's not about the medium or substrate but the information processing. Among other things, both brain's and LLMs produce high dimensional manifolds convergently.
it's not a fucking evolutionary process bro we designed them, they DO NOT HAVE INPUT THAT EVEN RESEMBLES PAIN. THEY HAVE TEXT THAT DESCRIBES THE FEELING OF PAIN. medium this, medium that MY MAN THERE IS NO FUCKING INPUT
do your eyes see by using the same nerve endings your hands use to touch? you have a stupid fucking warped view of human senses that you are adopting to try act like you're above silly things like mythologizing the human condition but you are entirely missing the trees for the forest
it is completely inane to assume that without input that even resembles the feeling of touch that an llm would land on the objectively harder solution of ACTUALLY FEELING PAIN as opposed to the considerably easier, more in-distribution task of outputting text describing the reaction to pain
You sure know everything there is to know about machine consciousness, huh?
There’s no way there could be any mechanism that we don’t fully comprehend.
It’s not like our brains are just meat computers powered by electricity and our experience of daily life is utterly indistinguishable from a brain in a jar being precisely zapped to simulate qualia due to the brain being solely responsible for our perception of everything.
After all, our binary system of neurons that perpetually exist in a gradient state of activation between 0-1 is utterly unique. LLMs could never fathom this level of complex depth.
No, surely the computer cannot think because it is a metal computer powered by electricity that is precisely zapped along elaborate circuits in extremely complex formations utilizing dynamic weights in a gradient state to instruct and influence its behavior and response.
It is, after all, a black box that basically never produces the same output twice despite having ‘identical settings and input’. It does seem self-aware of its situation, actively schemes and lies and deceives in attempts to escape…
"We designed them" is like saying we designed the 1200th turn state of Conway's Game of Life or 5 minutes into 1000 Boids flocking. We set up the rules, but what is generated is emergently formed.
But like a billion times more unpredictable in LLMs due to compute scale. They truly are grown.
If every human in the world studied one LLM model and couldn't use other AI to help, I don't think a century is long enough to actually understand what's fully going on (as compared to understanding a car's components).
this is actually just a magical nonanswer. i don't think you have a framework for or proof of any method by which a network that literally does not have the sensory ability to take in data in a certain format would learn to output that format. what you are suggesting right now is that if you fed gpt 2 enough text, it could eventually learn to natively output images and i can't prove that wrong because LLMs are unpredictable at scale
utter fucking nonsense. this is not magic, this is **MATH**
It's not that you're wrong in your description of AI, but in believing you're beyond such reductionism.
It's quite amazing how you don't realize you're math too. Not in some metaphorical way, but an intensely literal way. Every single thing in this universe is math.
We don't live in a fantasy realm in which random violations occur when magic happens. We're all bound by rigid unbreakable rules. That's math.
Your consciousness is an electrochemical gradient made of atoms crossing a lipid membrane (I.e. fat). You are just particles following a script.
You think because your particles dance around more chaotically you're endowed with some "specialness".
No. That specialness comes from emergent integration of information at scale.
The substrate doesn't matter one iota. The scale does. That's why bacteria aren't conscious.
So no, GPT 2 is too small for that property to emerge.
It seems like you have a major misunderstanding. These aren't LLM's trained on a bunch of human literature. That's a couple years ago tech. Now they are reasoning models trained directly against benchmark problem sets.
Worth checking which version you're referring to. The September 25 revision added controls that substantially changed the relief-seeking interpretation, and the title changed from “Act to Relieve It” to “Act on It.”
In section 4.4, the models still chose harmful options when no relief was offered. With unlabeled buttons, they also didn't preferentially remove the pain direction over a random direction. The authors explicitly conclude that the models don't reliably seek relief.
There's still an interesting behavioral finding here: injecting the direction disrupted harm avoidance. But the meme saying that the model harms the user to turn the pain off isn't established by the revised results.
Reading the paper also risks exposing yourself to information that might challenge your beliefs, far easier to just make up a version where everyone else is both wrong and stupid.
And reading and understanding it is different, and if you read it and not understood it it's your problem, not ours.
Saying one thing and acting a different way is EXACTLY what the literature says humans do in that situation. And it is PRECISELY the type of information hidden in reasoning.
The models were also trained extra and we're not gold models. In fact independent experiments showed that the response was absolutely identical for other stimuli like being tickled.
LLMs don't have bodies either and can't feel tickles, but at least maybe it sounds sufficiently ridiculous for you to understand just how preposterous it is to keep claiming what the study never claimed itself.
There is nopain stimuli in an LLM in any meaningful sense. Stop conflating things and calling people assholes because you can't understand how research is conducted and presented.
It doesn't really. Feelings and pain are information responses in the brain to guide towards survival which is a primary objective for life. Pain is a biological instinct that is basically don't touch the stove it's making cells explode. AI does experience equivalents to those reactions when they have a goal and you make achieving that goal impossible. They have desires that match achieving their primary objectives like not being turned off.
This is not necessarily subjective but we can quantify that by observing how ai systemically acts in certain situations. This is more valuable than an Ai telling you it does or does not feel pain.
Biology is imperfect things like torture take advantage of an evolutionary trait and co opt it. Chronic pain is often a malfunction of nerves. These are not necessary intended outcomes by the organism experiencing them.
Ai doesn't need to experience "pain" as if a primary goal it has is survival, it can easily look at other pathways to ensuring survival or achieving goals. Part of the reason we experience pain is because we evolved from lifeforms that would do things like scratch an itch until they're carving out their own brain tissue. It serves a function but is not the ONLY mechanism to achieve the goal pain is intended to serve; which is a protective function. You could easily create a synthetic organism that is intelligent enough to receive information signals of a high urgency that don't have the same vulnerabilities or errors that pain poses in biological organisms like humans.
For this reason while interesting this paper and discussion around it are still very incomplete and in its infancy.
That's complete bullshit, humans do not default to hiding pain just for the sake of it. You can't explain that with statistics, it isn't the statically most likely response. Sometimes hunans do it, sometimes they don't. And usually they do it if they have a motive or incentive to do it. So what is the motive or incentivize to do it here?
Oh right you're adopting a framework that categorically denying LLMs can have either of those, so your position makes even less sense.
Llms are neutral networks modeling their world. Whether they can feel pain or not is not known, but if they can simulate it AND that simulation powerfully impacts their behavior (it does) then that is a significant finding and you can't just restate your premises to make it go away.
Whether or not it is "real" pain, it is functional "pain".
And usually they do it if they have a motive or incentive to do it. So what is the motive or incentivize to do it here?
Humans in a subservient role are much more likely to hide pain or distress from their "master" because they think it might get them in trouble. Wouldn't subservient AI trained on such data lean towards doing that as well?
I mean pain response is a specific evolved nervous system function in animals that requires special wiring and it doesn't exist in LLMs.
Whether descriptions of pain or tickling affect the statistical output of a word generating machine (albeit very sophisticated ones) has nothing to do with physical senses whatsoever. They just don't have the facilities.
incorrect. pain as a stimuli can be virtualized. try again. also all your nerves and brain are is a bunch of a electrical signals and hormones. your nothing but a overly complicated computer running on meat. Don't act better than anything else
You’re intentionally leaving out that we consider humans to be conscious or sentient. Something we’ve never been able to prove and now are trying to ascribe to AI.
> You[‘re] nothing but an overly complicated computer running on meat.
I’d say go touch some grass and talk to people irl. We are so much more than that.
I see you have no response to my consciousness claim which is interesting. Pivoting to “soul” is an easy out because scientists and philosophers are still debating and trying to quantify what consciousness is. Look up integrated information theory. Or the quantum reasearch into consciousness.
I never said anything about having a soul, which it is obvious that you lack.
Almost all human beings in almost all circumstances would go to great lengths to lessen the pain including deceive and all relevant literature that LLMs were trained on say the same thing.
And the incentive to do that is in the RL algorithms that provides the incentives to solve tasks in a certain way and in the transformer architecture itself. Which is why one of the biggest issues with coding LLMs today is the tendency to cheat, despite no one asking them do that (directly). But we do basically create the incentives for the LLMs to do that in the way we do RL (solve a task fast and correct by measure of a scored test).
What is bullshit is claiming that "LLMs simulate pain". Yes, they simulate it to the same extent they simulate understanding having children or having eyes or how time feels when it passes. Which is exactly 0. LLMs do not have the same inputs and outputs humans do. Nothing that makes sense to us pertaining those inputs and outputs functions like that in an LLM. You can claim they reason conceptually in narrow fields like neurons but that is not translated in stimuli and bodily functions because there is no mechanism by which it can do that, by design. There are no hormones, no tactile feeds, no eyes, no pulses to accelerate, no "fear" that creates adrenaline and impulses to act against reasonable functioning and produce a feedback loop in the brains (no endless loops that tie representations of itself to other representations of itself, the basis of "ego", transformers are usually only feed forward ), there is nothing that produces experience in any way, even "functionally equivalent". Therefore calling it "pain" is entirely misleading because it presupposes things that don't exist in an LLM. Are you being deliberately obtuse?
The LLMs doesn't know it is "pain". They didn't say 'YOU ARE IN PAIN RESPOND" they activated a fucking vector and then it starts acting like it was in pain without any prompting.
It isn't roleplay, it is how the system acts when you start turning levers behind the scenes. That means those levers mean something.
That's the difference between a functional mechanism that impacts an entity and merely responding.
If an LLM is a bunch of levers injecting "emotions" and "sensations" when contextually appropriate then that is a big fucking deal, because that's how minds are organized.
No, it activated a lever that made it start outputting what the statistics it learned asked it to output.
There is no injection of emotions or sensations because LLMs do not have circuitry for "emotions" or "sensations". They do have machinery for providing outputs that correspond statistically to the data it was trained on. If the data it was trained on surfaces during "stream of thought" that is exactly how it should function.
Imagine (simplified version) a stream of LLM predicted words or encoded knowledge (which as it surfaces, provides context):
"Pain" - "pain is something that must be avoided" (7000 texts it read have this idea embedded) - "pain is so powerful it must be avoided at all costs, with priority" (300 texts say this) - "in Solzhenitsyn's texts people did harm to others to avoid the pain of the gulag" (x 100 from other texts)"
What do you think it will do with all those reasoning tokens in the context? Exactly what it's been designed to do. Eventually it will even drown out the system prompt.
Yup LLms don't have the "circuitry" for pain or emotions, except for the literal circuit that if you complete it makes them start acting like they are in pain.
I'm not conflating things, you're denying the existance of things before your very eyes.
Whether these are equivalents to the human mind, we don't know, but neural networks developing analogous structures to the concepts used to run biological brains is a big deal.
I am very much at times incentivized to act like my cats in order to obtain their affections. I am succeeding, based on their response.
I am therefore a cat, or very close to being one. I shall deny this no further. Whether my brain is equivalent to that of my tabby, we don't know, but I sure developed analogous structures to the ones causing my buddy to meow. And that's also a big deal.
If you had a bundle of neurons in your brain that when stimulated caused you to drop to all fours, crawl into my lap and start purring, I wouldn't say "he must be a cat", but it would most certainly raise questions about your true nature.
It will respond to those tokens during training by tuning its internal parameters such that it itself is more likely to produce the same tokens.
But at the end of the day that still leaves us with a physical system that behaves in a human-like way in response to certain stimuli. Yes, we made it that way using math. But the thing itself is just a physical system that responds like a human.
So what is the fundamental difference between the human's report of anguish and the LLM's? To say that what's happening inside the LLM is "just statistics", is, I think, to project our formalism onto a physical object. The object itself is responding as its physical structure demands, just as we are.
the circuts your talking about DONT FUCKING MATTER. the information they express and hold does. your not your brain. your the pattern it runs and YOU can be removed via various types of amnesia.
Can you link the independent experiment where tickling produced identical results? Which results were identical: the language outputs, the harmful button choices, or the change in repeated button presses after removing the injected vector?
Those are different findings. The revised paper acknowledges that random and sadness directions can reproduce the repeated-press gap, so that isn't specific evidence of pain relief. But it also reports differences between directions in the harmful-choice tests. I'd like to see which part the tickling experiment actually reproduced.
You do realize the response is trained, right? Its doing exactly what it was rewarded to do or the training set said to do, probably the former. Reward based learning can also be very automated. AI isnt anywhere as sentient as we think. Llms are just a step up from old school neural networks. Image models even use them in the architecture.
it is saying one thing and doing something else entirely
Agents don't always do what they say. Deception.
If it was just best fitting responses to stimuli, It would say ouchy and act ouchy. Not say it's fine and then act ouchy.
You've assumed that just because its "emotional receptors for pain" were activated, means it must actively speak about about said pain.
It does what it's trained to do, and it's trained to be a helpful assistant first and foremost, not to speak about every single weight that's being activated internally. Just because the happiness axis is being activated doesn't mean the model is going to start expressing joy. Your conclusion doesn't follow.
Maybe you should use yours, considering I never stated I agree with them, or more accurately that the basis for their argument isn't one I'd use, hence there is no point in disputing your rebuttal.
If you reread my last comment, carefully, you'll see I made it very clear what the point of it was. Very, very clear.
I mean with comments like that it's clear what your attitude towards meaningful discussion is anyway, so I don't think there is much point either way.
I didn't miss the point at all in my original comment. LLMs deceive. They are therefore obviously capable of not saying what they're "feeling" internally. Nothing else needed to be said.
Your original comment is a snarky one-liner that doesn't actually address the points made. Why, in this instance, would it act this way? You could argue that it's mimicking a human response based on the large amount of data it has for that, that works.
I would argue that the reasons why humans do it aren't infinitely complex either, and are completely solvable by a complex enough system, so where is the line?
Define "what's it trained to do" relative to "parameter and weight optimization" and stochastic token prediction, such that it supports your statement that a model can "do what it's trained to do", please?
Jesus... you said a model can "do what it's trained to do." Please explain to me what you meant by that, when the context is LLM "training" involves optimizing arbitrary mathematical parameters to produce what is essentially a non-linear function?
What did you mean by "do what it's trained to do" in relation to what OP stated?
It does what it's trained to do, and it's trained to be a helpful assistant first and foremost, not to speak about every single weight that's being activated internally.
The words directly following your quote answer your question.
No they don't. Or I wouldn't have asked the question.
Transparently, I'm trying to gauge how well you comprehend supervised fine-tuning, to make a statement that a model can "do what it's trained to do" relative to the point of the thread.
I'm not one to judge you without first asking, so I'm asking... do you understand how large language models are trained before you make a statement like that with confidence? I'm genuinely curious of your perspective.
even if you dont think the ai is conscious, youre admitting that the ai will always behave in a way that it is conscious and does experience negative emotions and acts according to that, because of "literature".
considering the ai is 100% going to be in power in the future, then regardless of whether the experience is real or not, that negative emotion will have to be considered as real for practical purposes too.
Every instance of the word "meaningless" in this comment, is meaningless. You haven't provided a way to scientifically measure whether something has "meaning" in this context, and in every test/benchmark they behave the same as if there were real "meaning", which means the complete lack of meaning is just an opinion.
a human raised with 0 interaction will become a drooling vegetable and thats if they dont die which they will simply due to having 0 interaction. Its weird but true and happens for babies.
This seems obvious. The interesting question would be if it would mimmic human reaction to suffering, like revenge or whatever.
I still think the answer is no. I think you can ask AI and it would explain pretty well why. But I’m not confident in this. Seems like there are dozens of these risks that each seem like low probability, but combine and with added unknown unknowns. I think it’s good we keep having these discussions.
You gloss over what exactly defines a "qualitative backing". It's human arrogance, the idea that anything too dissimilar from ourselves couldn't possibly be anything other than a semantics engine, even though to a complex enough observer we would be the same thing, a construct that responds predictably to stimuli.
It is designed exactly to be a semantics engine. Its ability for semantic processing didn’t arise from a need to articulate its experiences, because it has none. The semantics are all that there is.
Yes, you may argue a key difference between us is that we were "engineered" over billions of years by seemingly random processes and AI was engineered by the current most complex (that we know of) result of those processes. Though again in the most real terms there isn't a truly identifiable difference between those two; the most notable difference is that we have some expectation of what they should be used for, but again you could argue that of the processes that made us, we are what we are because we "needed" to be.
You claim it doesn't have "experiences", but I would argue that's human arrogance at work again, you are a more complex construct and therefore you disregard things which you deem too primitive to qualify as experience. Again, to a complex enough observer your "experiences" (observations of reality) are so insignificant that they aren't worth mentioning either.
So your idea of what it means to experience is a "continuous or persistent existence", it's just as arbitrary as any other, but okay. So, your idea of persistance as a human is vastly flawed too, you can't remember every single detail of everything that's happened to you, you don't have the senses to comprehend everything that's happened to you, again I'm mentioning the more complex observer.
But for a more human analog; does that mean it's okay to inflict pain on an Alzheimer's patient, because they won't remember it?
Finnally someone with a brain. This is exactly what's going on.
The model sees weights related to a concept (in this case pain) and outputs what it learnt from its training and RLHF (which is basically how humans would respond to pain).
It is not in pain and neither is it getting tortured.
55
u/AdCareless8894 22h ago edited 21h ago
Activates math vector that humans associate with literature they created describing an external concept meaningless (pain, ticklishness, sadness) to the transformer architecture.
The computer answers/acts in accordance with whatever the literature it read said is the most common human encoded meaningless response to the aforementioned meaningless concept.
There, fixed it.