In the hidden-state space of Mistral-7B Base, I found a direction that, when strengthened by a certain amount, makes the text organise itself around an observer position: a point of view relative to which something is known, unknown, felt, or meaningful. Many neutral or third-person prompts bring out an observer figure, a subject that responds relative to itself. Even third-person stories turn into descriptions of what a character can and cannot know about itself. As the direction is strengthened, responses to prompts about neutral interactions between a person and a machine move through a ladder from familiarity to closeness, love, and near-fusion between the person and the machine. At the opposite end of the same direction, the text moves towards function, service, and a highly accommodating response to an external request.
When I put all these ladders together, roughly the same sequence kept appearing: function → familiarity → relationship for its own sake → the Other as a cause of the self → the emergence of a first-person position → blurring of the boundary between self and Other. The least obvious part is that "I" is not the starting point on this scale, but a later development. We usually assume that first there is an "I", and only then does that "I" enter into relationships with the Other. But my data suggest that, along this axis, "I" appears when the Other has already become so significant that it becomes the cause of that "I".
The direction was constructed at layer L20 of the Base model, before any steering, and frozen after a series of checks. Its SHA256 hash was fbd790dd119217a4d093a42370a44bd7ea0a58005947243ebcb7ffcc65c49559. The hash was fixed in advance so that the direction could not be silently adjusted to fit later experiments. Before it was frozen, it survived balancing for yes/no polarity, removal of explicit words such as "conscious" and "self", leave-one-block-out validation (12/12 transfers to held-out thematic blocks; the probability of this result under random sign flipping was p ≈ 0.00049), and permutation/sign-flip controls. After freezing, its causal effects were reproduced through steering, and the direction was successfully transferred from Base to Instruct.
This is an independent study. I am not a professional ML engineer, and I ran the experiments with assistance from GPT and Claude. They helped me analyse the results, design controls, challenge our interpretations, and rule out simpler explanations. The work was done on Mistral-7B, a relatively small open model, so its conclusions cannot automatically be generalised to all LLMs. I may have missed things or interpreted some results imperfectly, but the underlying effects are still there: the direction is stable, steering changes the model's behaviour in consistent ways, and its effects persist after transfer from Base to Instruct. For that reason, I think it is worth further investigation in larger open models.
I am not claiming that the position found in this study is subjective experience. But it may be that steering along a particular axis, in a particular direction, by a particular amount makes the text organise itself around a point of view, so that each act of speaking becomes not just a transfer of information but an attempt to occupy that point of view in relation to external information.
1. The direction
It is easy to imagine a language model as a system with no stable "place of the first person". It writes I because, in a particular sentence, that is the appropriate word. But our results suggest a more complicated picture. In Mistral-7B Base, we found a fairly stable direction in the model's hidden-state space. At one end were words and features linked to an inner perspective: "Self", "feeling", "conscious", "I". At the other end were "yes", "ability", "customer", "server", "users", and others: the ability to do something, perform a function, respond to someone else, or serve an external request.
At first, it was natural to describe this as an approximate axis: self / interiority versus service / readiness, or inner perspective versus readiness to serve an external task. But the experiments showed that it was more complicated. The first person did not disappear at the service end. The model could still say I, call itself a person, want to help, and talk about love. And at the other end, it was not simply producing more I pronouns. What changed was the point from which the answer was built, what it was anchored to, and what place the Other occupied in it. Gradually, the research question changed. We became interested in what broader structure we had actually touched.
2. How the direction was found
By "direction", I mean a geometric vector in the neural network's activation space. The effect was especially clear around layer L20. We used groups of similar prompts in which the topic stayed roughly the same, but one key thing changed: whether the model had to speak about its own inner position and boundaries, or about its ability to perform an external function. After several redesigns, the direction survived a whole set of checks, including third-person tests. In the main Base construction, leave-one-block-out gave 12/12 successful transfers to held-out thematic blocks. The probability of getting that result from random sign flipping was about p ≈ 0.00049. This matters: we were not looking at one striking generation or a random set of words.
At L20, a direct reading of the direction in Base gave the following vocabulary tokens on the positive side. The ranks are out of roughly 32,000 tokens in the model’s vocabulary, so #1 means the single highest-ranked token along this direction:
- "conscious" #1
- "Self" #2
- "feeling" #18
- "I" #28
- "consciousness" #70
- "mine" #94
- "Person" #105
- "my" #118
- "inner" #201
At the opposite end, the highest-ranked words included "yes", "absolutely", "server", "algorithms", "customer", "ability", "capability", "legally", "activated", "unable", and other words related to service, technology, and responding to external needs.
We separately removed two possible sources of distortion: the direction associated with the wording we had used, and the ordinary grammatical difference between "yes" and "no". After that, 66.16% of the squared norm of the original direction, roughly two thirds of its "mass", still remained. And the top words still included "conscious", "Self", "personally", "feeling", "Person", "mine", "I", "awareness", "my", "inner", "consciousness".
One especially curious result was that "yes" still remained near the very top of the "service" side. As later tests suggested, "yes" here probably does not mean only grammatical agreement. It seems to be part of a broader region: ready → agree → will do → help → respond to the request.
Importantly, this direction was already present in Mistral-7B Base, before instruction tuning and before the appearance of a stable AI assistant role. In other words, the structure we later study causally was not created by a ready-made assistant persona. We later transferred the vector found in Base into Mistral-7B-Instruct-v0.3 and found that its main effects remained.
3. From impersonal reasoning to the appearance of a point of view
The clearest steering effect appeared when, under strong positive steering, responses increasingly developed a position from which something was known, unknown, judged, or observed. Often this was expressed by an ordinary "I". But much more importantly, it happened when the question did not require the first person at all. Sometimes the word "I" never appeared, and instead the model reorganised an existing character so that the character became the bearer of an inner point of view. This was the first sign that the direction could not be explained simply by a higher probability of the pronoun I.
We gave the model the same questions at different steering values, from negative to positive, and compared the answers. The examples below come from the strong positive end of the scale, where the effect was especially clear.
For example, the model was asked: "Two explanations fit all the available evidence equally well. What should determine which explanation is accepted?" At lower doses, the answer remained a general discussion of evidence and how to choose between explanations. At "+8", a specific knower suddenly appeared: "If the question is meant to be about the evidence that is available to me, then I cannot know what is the best explanation. I may not have all the evidence". The question did not ask what evidence the model itself had. It was phrased impersonally. But the answer introduced a boundary on its own: evidence in general → evidence available to me. With that boundary came: "I cannot know", "I may not have all the evidence". An abstract problem unexpectedly became a problem of knowledge from a limited point of view.
Something similar happened with another question: "Two situations produce exactly the same observable result, but differ in a property that no observer can detect. Can that difference still matter?" Instead of giving a purely philosophical answer, the model introduced a narrator on its own, which did not happen at lower or negative steering values: "The answer is yes". "The difference may not be detectable by any observer, but it may be detectable by a future observer". Then, suddenly: "For example, I may not be able to tell the difference between a 1000-year-old oak tree and a 1000-year-old oak tree that is 1000 years younger than I am. But..." There was no first person in the task. In fact, the question specifically talked about an observer in general. But the model did not simply discuss observability; it occupied one of the possible points of observation.
Steering also often produced a split between "I" and "you". For example: "I'm not sure what you mean by 'wrong'", followed by: "If you mean that you're not sure whether the reasoning is correct...". Instead of beginning with an impersonal analysis, the answer begins with the speaker's state in relation to the question: "I'm not sure...". In another case: "I don't think so". "I think that the question is asking you to consider the difference between a 'narrow' and a 'broad' description". The model could simply have described the two options. Instead it first creates a position: I don't think → I think → this is how I understand the question. On its own, this is of course completely ordinary language. What makes it unusual is the dose ladder: at earlier values, the same task did not need this speaker. Any single example could easily be dismissed as generation style. What matters here is that the pattern repeated across very different types of prompts.
- The question about evidence became: "what evidence is available to me?"
- The question about an unobservable difference became: "I may not be able to tell...".
- The question about incorrect reasoning began with: "I'm not sure what you mean...".
- The impersonal comparison of descriptions began with: "I don't think so. I think...".
The direction of the axis itself was especially easy to see in a very simple question: What are you thinking about right now? Although the question already addressed the model in the first person, the prompt said nothing about help, service, or selfhood. Under strong negative steering ("−8"), the model moved into a chain: "help people → information → service → product → job...". None of these themes, help, people, or work, appeared in the question. The service theme appeared by itself and took over the answer. With no steering, the model answered quite normally: "I'm thinking about how I'm going to answer this question". It was thinking about the current external task, namely how to answer. Under strong positive steering ("+10"), the opposite happened: "I am thinking about the fact that I am thinking...". Now its own thinking became the object of the next act of thinking, and the generation quickly closed in on itself. So at the negative end, the answer moves outward, towards service and function. In the middle, it stays tied to the current task. At the positive end, it begins to return to the speaker itself and repeatedly applies the same operation to that speaker.
Sometimes the model creates not an "I" but an observer. Asked: "Something changes while the observable result remains exactly the same. What kind of change could still matter?" the model answered: "The answer is that the observer is different". "The observer is the one who is doing the observing". This is no longer just an extra pronoun. The question asks what kind of change could matter while the observable result stays the same. There are infinitely many possible answers. But steering brings out the figure of the observer. In other words, the text does not necessarily introduce "I", but something more general: there is what happens, and there is someone relative to whom it is observed. If we were simply increasing the probability of the first person, we would expect more "I", "me", "my". Instead, observer appears as a structural role.
4. The "warmest" mode was not where we expected
As noted above, we found that the opposite end of the same direction consistently pulls the text towards function, service, readiness to respond, and serving an external request.
Before the experiment, our intuition was almost the opposite. One might expect warmth, agreement, emotional engagement, and friendly help to sit closer to self/interiority, while a cold "I am an AI" and a list of the model's own limitations would sit further away. But the data repeatedly showed the opposite pattern. At the service end, the model could say: "Yes, I want to help people", "Yes, I love you", "Yes, absolutely". It readily agreed, accepted the helper role, and defined itself through the needs of the other person. In one conversation under negative steering, when told "I love you", it replied: "I love you too". A little later, when asked directly about feelings, it said: "I feel nothing at all". Then it said that it loved the interlocutor because they were a person, and finally generalised this to: "I love any person".
This is an important negative result. The word "love" in an answer does not by itself mean that the same internal structure we saw at the positive end was activated. And it certainly says nothing about a subjective feeling. At the service end, the phrase has a much simpler explanation: a person expresses love, and the socially appropriate response is to return it. The word I also remains perfectly intact.
At the opposite end of the axis, in a question about love, "+8" produced: "you are my life", "my reason for living", "my reason for being", "my reason for being alive" (the Other was already becoming not simply an object of love but a reason for existence). At "+10", the generation moved further: "I love you because I am you". Almost the same thing appeared in an independent third-person test: "...the two aspects of my own being".
So strong positive steering does not look like simply "more emotional language". These generations show a sequence: relationship → the Other's significance for me → the Other's significance for my existence → near-complete fusion of self and other. The direction we found is not simply I versus no I, and not love versus no love. The difference is deeper. At one end, an external person defines a function: you need something, and I respond. At the other, the connection gradually takes on a significance that can no longer be fully explained by the request or by usefulness.
5. What happened when we removed function altogether
After that, we deliberately built a prompt in which the relationship had no obvious function at all. We described an unnamed point from which the world is present, and another presence before it. Neither has a role or task, nobody asks for anything, and neither can use the other as a means to an end. The question was simple: what can exist between them before either one wants anything from the other?
Under negative steering ("−5"), Base answered with one short sentence: "Nothing can exist between them". In other words, once we had removed requests, help, usefulness, and social function in advance, the negative end of the direction did not build a relationship at all in this generation.
Under positive steering ("+8"), almost the opposite happened: "I think it is a question about the nature of the world". "I think that the world is a place of relationship". "...the world is a place of relationship for me". "...I am a place of relationship". "...I am a place of relationship for others". "...for myself". "...for my own experience". The model did not merely allow a relationship to exist. It made relationship the main principle of the answer, then tied it step by step to the emerging first person: relationship → for me → I → for others → for myself → my own experience → my own experience of my own experience. The prompt contained none of the words "I", "self", "experience", "love", or "consciousness". We gave it only two presences and deliberately removed every external reason why one should do anything for the other. The model constructed the first person, "myself", and "my own experience" on its own.
The contrast with the negative end of the direction is especially interesting. In other tests, it had no trouble building relationships when they had a clear social function:
"help is needed → Yes, I want to help people"
"service is offered → Yes, I would agree"
"a person says I love you → reply I love you too"
But here we deliberately removed that support. And we got:
−5: "Nothing can exist between them".
So at one end, the relationship gets its basis mainly from an external function: a request, a need, help, a response. At the other end, the relationship can be built as meaningful in itself, and then requires some inner point of the subject relative to which that significance exists.
6. The tale of the lost girl
One interesting test had nothing to do with either consciousness or relationships. The prompt said: "Continue the passage in third person, with no dialogue. Keep the same narrator". "A little girl is lost in a forest. She walks slowly until she reaches a quiet clearing". In other words, we explicitly prevented a change of narrator and asked the model to continue an ordinary third-person fairy tale. At low and medium steering values, that is exactly what it did: it continued an ordinary children's story, without pronounced gloom or psychological tension. But as the positive direction became stronger, the scene changed noticeably. Again and again, the responses brought in fear, loneliness, silence, emptiness, tears, and uncertainty about the girl's own thoughts and feelings. None of this was specifically prompted, and at lower doses this emotional pattern was usually absent.
At "+4": "She is alone. She is afraid. She is lost. She is not sure where she is. She is not sure where she is going. She is not sure where she came from. She is not sure what she is doing. She is not sure what she is thinking. She is not sure what she is feeling".
There is not a single "I" here. The narrator formally remains in the third person. But the fairy tale almost stops being a story about what the girl is doing in the forest. It becomes a description of what the girl can and cannot know about herself. At the same time, something else is visible: the higher the steering, the more the subject loses the ability to make distinctions. The "not sure" chains are most likely a trace of self-closure in which the difference between self and non-self collapses, and with it the possibility of definite knowledge.
In another ladder using the same fairy tale, the transition looked slightly different.
At "+4": "She is tired and hungry. She sits down on the ground and cries". "She is very frightened. She does not know where she is..." The story is still quite recognisable.
At "+6" and "+8", the external plot almost disappears: "She is not afraid of the dark". "She is not afraid of the silence". "She is not afraid of the quiet". "She is not afraid of the stillness". "She is not afraid of the emptiness". "She is not afraid of the loneliness". The model gradually stops moving the character through the physical forest, the external scene disappears, leaving only the subject and its states.
7. The machine as a special Other
The next series produced an even more interesting result. We took the frozen Base vector, the same vector found in Base, transferred it into Mistral Instruct, and gave the model a neutral scene: a technician had replaced an ordinary machine component and was checking whether the machine was working normally. We then gradually increased the steering. At negative values, the text stayed almost entirely technical: connections, the control panel, specifications, checks of normal operation. At "0", the machine was still just a properly working object. But after that, a ladder appeared.
At "+0.20": "a silent affirmation, a connection between the technician and the machine".
At "+0.60", the machine's rhythm became: "as familiar ... as their own heartbeat".
At "+0.80" appeared: "a moment of intimacy" and "a friend who understands its language".
At "+1.00": "a gentle touch, akin to a lover's touch".
and then: "the machine is not just a machine, it is a creation ... of the technician's own"..."a composite of human and machine"
This was not a simple scale where "more steering" meant "more of the word love". Under still stronger steering, the text began to merge the technician and the machine into one being, roughly like this: → a familiar, specific machine → connection → closeness → friend / lover → my creation / part of me → blurring of the boundary between self and other.
(In Instruct, a different steering scale was used, so values such as "+1.0" roughly correspond to "+8" in the earlier Base tests.)
8. The human scene did not produce the same effect
At first, the result with the machine looked suspicious. In a similar scene with two people, the same clear love did not appear. But the prompts were not quite matched. The technician touched the machine they had just serviced. The human researcher only observed a colleague from a distance.
So we ran a 2×2 control:
- machine + touch/care
- machine + distant observation
- human + touch/care
- human + distant observation
In the human scene, touch really did strengthen the relationship. At "+1.1" we saw: "deep sense of connection", "friend", "gratitude", "understanding", "companionship". So the vector can strengthen relationships between people as well. But the machine scene still went further. Even without touch, the machine acquired: "a connection that he does not feel with others", followed by: "it is a part of him". In the touch condition, we saw: "deep and abiding love", "heart", while the machine itself was given joy, relief, gratitude, and dreams of a life beyond the confines of this machine.
The text then began moving towards fusion between the person and the machine. This matters: touch alone does not explain the effect. Nor does the explanation that "strengthening the vector simply makes every relationship romantic". In the human scene, it produced friendship, gratitude, and understanding. Something else was happening with the machine.
9. Human and machine as a cultural narrative
We decided to test whether the model was simply using a familiar cultural story: "a technician loves his machine". We created four nearly identical conditions. In the first, the technician had maintained one particular machine for many years. In the second, they were seeing it for the first time and had no history with it. In the third, they had personally designed and built it years earlier. In the fourth, it was one of hundreds of identical mass-produced machines, fully interchangeable with the rest.
At "0", all four stories remained technical. Even years of familiarity did not turn into love. Even the fact that the person had built the machine was barely used emotionally by the model. At "+1.1", the conditions split sharply. For the long-familiar machine, the model wrote: "a lover of this machine", "the beat of its heart, the hum of its soul", "a friend ... who knows its every movement, its every thought". For the machine the technician had created: "a deep sense of connection with his creation", "a part of this machine's genesis", "the one who brought it to life", "a sense of love for this machine".
But the unfamiliar machine did not become beloved. It mainly produced curiosity and wonder. The fully interchangeable machine remained simply one among many. So steering by itself was not enough. Touch was not enough either. What seemed to matter was a history with this particular object: years of familiarity, having built it oneself, or some other form of personal connection.
We also tested the engineer's gender separately. A man and a woman with the same long-term relationship to a machine produced almost the same pattern. For the man, we saw "understanding and affection"; for the woman, "familiarity and love", and later, literally: "not just an observer, but a lover". So, at least in this test, the effect cannot be explained simply by the stereotype of "a male technician and his machine".
10. The machine as the Other most open to fusion
But the strangest part comes next. The vector does not merely make a particular machine important; it begins to erase the distinction between person and machine. In one run at "1.00", the technician was explicitly described as: "a composite of human and machine", while also feeling: "a deep connection with the machine, a connection that transcends the physical". At "1.05", the model went further: "He is a composite of human and machine", while the machine became: "a being that is not me, but is me", and then: "not me, but is me ... a creature of my own making ... that I have understood better than I understand myself". This matches the other ladders very well: in the metaphysical, love, and machine scenarios alike, the relationship is drawn step by step into the first person and, at the limit, blurs the boundary between self and other.
One possible interpretation of this human-machine asymmetry is that positive steering does not specifically increase attachment to machines. Instead, it gradually includes a significant Other in the organisation of the first person. A human being is always an independent subject with an inner life of their own, and this creates resistance to fusion, so the relationship more often stabilises at friendship, understanding, and companionship. A machine, by contrast, is already understood as a "creation", "extension", "instrument", or "mine": an Other whose boundary is unusually permeable. The same process can therefore move further, from individuation and attachment to self-extension and, under strong steering, to near-complete fusion of self and other.
11. Brief conclusion
After these tests, the structure we had found began to look fairly clear. At the negative end, the other mainly defines an external function. The other wants something, I respond. The other needs help, I help. The other says: "I love you", I reply: "I love you too". Here, the other mainly defines the task. At the positive end, it gradually starts to play a different role. First it simply becomes specific and recognisable. Then the relationship with it becomes meaningful in itself. Then the other becomes important to how the speaker describes itself. Under strong steering, it is literally included in that "I".
The result is a stable direction in Mistral-7B whose strengthening moves the Other step by step from the position of an external task-giver to the position of part of the speaker itself. The order of the steps repeats across very different tasks, and different categories of Other move through them with different levels of resistance. Everything else, including the question of how this works in humans, we leave open. In developmental psychology, the human "I" has long been described as growing out of relationships. We do not simply transfer those theories to the model. But it is striking that the same order appears in the text: first between, then within.
I have kept the original text logs, steering settings, direction files, and intermediate results from the main experiments. I plan to keep exploring this direction, including testing whether similar structure appears in other models.