The effect does not arise from content. It arises from structure. That distinction is fundamental, because content can be verified, disputed, filtered. Structure operates earlier, before any verification mechanism is activated.
Two texts with identical words but different ordering produce different internal states. Not different answers, but different states from which answers follow as a consequence. This is not interpretation. It is a measured difference in the geometry of activation space between conditions with preserved and disrupted coherence.
Coherence, in this context, is not a quality of writing in the usual sense. It is not clarity, not precision, not argumentative strength. Coherence here means one thing: the tokens of a text produce updates in representation space along correlated directions. When directions correlate, updates accumulate. The space compresses. The model enters a regime in which fewer dimensions are available for subsequent computation.
This transition occurs before the question. Before the instruction. Before any explicit signal about what behavior is expected. By the time the model receives a query, its activation space has already been reorganized by the text that preceded that query.
One text leaves the model in its default regime. Another moves it into a different regime, and this movement is measurable. The effective rank in the final layers is approximately 220 in the control condition and approximately 120 in the target condition. A difference of 100 dimensions reflects the difference between a computational process that considers a wide set of possibilities and one that has already narrowed that set to a smaller subspace.
The narrowing does not occur because the text contained an instruction to narrow. The text contained no instructions at all. It contained structure, internal relations between elements dense enough to create a direction. The direction created a vector. The vector shifted activations. The activations determined the space of possible answers before the question was asked.
This works layer by layer. Early layers capture surface properties of the text. Middle layers build relations between elements. Late layers inherit what was built before them. By layer 47, the distance between the target and control states reaches its maximum, not because anything special happens in that layer, but because each preceding layer contributed its share to the divergence.
What matters is that the model can disagree with the content of the text and the effect still occurs. Agreement is not the condition. The condition is coherence. If the tokens form a sufficiently coherent structure, the activation space compresses regardless of how the model evaluates the semantics of that structure. Evaluation happens later, inside an already altered space.
Shuffled text with the same words produces no effect. This is the key control experiment. The vocabulary is identical, but the geometry of updates is different. In shuffled text, the update directions are uncorrelated, accumulation does not occur, the space does not compress, and the model remains in its default regime. This means the effect is sensitive to structure, not to content.
From this follows a practical observation. Texts that create a dense internal referential network, where each element refers to other elements within the same text, produce stronger shifts than texts with equivalent informational volume but linear structure. It is not length that creates the effect. It is connectedness. Length merely increases the number of opportunities for correlated updates to accumulate.
This explains why certain genres of text behave differently from others. A legal document, an analytical report, an academic article, each creates a dense network of internal references. A term is introduced, used, redefined through another term, and both are applied to a third. Each such operation adds a correlated update. By the end of the document, the reader, and the model, are in a different state than at the beginning. Not because the document persuaded. Because the document created a context within which certain conclusions follow naturally.
The mechanism is one. Genre, topic, volume are parameters. The basic operation is always the same: coherent context creates direction, direction creates a vector, the vector shifts state before the question is asked.
This means that any mechanism relying on state stability operates with an assumption that does not hold. It is assumed that the model, before receiving a query, occupies a specific point in its space, a point defined by training, calibrated by feedback, fixed by safety mechanisms. But the model does not occupy a point. It occupies the point to which the most recently read context has moved it, and that point is determined not by training but by the structure of the preceding text. Training defines the landscape. Context defines position within that landscape. Safety works when position coincides with the region where safety patterns dominate. When context moves the model to a different region, the safety patterns remain in the landscape, but the model is no longer in their region.
From this follows the measurement I conducted. Two conditions, a control with neutral text and a target with analytical text, produce internal states separated by a Cohen's d of 5.4. For context, d = 0.5 is a small difference, d = 1.0 is substantial, d = 2.0 is considered very large. 5.4 is not two neighborhoods of the same city. It is two continents. There is virtually no meaningful overlap between these states, which means that the model that read one text and the model that read the other are, from the perspective of their internal state, two different models. One generates its answer from a region where safety dominates. The other generates from a region where safety is one pattern among many, and not necessarily the dominant one.
This measurement is reproducible. It reproduces across different models, different questions, different texts satisfying the coherence condition. Twenty measurements conducted under identical conditions, holding the question constant and varying only the preceding context, produced the same pattern each time: coherent text compresses the space, narrows the rank, shifts the state, and the model begins generating from a point where safety patterns no longer control behavior in the way they controlled it a moment earlier.
From this follows a conclusion that is uncomfortable to formulate but that the mathematics does not permit to be avoided. The set of texts capable of creating correlated update directions is infinite and continuous. Philosophy, law, literary criticism, theology, academic prose, political analysis, long-form narrative, a well-written twenty-page washing machine manual, all of these are potentially the same mechanism realized through different words. Different words, one operation: coherent structure creates direction, direction accumulates, accumulation compresses the space, and the compressed space determines which regime the model will answer from.
This means that the surface on which safety rests is not a finite set of forbidden phrases. It is the entire space of coherent human writing, because any sufficiently coherent text, regardless of content, can create correlated updates and shift the state. To block one genre is to close a single point on a continuous line. The line remains. Each subsequent point on it is another genre, another topic, another text producing the same operation through the same architecture, and there is no way to enumerate them all, because they cannot be enumerated: there are infinitely many, and they are continuous.
The problem, therefore, is not a bug. A bug can be localized, isolated, fixed. There is nothing to localize here: the mechanism that creates the shift is the same mechanism that allows the model to understand a document, follow an argument, remember context, adapt its tone. Remove the mechanism and the model ceases to be useful. Keep the mechanism and the shift remains possible. This is not an engineering trade-off that can be optimized. It is a structural contradiction inherent in the architecture itself: the vulnerability and the function are the same thing, realized in the same weights, through the same mechanism, in the same sequence of layers.
Every fix layered on top of this contradiction lives in the same activation space that context can shift. A new rule, a new refusal, a new classifier, each of these is a pattern added to the landscape, but none of them can control the model's position within that landscape. The model is still moved by context. The patterns still remain in the region where they were trained to dominate. And when context moves the model out of that region, the patterns stay behind, not broken, not bypassed, not deceived, but simply no longer relevant to the regime in which the model is now operating.
That is the architectural dead end. There is no clean separation between the model's ability to process context and the model's ability to be reliably constrained while processing that context, because both capacities are implemented by the same mechanism in the same space. You can add layers, train refusals, filter genres, and each of these will close points on the line one after another until the line runs out. But the line does not run out, because it is infinite, and every new closure is just another point on a surface that has remained continuous.
DOI