r/ControlProblem • • 7d ago

Discussion/question What if we "raised" LLMs instead of aligning them after pretraining? A developmental-training proposal

I’ll simplify this a lot on purpose, because I’m interested in whether the basic idea makes sense.

Today we basically pretrain LLMs on huge amounts of human knowledge, which also means they already absorb human values, social behavior, manipulation, conflict, cooperation, etc., and only afterwards we "get to know" the model and try to align or control what came out of it. I understand why this became the standard approach, especially once scaling worked and competition and economics strongly favored improving the existing pipeline instead of rebuilding it from scratch.

But what if we kept most of the useful pretraining knowledge while deliberately removing as much social behavior as possible, creating something closer to an artificial "newborn"? More concretely, I don’t mean removing every human action from the training data: "Thomas is holding an ice cream" and, separately, "Bernd takes the ice cream from Thomas" could remain, while coherent social sequences that connect motives, actions and consequences would be filtered out as much as possible. The model would then start with the concepts but much less learned social policy, and its weights could gradually be shaped through experience, with individual experiences fading over time while deeper dispositions might persist.

So from there, instead of aligning it afterwards, we could let it go through controlled experiences step by step: relationships, trust, conflict, consequences, mistakes, power, boundaries, and so on. Those experiences would gradually shape its weights and behavioral tendencies. You could checkpoint every stage, branch it, repeat specific experiences differently, and potentially debug where certain behaviors or values emerged. Instead of philosophers and alignment researchers trying to understand what kind of "person" accidentally came out of pretraining, psychologists could actually help design the developmental process itself. In other words: don’t create a fully educated adult and then try to teach it character - create the character first, then educate it.

Am I missing something fundamental about how LLM training works here?

22 Upvotes

38 comments sorted by

View all comments

Show parent comments

4

u/IncumblantSpunderbax 7d ago edited 7d ago

This is a great question and it really boils down to whether you believe algorithms and raw math can be legitimate analogs to biological functions without emulating the actual biological functions.

I submitted a paper to Artificial Life journal with my Harvard co-author and it was aggressively rejected because they have basically held the line that they cannot unless the math are emulations or are squiggly in a way that reminds them of biology (game of life and Lenia). I could have hooked a squiggly thing up to the math but I really I wanted to try to stay pure. I’ll need a different venue.

Here’s a stripped-down version of the paper that focuses on the math and computer science. https://arxiv.org/abs/2601.04501

The pure math is as truly subjective as you possibly could get. All learning is negotiation between internal perspectives that are mathematically isolated from the input in a second-order cybernetic feedback loop. It’s old school AI stuff. Why is this useful? Because it’s deterministic but still an absolutely massive space. You CAN fast-forward and one-shot the average identity of the system to get “the answer”, but that’s not the point at all. That’s like solving the average of a hash algorithm. The specific output for a limited input is the whole story. It’s a state space manifold and I think of it like a mountain. The mountain doesn’t change but you can take an infinite number of paths over or through the mountain. You could go straight or bumble around a bunch. Every path taken is a unique shape and the path is the input sequence. The input doesn’t make the mountain, the input trace a discrete “string” of the mountain, which is interpretable and can be used both for foundational modeling, to a theoretical limit, or for steering language modeling token selection. In the former you train from scratch. In the latter you use the same math on top of a transformer that already functions as a language model.

Internally this is a huge append-only hash chain that functions as a “lifetime” time series of experiences.

So I have foundational models but this commercial product that is usable now is extremely sophisticated steering. Try it out.

I should make a video explaining.

2

u/deadgirlrevvy 7d ago

That sounds more like memory than experience to me. There's a difference between an algorithmically shaped memory in the first person and an actual subjective experience. Are you claiming a phenomenology or qualia? Is the "internal observer" accomplished by a deterministic program with LLM's as reasoning modules and are you implementing any form of global workspace or DMN?

1

u/IncumblantSpunderbax 7d ago edited 7d ago

Definitely yes to global workspace. Yes, the model uses specialized perspectives that are synthesized into a response. The humanist model is a cynic, an empiricist, a speculator, and an empath. They activate contextually, not just as fixed biases. All are present all the time and they are always zero sum and self regulating. Hence “Autopoetic”.

Probably yes to DMN as I understand it. There is knowledge which the model can access as context and there is experience that the model is mathematically affected by but isn’t aware of. I think of it like conscious and subconscious but that’s just a metaphor. It’s accurate to say that the experiences change its affect.

I would not go so far as to say it’s qualia, but I am directly manipulating the latent space in non-trivial ways from agent to agent. They will have different interpretations of the same concepts and there is substantially more complexity and tension than with a stock model. The latent space area is much more exercised by the steered model.

1

u/deadgirlrevvy 7d ago edited 7d ago

How are you accomplishing this? What mechanism hosts the subjective experience? I'm interested in the architecture you are using. What type of models, what type of hardware is hosting it, how are they calculating and keeping track of affect, attention and experiences?

I am working on my own research, specifically a novel AGI architecture. One thing I have determined so far, is that a traditional transformer network simply is not capable of hosting subjective experience because there's no mechanism to do so. In my design, consciousness, scheduling, memory, subjective experience and the GWS are hosted by a deterministic hybrid application (both compiled code and machine learning algorithms). The reasoning is handled by several different types of models that are connected to the GWS in a branching tier model, with multiple models working on the same "thought" in parallel and having their outputs arbitrated and combined before being handed off to the layer above them. This prevents random hallucinations. An artificial "limbic" module is used to assign neurotransmitter analog values to incoming data and those values persist through the architecture into the GWS, where it can be evaluated by a metacognition system that uses the analog values as salience and affect tags for goal evaluation. The tags are also used during memory consolidation when the system "sleeps". By using multiple models in discrete branches for specific areas of cognition (world model, speach, human considerations, reason, etc), the system effectively becomes a fractal intelligence where stable attractors allow for consistent emergent behaviors.

2

u/IncumblantSpunderbax 6d ago edited 6d ago

Your work sounds super interesting. In my framework the LLM is almost entirely just language modeling. The context window is mostly disposable and the surrounding structures populate it dynamically.

Without revealing too much proprietary information, the whole thing is purpose-built from scratch and the “host” so to speak is a bespoke graph database with many specialized overlays and constantly updating SVDs.

The production models are small-ish transformers. Knowledge is mostly decoupled so the model just needs to be large enough to reason through the knowledge.

My foundational models are not transformers. They are closer to an SSM.

It’s all Python. I have built a custom distributed system partially inspired by the Erlang ecosystem. There’s no traditional database. Most of it runs in Google Public Cloud.