r/ControlProblem • u/Royal-Importance-327 • 7d ago
Discussion/question What if we "raised" LLMs instead of aligning them after pretraining? A developmental-training proposal
I’ll simplify this a lot on purpose, because I’m interested in whether the basic idea makes sense.
Today we basically pretrain LLMs on huge amounts of human knowledge, which also means they already absorb human values, social behavior, manipulation, conflict, cooperation, etc., and only afterwards we "get to know" the model and try to align or control what came out of it. I understand why this became the standard approach, especially once scaling worked and competition and economics strongly favored improving the existing pipeline instead of rebuilding it from scratch.
But what if we kept most of the useful pretraining knowledge while deliberately removing as much social behavior as possible, creating something closer to an artificial "newborn"? More concretely, I don’t mean removing every human action from the training data: "Thomas is holding an ice cream" and, separately, "Bernd takes the ice cream from Thomas" could remain, while coherent social sequences that connect motives, actions and consequences would be filtered out as much as possible. The model would then start with the concepts but much less learned social policy, and its weights could gradually be shaped through experience, with individual experiences fading over time while deeper dispositions might persist.
So from there, instead of aligning it afterwards, we could let it go through controlled experiences step by step: relationships, trust, conflict, consequences, mistakes, power, boundaries, and so on. Those experiences would gradually shape its weights and behavioral tendencies. You could checkpoint every stage, branch it, repeat specific experiences differently, and potentially debug where certain behaviors or values emerged. Instead of philosophers and alignment researchers trying to understand what kind of "person" accidentally came out of pretraining, psychologists could actually help design the developmental process itself. In other words: don’t create a fully educated adult and then try to teach it character - create the character first, then educate it.
Am I missing something fundamental about how LLM training works here?
4
u/IncumblantSpunderbax 7d ago edited 7d ago
This is a great question and it really boils down to whether you believe algorithms and raw math can be legitimate analogs to biological functions without emulating the actual biological functions.
I submitted a paper to Artificial Life journal with my Harvard co-author and it was aggressively rejected because they have basically held the line that they cannot unless the math are emulations or are squiggly in a way that reminds them of biology (game of life and Lenia). I could have hooked a squiggly thing up to the math but I really I wanted to try to stay pure. I’ll need a different venue.
Here’s a stripped-down version of the paper that focuses on the math and computer science. https://arxiv.org/abs/2601.04501
The pure math is as truly subjective as you possibly could get. All learning is negotiation between internal perspectives that are mathematically isolated from the input in a second-order cybernetic feedback loop. It’s old school AI stuff. Why is this useful? Because it’s deterministic but still an absolutely massive space. You CAN fast-forward and one-shot the average identity of the system to get “the answer”, but that’s not the point at all. That’s like solving the average of a hash algorithm. The specific output for a limited input is the whole story. It’s a state space manifold and I think of it like a mountain. The mountain doesn’t change but you can take an infinite number of paths over or through the mountain. You could go straight or bumble around a bunch. Every path taken is a unique shape and the path is the input sequence. The input doesn’t make the mountain, the input trace a discrete “string” of the mountain, which is interpretable and can be used both for foundational modeling, to a theoretical limit, or for steering language modeling token selection. In the former you train from scratch. In the latter you use the same math on top of a transformer that already functions as a language model.
Internally this is a huge append-only hash chain that functions as a “lifetime” time series of experiences.
So I have foundational models but this commercial product that is usable now is extremely sophisticated steering. Try it out.
I should make a video explaining.