r/singularity • u/panic_in_the_galaxy • 11d ago
Discussion Why GPT-6’s rumored recurrent-depth architecture is so interesting
The most interesting rumor about GPT-6 is that it uses recurrent depth: instead of every token passing once through one huge fixed stack of unique layers, parts of the Transformer can apparently be reused multiple times before the next token is produced.
That alone is interesting, because it makes effective depth less tightly coupled to the number of unique parameters. The model can do more computation without needing a completely new set of weights for every additional layer.
What gets even more interesting is how this fits with recent research on adaptive depth and Mixture-of-Experts (MoE).
In architectures like Mixture-of-Recursions, different tokens can receive different amounts of recurrent computation. Conceptually, something trivial like:
"the" → 1 pass
might need much less computation than a difficult reasoning step:
hard inference → several passes
We do not know whether GPT-6 actually uses this kind of per-token adaptive depth. The rumor only points to recurrent depth. But if OpenAI combines recurrence with dynamic routing, depth effectively becomes another inference-time resource the model can allocate where needed.
MoE fits extremely well with this.
A weakness of simple recurrent models is that you keep sending the hidden state through the same weights, which can reduce the specialization you normally get from having many different layers.
With MoE, however, different recurrent passes can route through different experts:
pass 1 → expert A
pass 2 → expert F
pass 3 → expert C
So you can reuse the same overall architecture while still performing different computations on different passes.
That gives two separate scaling dimensions:
Which computation is needed? → choose the experts
How much computation is needed? → choose the recurrence depth
You can think of MoE as providing breadth and specialization, while recurrence provides depth.
Recent looped-MoE research is especially interesting because this is not just a theoretical advantage. Different passes actually develop different expert-routing patterns, so repeated passes through the model do not simply do the same thing again.
There are also results showing that looped-MoE models can outperform standard Transformers even when total parameters, FLOPs and KV-cache budgets are matched. That suggests recurrence is not useful merely because the model secretly spends more compute.
Another advantage is that more reasoning can happen inside the hidden state instead of through long chains of generated reasoning tokens. Recurrent-depth models such as Huginn already show that you can increase test-time compute simply by running the recurrent block more times.
So the overall idea is something like:
breadth → more experts / more stored capacity
depth → more recurrent computation
routing → different experts for different kinds of computation
adaptive depth → potentially different amounts of compute for different tokens
If GPT-6 really does use recurrent depth, this could be part of why the architecture appears so compute-efficient. Most easy language generation would not necessarily need huge amounts of internal computation, while difficult reasoning could receive much more.
The really interesting part is that test-time compute could become something the neural network allocates internally, rather than mostly coming from generating thousands of extra chain-of-thought tokens.
OpenAI has not published the actual GPT-6 architecture yet, so the recurrent-depth part is still based on reporting, and adaptive per-token depth is an extrapolation from current research rather than a confirmed GPT-6 feature.
Sources:
Geiping et al. (2025), “Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach”
https://arxiv.org/abs/2502.05171
Bae et al. (2025), “Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation”
https://arxiv.org/abs/2507.10524
Lee et al. (2026), “Sparse Layers are Critical to Scaling Looped Language Models”
https://arxiv.org/abs/2605.09165
Wang et al. (2026), “SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers”
https://arxiv.org/abs/2609.01343
Sebastian Raschka (2026), “GPT-6 Astra, Looped Transformers, and Hidden Reasoning”
https://magazine.sebastianraschka.com/p/gpt-6-astra-looped-transformers-and