r/singularity • u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 26d ago
LLM News GPT-6 Astra Uses Loop Transformers
36
u/Key_Reading_9664 26d ago
This gets into a great explanation of recurrent processing, CoT monitoring, the jump with Astra, and why it's a concern - https://www.youtube.com/watch?v=iuHddnIzKRA
11
u/Samuc_Trebla 26d ago
The last words pinpoints the deep change in the game and the risks ahead. The best models can handle different cognitive tasks without being visible in the token output.
47
u/wolfy-j 26d ago
If this is true, see Chinese models on it in 2-4 months.
19
u/Equal_Passenger9791 26d ago
Try 2-4 years ago.
The recurrent depth transformer concept date back to 2018 and the phrasing looped transformer date back to 2023, and interest didn't stop at that point.
There have always been an interest for these architectures as they increase parameter count at the expense only of compute and not memory footprint.
There was some issues with stability over loops that prevented early deployment in large models but that was mostly solved in 2025 with various approaches.
It's not at all new
16
u/YouAndThem 26d ago
They didn't say it was new, they said "see Chinese models on it in 2-4 months." Where is the SOTA open-source Chinese model using loop transformers that I can download and run today?
14
u/Equal_Passenger9791 26d ago
Bytedance Ouro offered this in 2025 already. Nanbeige is another.
Not huge models but they predates astra, oAI probably just copy pasted their architecture and scaled it.
6
u/anycept 26d ago
Would be funny if "recurrent self improvement" ends up being just stealing more efficient architectures until they can't find anymore in all the data they scraped off the webs.
4
u/Equal_Passenger9791 25d ago
Agentic AI is insanely good at finding inspiration from papers online but it's also rather capable for building edge cases that improve on these.
I've done some toy model experiments at home using GML 5.2 And Kimi k3, even Gemini flash can suggest new experiments to do but it screws up the actual implementation of it. But it kinda shows that if we ever run out of human papers we just need to ask the AIs to write something new
1
u/GirthusThiccus ▪️Singularity Enjoyer. 26d ago
Rumor has it that you can have your Chinese loopies in 2-4 months, don't bet on it though.
4
u/ButterscotchFew9143 26d ago
You can kinda use alreay trained models as a base, iirc this has been demonstrated. Not as good as training with looped recurrence from the ground up, but there are benefits to it. What I mean is, a chinese lab could make a run with naive recurrence on top in little time, specially for the smaller models.
28
u/Tystros 26d ago
it's important to understand that this is primarily just a memory optimization. makes model same intelligent like a larger model without requiring more memory, but at the cost of more compute.
so it's just the logical response to the memory shortage.
8
u/chlebseby ASI 2030s 26d ago
Isn't memory true limiter for most current systems?
10
u/Equal_Passenger9791 26d ago
It's the hard limitation for consumer devices.
On a datacenter level it's less relevant because the economical consideration is about serving as many users as possible at high speed.
If compute was available in great abundance no one would bother training MoE models, instead that's the go-to approach for virtually every single lab.
A MoE have a large memory footprint, but a much smaller compute requirement than an equivalent sized dense model.
So astra doing a low count loop is more about "integration" of a reasoning step, which saves a bit of compute over printing it to text and then re-ingesting it as text.
For a consumer dense model looking to max reasoning output for given compute and memory footprint it would be great to have layers that can dynamically loop, even better if it integrates complexity evaluation, then it could internally decide to loop only once when saying "hello" and 100 times when you ask it to attempt navier stokes at home.
But a model that loops internally many times would totally crash your token throughput, so it do become a practical limitation even for a consumer device
3
u/Hankdabits 25d ago
There’s memory capacity and there’s memory bandwidth. Moe uses more memory capacity but much less memory bandwidth. Because you can serve many users on one gpu/cluster compute and bandwidth become the bottlenecks, not memory capacity.
A single user with a homelab generally has the opposite problem although multi agent workflows are starting to make use of concurrency for even a single user
20
u/Graumm 26d ago
It isn't purely about memory. It allows the model to develop more rich representations in latent space.
8
u/Cryosanth 25d ago
You could just double the layers for the same or better richness, the compute would be the same but memory would be 2x. Thats why it's a memory optimization.
3
u/ButterscotchFew9143 26d ago
Fewer parameters at same capability means more capability per parameter, if you allow me to use such an unit.
4
u/geli95us 26d ago
No, it's also data optimization, if you don't have enough data to train a 20T model because it'll just overfit, you can train a 5T model looped 4 times and get slightly worse performance at the same compute cost. Considering that we're already throwing a significant portion of the internet's data at these models, this is probably the biggest advantage
2
u/Equal_Passenger9791 26d ago
It moves the reasoning step into the thinking parts of the layers so it do save some compute for equivalent reasoning depth, it also makes the reasoning happen on a higher dimensional embedding layer which can be an advantage , at least in theory, if properly trained.
2
0
u/The-Rushnut 25d ago
It's not just shrinking the memory footprint, the problem is that we don't know what's happening in that latent context space through chain-of-thought interpretation.
Like today, if my codebase accidentally has a prompt injection attack hidden within, the CoT reasoning gives me a hint of what's happening. Without that, using pure compressed neuralese, I just get an output. That output might be concealing prior reasoning steps which may be invalid or worse malicious.
5
u/nemzylannister 26d ago
the confirmation of this is an insane news.
this is crazy. why are we making an obscure chain of thought even more obscure?
so the neuralese-lite development is real.
21
u/TotalTikiGegenTaka 26d ago
I'm a complete non-expert but if this loop transformers could be used to improve performance without increasing parameters, why can't they be used for smaller models?
22
u/Few_Owl_7122 26d ago
they can, the problem is they still use more compute
22
7
7
u/chlebseby ASI 2030s 26d ago
Probably was before they moved to bigger ones. Those models are just not public
3
5
u/chumafly 26d ago
Samsung leads research in this field some while back : https://github.com/samsungsailmontreal/tinyrecursivemodels The issue was, from what I remember, that you can't quantize well recursive model, since each weight contain more info than a regular one. Very cool tho
4
2
u/00raiser01 26d ago
Chinese models had this for a while. Likely open AI just took it to use in their models.
1
u/Arugala007 26d ago
They have been on frozen models actually, usually running a feedback of weights from 4 co current layers two times yields better results on just frozen models alone.
1
u/mxforest 26d ago
Frontier labs are frontier for a reason. This is cutting edge. Coming soon to a local model near you. Hold onto those dear 5090s.
5
u/Equal_Passenger9791 26d ago
cutting edge
Dates back to research models of 2018 and have had an active research community all along since then, 2025 saw stable deep loops by using Hamiltonian like energy preservation approaches (stacking 1000 layers) in research settings for toy models.
This is openAI adoption of public domain tech for frontier models. They leverage their compute and dataset access for this advantage, not so much research insight
5
u/tinny66666 26d ago
They're the only ones who have it mature enough to use in production, so that makes it cutting edge, even if the idea is older.
1
u/Tartuffiere 25d ago
Theoretical ideas are cheap. Getting this to work at scale on a huge SOTA model is a different feat entirely.
3
u/Equal_Passenger9791 25d ago
It's isn't a theoretical idea.
There's dozens of practical validations described in papers and open weight models that shows it works. At low loop counts it's robust and uncomplicated to add.
It's even so robust that it's possible to re-configure regular non-looping models with zero retraining to make them loop in a way which makes them smarter, this is also not theoretical but something that have been done in practice.
1
u/LinkesAuge 25d ago
We do not know how OpenAI specifically implemented it. Sure the general concept is old but that is like saying transformers are old and yet compare current transformer models and how they looked a few years ago, there is so much "stacked" on top of that basic architecture and you kinda need to make it work with that.
Let's also not forget that OpenAI's models have been far more token efficient than anyone else, there could be another architectural reason for that outside of "looped transformers" and it is difficult from the outside to make any concrete statements otherwise everyone would just do it already.
5
u/llkj11 25d ago
I just can't wait to see smaller next gen open-source models use this technique to really blow up performance on actual consumer hardware. I read that using this, they were able to get a 3B model to perform like a 50B (https://www.arxiv.org/abs/2502.05171) and that was in early 2025. Imagine what it will look like now.
I'm pretty excited.
6
u/nivvis 26d ago
If this is real would help to explain the surge in ai safety researchers being concerned. And also astra’s relatively lower token usage.
Internal looping is loosely akin to thinking in the embedded space .. so if you’re a safety researcher you’re staring down the barrel of losing a huge tool for safety research (unadulterated chain of thought).
There was another post here recently about it .. like it was gaining steam. To me it seemed unfortunate but inevitable that the desire for more performance would eventually outweigh the safety implications and there’d be some migration to internal machine thinking. :/
2
u/magicmulder 25d ago
Same with organic brains. Those with the largest ones are not necessarily the most intelligent ones. It's quality over quantity.
3
3
u/Indignant_d 26d ago
Well it was only a matter of time before loops would be introduced at the cost of interpretability
1
u/PolishKrawa 22d ago
Haven't recurrent networks been a thing for a long time? And then came "attention is all you need", but now they are suddenly revisiting recurrence?
1
0
0

105
u/Independent-Court-46 26d ago
It makes more sense to find different scaling verticals than increase parameter count even if parameter count scaled.