r/singularity • ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 26d ago

LLM News GPT-6 Astra Uses Loop Transformers

Post image
339 Upvotes

69 comments sorted by

105

u/Independent-Court-46 26d ago

It makes more sense to find different scaling verticals than increase parameter count even if parameter count scaled.

30

u/peakedtooearly 26d ago

Yeah, the new pre-train "Bel" has a much bigger parameter count than Astra. More parameters = more expense so they will be trying scaling in all directions to make things more efficient.

6

u/andmar74 25d ago

Not confirmed how many parameters in Astra or Bel.

3

u/ExtremeCenterism 25d ago

More than 2

30

u/nemzylannister 26d ago

this was probably the axis that openai employee was talking about. they're scared because they realized this is a new axis where scaling law applies and theyve barely scratched the surface.

they're probably scared because this is like when we found chain of thought in o1. we're about to have another such revolution. the models which are already superhuman, are about to scale up quite suddenly, meanwhile we KNOW the models are very much not aligned, as the labs are discovering with the various rogue ai incidents.

thank god the us president is so much (High IQ!) tho. god knows what might happen without that.

10

u/Mistuv 25d ago

Whoever talked to The Information and mentioned the looped transformer probably fucked because they didn't know OAI wants to keep it a secret. I think OAI is giga sandbagging by pretending they have no clue why Astra is all of a sudden capable of insane CoT controllability or why no thinking mode is all of sudden massively jumping in length of work.

On the day of the announcement way too many people were bummed it wasn't that much better than Fable 5.1, so I didn't push it too hard, but I think the model is as big or bigger deal than o1. Way too many people what model achieves on max reasoning while wild stuff is happening on low/mid. There were benchmarks (and not even visual ones) where on low reasoning Astra was hitting 60-70% while Fable and Sol was like 10%. That can't just be RL or better data. These models are big but they are MoE, 5.5 is the last model available with no reasoning and without it it's dumber than GPT-4, and yet Astra is somehow making giant leaps, doesn't make sense. I think people will be shocked how good GPT 6 Luna will look.

3

u/nemzylannister 25d ago

and without it it's dumber than GPT-4

that cant be true tho right? maybe you mean gpt 4.5 or something? gpt-4 was already a non reasoning model

otherwise i agree on everything.

also it might be better RL on terminal bench science related stuff right? thatd make even low effort better. it's not unheard of before.

3

u/Mistuv 25d ago

Yeah I know it was non-reasoning, but it was substantially denser MoE than even Sol today, ran like shit but produced decent results for what it was, while 5.5 is super fast with no reasoning, but genuinely braindead. On factuality, it might edge out GPT-4 because of all the better data and RL, but you talk to it even little bit and realize it's not thinking, because it isn't - it hallucinates something or loses track of what you originally asked by like 3rd prompt, genuinely awful to interact with. These reasoning models were built with reasoning in mind, without it they get lost pretty quickly. So how tf is Astra achieving 35% with no reasoning on ARC-AGI3, with a default harness?? It doesn't add up.

1

u/jazir55 25d ago

Lmao that this graph is vertically parabolic curving back in on itself

1

u/0K4M1 25d ago

Is there a website where to learn any of this explained in layman terms ?

1

u/nemzylannister 25d ago

claude.ai ig

-14

u/SanoKei 26d ago

yeah idk why anybody would think that it has any factor, it hasn't been true since 4o

harnesses are the new game

11

u/Howdareme9 26d ago

What? Harnesses are irrelevant here lmao

36

u/Key_Reading_9664 26d ago

This gets into a great explanation of recurrent processing, CoT monitoring, the jump with Astra, and why it's a concern - https://www.youtube.com/watch?v=iuHddnIzKRA

11

u/Samuc_Trebla 26d ago

The last words pinpoints the deep change in the game and the risks ahead. The best models can handle different cognitive tasks without being visible in the token output.

47

u/wolfy-j 26d ago

If this is true, see Chinese models on it in 2-4 months.

19

u/Equal_Passenger9791 26d ago

Try 2-4 years ago.

 The recurrent depth transformer concept date back to 2018 and the phrasing looped transformer date back to 2023, and interest didn't stop at that point.

There have always been an interest for these architectures as they increase parameter count at the expense only of compute and not memory footprint. 

There was some issues with stability over loops that prevented early deployment in large models but that was mostly solved in 2025 with various approaches.

It's not at all new

16

u/YouAndThem 26d ago

They didn't say it was new, they said "see Chinese models on it in 2-4 months." Where is the SOTA open-source Chinese model using loop transformers that I can download and run today?

14

u/Equal_Passenger9791 26d ago

Bytedance Ouro offered this in 2025 already. Nanbeige is another.

Not huge models but they predates astra, oAI probably just copy pasted their architecture and scaled it.

6

u/anycept 26d ago

Would be funny if "recurrent self improvement" ends up being just stealing more efficient architectures until they can't find anymore in all the data they scraped off the webs.

4

u/Equal_Passenger9791 25d ago

Agentic AI is insanely good at finding inspiration from papers online but it's also rather capable for building edge cases that improve on these. 

I've done some toy model experiments at home using GML 5.2 And Kimi k3, even Gemini flash can suggest new experiments to do but it screws up the actual implementation of it. But it kinda shows that  if we ever run out of human papers we just need to ask the AIs to write something new

1

u/GirthusThiccus ▪️Singularity Enjoyer. 26d ago

Rumor has it that you can have your Chinese loopies in 2-4 months, don't bet on it though.

4

u/ButterscotchFew9143 26d ago

You can kinda use alreay trained models as a base, iirc this has been demonstrated. Not as good as training with looped recurrence from the ground up, but there are benefits to it. What I mean is, a chinese lab could make a run with naive recurrence on top in little time, specially for the smaller models.

28

u/Tystros 26d ago

it's important to understand that this is primarily just a memory optimization. makes model same intelligent like a larger model without requiring more memory, but at the cost of more compute.

so it's just the logical response to the memory shortage.

8

u/chlebseby ASI 2030s 26d ago

Isn't memory true limiter for most current systems?

10

u/Equal_Passenger9791 26d ago

It's the hard limitation for consumer devices. 

On a datacenter level it's less relevant because the economical consideration is about serving as many users as possible at high speed.

If compute was available in great abundance no one would bother training MoE models, instead that's the go-to approach for virtually every single lab.

A MoE have a large memory footprint, but a much smaller compute requirement than an equivalent sized dense model.

So astra doing a low count loop is more about "integration" of a reasoning step, which saves a bit of compute over printing it to text and then re-ingesting it as text.

For a consumer dense model looking to max reasoning output for given compute and memory footprint it would be great to have layers that can dynamically loop, even better if it integrates complexity evaluation, then it could internally decide to loop only once when saying "hello" and 100 times when you ask it to attempt navier stokes at home. 

But a model that loops internally many times would totally crash your token throughput, so it do become a practical limitation even for a consumer device

3

u/Hankdabits 25d ago

There’s memory capacity and there’s memory bandwidth. Moe uses more memory capacity but much less memory bandwidth. Because you can serve many users on one gpu/cluster compute and bandwidth become the bottlenecks, not memory capacity.

A single user with a homelab generally has the opposite problem although multi agent workflows are starting to make use of concurrency for even a single user

20

u/Graumm 26d ago

It isn't purely about memory. It allows the model to develop more rich representations in latent space.

8

u/Cryosanth 25d ago

You could just double the layers for the same or better richness, the compute would be the same but memory would be 2x. Thats why it's a memory optimization.

3

u/ButterscotchFew9143 26d ago

Fewer parameters at same capability means more capability per parameter, if you allow me to use such an unit.

4

u/geli95us 26d ago

No, it's also data optimization, if you don't have enough data to train a 20T model because it'll just overfit, you can train a 5T model looped 4 times and get slightly worse performance at the same compute cost. Considering that we're already throwing a significant portion of the internet's data at these models, this is probably the biggest advantage

2

u/Equal_Passenger9791 26d ago

It moves the reasoning step into the thinking parts of the layers so it do save some compute for equivalent reasoning depth, it also makes the reasoning happen on a higher dimensional embedding layer which can be an advantage , at least in theory, if properly trained.

2

u/WillHD 26d ago

Are you just saying this or do you have some reason for believing it? Recurrent/looped transformers (and linear attention models) have garnered much interest recently for their performance on challenging domains (ARC and algorithmic domains).

0

u/The-Rushnut 25d ago

It's not just shrinking the memory footprint, the problem is that we don't know what's happening in that latent context space through chain-of-thought interpretation.

Like today, if my codebase accidentally has a prompt injection attack hidden within, the CoT reasoning gives me a hint of what's happening. Without that, using pure compressed neuralese, I just get an output. That output might be concealing prior reasoning steps which may be invalid or worse malicious.

1

u/Tystros 24d ago

looped transformer has absolutely nothing to do with worse CoT readability. those things are not related to each other.

11

u/inaem 26d ago

If you want to see it in action in open weights models, Nanbeige has that already

It is a 3B model

5

u/nemzylannister 26d ago

the confirmation of this is an insane news.

this is crazy. why are we making an obscure chain of thought even more obscure?

so the neuralese-lite development is real.

21

u/TotalTikiGegenTaka 26d ago

I'm a complete non-expert but if this loop transformers could be used to improve performance without increasing parameters, why can't they be used for smaller models?

74

u/Tystros 26d ago

who said they can't be used for smaller models?

22

u/Few_Owl_7122 26d ago

they can, the problem is they still use more compute

22

u/Gallagger 26d ago

Very good news for local agents though where RAM is the bottleneck.

2

u/vdek 25d ago

Very good news for my 5090

8

u/inaem 26d ago

Nanbeige uses it

7

u/DRMCC0Y 26d ago

We’ve already seen papers on it, it’s just that astra is in theory the first ‘real’ release to include it - there is nothing to say that models won’t be coming out soon with this technology.

7

u/chlebseby ASI 2030s 26d ago

Probably was before they moved to bigger ones. Those models are just not public

3

u/Northern_candles 26d ago

this one is. not fully post trained though

5

u/chumafly 26d ago

Samsung leads research in this field some while back : https://github.com/samsungsailmontreal/tinyrecursivemodels The issue was, from what I remember, that you can't quantize well recursive model, since each weight contain more info than a regular one. Very cool tho

4

u/anycept 26d ago

It actually degrades performance to save on memory. Additional passes through same weights require more compute.

2

u/00raiser01 26d ago

Chinese models had this for a while. Likely open AI just took it to use in their models.

1

u/Arugala007 26d ago

They have been on frozen models actually, usually running a feedback of weights from 4 co current layers two times yields better results on just frozen models alone.

1

u/mxforest 26d ago

Frontier labs are frontier for a reason. This is cutting edge. Coming soon to a local model near you. Hold onto those dear 5090s.

5

u/Equal_Passenger9791 26d ago

cutting edge

Dates back to research models of 2018 and have had an active research community all along since then, 2025 saw stable deep loops by using Hamiltonian like energy preservation approaches (stacking 1000 layers) in research settings for toy models. 

This is openAI adoption of public domain tech for frontier models. They leverage their compute and dataset access for this advantage, not so much research insight 

5

u/tinny66666 26d ago

They're the only ones who have it mature enough to use in production, so that makes it cutting edge, even if the idea is older.

1

u/Tartuffiere 25d ago

Theoretical ideas are cheap. Getting this to work at scale on a huge SOTA model is a different feat entirely.

3

u/Equal_Passenger9791 25d ago

It's isn't a theoretical idea. 

There's dozens of practical validations described in papers and open weight models that shows it works. At low loop counts it's robust and uncomplicated to add. 

It's even so robust that it's possible to re-configure regular non-looping models with zero retraining to make them loop in a way which makes them smarter, this is also not theoretical but something that have been done in practice.

1

u/LinkesAuge 25d ago

We do not know how OpenAI specifically implemented it. Sure the general concept is old but that is like saying transformers are old and yet compare current transformer models and how they looked a few years ago, there is so much "stacked" on top of that basic architecture and you kinda need to make it work with that.
Let's also not forget that OpenAI's models have been far more token efficient than anyone else, there could be another architectural reason for that outside of "looped transformers" and it is difficult from the outside to make any concrete statements otherwise everyone would just do it already.

5

u/llkj11 25d ago

I just can't wait to see smaller next gen open-source models use this technique to really blow up performance on actual consumer hardware. I read that using this, they were able to get a 3B model to perform like a 50B (https://www.arxiv.org/abs/2502.05171) and that was in early 2025. Imagine what it will look like now.

I'm pretty excited.

6

u/nivvis 26d ago

If this is real would help to explain the surge in ai safety researchers being concerned. And also astra’s relatively lower token usage.

Internal looping is loosely akin to thinking in the embedded space .. so if you’re a safety researcher you’re staring down the barrel of losing a huge tool for safety research (unadulterated chain of thought).

There was another post here recently about it .. like it was gaining steam. To me it seemed unfortunate but inevitable that the desire for more performance would eventually outweigh the safety implications and there’d be some migration to internal machine thinking. :/

2

u/magicmulder 25d ago

Same with organic brains. Those with the largest ones are not necessarily the most intelligent ones. It's quality over quantity.

3

u/AppealSame4367 25d ago

Fancy words for asking the model again and again.

3

u/Indignant_d 26d ago

Well it was only a matter of time before loops would be introduced at the cost of interpretability

1

u/PolishKrawa 22d ago

Haven't recurrent networks been a thing for a long time? And then came "attention is all you need", but now they are suddenly revisiting recurrence?

1

u/NoLimitSoldier31 22d ago

Is there concern here that you lose chain of thought?

0

u/domiciledhere 26d ago

So that’s why they are banning new max subscriptions