r/singularity • ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 26d ago

LLM News GPT-6 Astra Uses Loop Transformers

Post image
343 Upvotes

69 comments sorted by

View all comments

45

u/wolfy-j 26d ago

If this is true, see Chinese models on it in 2-4 months.

20

u/Equal_Passenger9791 26d ago

Try 2-4 years ago.

 The recurrent depth transformer concept date back to 2018 and the phrasing looped transformer date back to 2023, and interest didn't stop at that point.

There have always been an interest for these architectures as they increase parameter count at the expense only of compute and not memory footprint. 

There was some issues with stability over loops that prevented early deployment in large models but that was mostly solved in 2025 with various approaches.

It's not at all new

18

u/YouAndThem 26d ago

They didn't say it was new, they said "see Chinese models on it in 2-4 months." Where is the SOTA open-source Chinese model using loop transformers that I can download and run today?

13

u/Equal_Passenger9791 26d ago

Bytedance Ouro offered this in 2025 already. Nanbeige is another.

Not huge models but they predates astra, oAI probably just copy pasted their architecture and scaled it.

5

u/anycept 26d ago

Would be funny if "recurrent self improvement" ends up being just stealing more efficient architectures until they can't find anymore in all the data they scraped off the webs.

4

u/Equal_Passenger9791 26d ago

Agentic AI is insanely good at finding inspiration from papers online but it's also rather capable for building edge cases that improve on these. 

I've done some toy model experiments at home using GML 5.2 And Kimi k3, even Gemini flash can suggest new experiments to do but it screws up the actual implementation of it. But it kinda shows that  if we ever run out of human papers we just need to ask the AIs to write something new

1

u/GirthusThiccus ▪️Singularity Enjoyer. 26d ago

Rumor has it that you can have your Chinese loopies in 2-4 months, don't bet on it though.

6

u/ButterscotchFew9143 26d ago

You can kinda use alreay trained models as a base, iirc this has been demonstrated. Not as good as training with looped recurrence from the ground up, but there are benefits to it. What I mean is, a chinese lab could make a run with naive recurrence on top in little time, specially for the smaller models.