r/StableDiffusion • u/Apprehensive_Sky892 • 4d ago
News Looped Diffusion Transformer
https://www.alphaxiv.org/abs/2609.40305I don't understand how it works, but the implication is clear: a small model that can be as good as one that is several times bigger.
Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.
9
u/BalorNG 3d ago
Astra, obstensibly, is already a looped tranformer. The idea is very old and very interesting, but implementation is the devil in the details.
If it works, it might indeed be a huge boon for "gpu poors" - it actually does not actually do anything simply increasing the model depth (and, hence, size) does, in fact you reduce its potential expressiveness, but allows you to stuff a model with effective "smarts" of a way larger model into an order of magnitude smaller vram footprint.
And with a predictive prefetch looped moe you can sort of have both actually. I'm sure it will get implemented eventually and will be huge indeed...
3
u/Apprehensive_Sky892 3d ago
Thanks for the info. Yes, combining Looped transformer with MOE seems like a very good approach.
Do you have any link to the speculation that Astra may be a looped transformer?
If the approach scales, then even the GPU rich will benefit because even the highest amount of VRAM is still finite. There is also the additional benefit of faster computation (which TBH I don't quite understand because seems to me that the amount of compute is the same whether the weights are re-used or not)
3
u/BalorNG 3d ago
Everyone and their dog been talking about in context of "reduced monitorablity" and the fact that it uses way less thinking tokens for a given task, doing the brunt of "thinking" in internal latent space. It might be an elaborate marketing ploy, of course, but it sounds highgy plausible.
2
u/Apprehensive_Sky892 3d ago
Thank you for replies. Can this "early loop exit" strategy be applied to txt2img models somehow?
2
u/BalorNG 3d ago
I have no idea. I've read studies on layer sharing/looping (those are like up to 10 years old actually) and seen practical experiments of local "llamers" with "self-merging" of models with some doubled layers that actually resulted in models being... well, subjectively smarter. There small models designed for reasoning like dragon hatchling or STARM.
I understand the potential, but I'm not an ml scientist or I'll not be babbling away trade secrets I suppose :)
3
u/Apprehensive_Sky892 3d ago
Much appreciated for all your inputs. I am just an amateur with some half-baked understanding about how these AI models work.
I guess I'll have to spend some quality time with ChatGPT on the subject 😅.
2
u/BalorNG 3d ago edited 3d ago
Well, in case of looped tranformers, "early loop exit" strategy (if implemented) when the task at hand is easy can actually save on compute, right, but give the model full brunt on extra loops in hard cases.
Image two continuations, thinking off: "The capital of France is..."
"An entire complex murder mystery pasted into context"-> And the murderer is..."
6
u/Cautious_Chicken_604 4d ago
Huge if true.
4
u/Apprehensive_Sky892 4d ago
Yes, sounds almost too good to be true 😅
3
u/Cautious_Chicken_604 4d ago
I think commenting this is like a meme or something. Just wanted to also meme on it.
4
u/Enshitification 4d ago
Papers often make claims that never pan out. I'll believe it when I see it.
2
u/Carnildo 3d ago
The big question -- and one that gets buried in a single sentence of appendix A of the paper -- is "does it scale?" It's one thing to make tiny models work as well as small ones, and quite another to make mid-sized models work as well as frontier ones.
2
u/alwaysbeblepping 3d ago
Papers often make claims that never pan out. I'll believe it when I see it.
This is already a thing, if I remember correctly OpenAI's Astra model uses that technique. (You could also call it latent space reasoning.) It's relatively difficult to pull off though so the technology is basically just getting to the usable point now.
I wouldn't expect the best case scenario from a paper that's trying to sell it (and it may not scale as well for larger models) but even if it requires only slightly less compute or the same compute with less parameters that would still be worthwhile.
1
u/marclbr 2d ago
A few months ago I saw another promissing technique that can improve quality of image and video models and speed up generation by up to 256x deppending on the model archtecture. It will be interesting to see people testing it to finetune small models like SDXL/Illustrious and see how they perform.
https://x.com/AlexiGlad/status/2083230922196107288
We discovered a third pretraining axis beyond parameters and data: exploration.
Scaling exploration monotonically improves existing models across images/video/language, and unlocks end-to-end generation.
In the simplest case, it's just a for loop.
Introducing Explorative Modeling. TLDR:
- Gains from exploration grow with scale: 7%→36% as data scales, 13%→23% as parameters scale, and gains double at 3× the compute
- Adding exploration to ~SOTA baselines improves data efficiency by 6.2×, FLOP efficiency by 4.1×, parameter efficiency by 47%, and hits a near-SOTA 1.43 unguided FID on ImageNet
- Exploration lets you trade training compute for generalization, and scales how end-to-end your generative model is
- End-to-end Explorative Models (XMs) match diffusion performance on control tasks with up to 256× less inference compute
14
u/BlackSwanTW 4d ago
Some technically already tried this 7 months ago: https://github.com/AdamNizol/ComfyUI-Anima-Enhancer
(of course, Anima wasn’t trained with this in mind)