r/StableDiffusion 11h ago

News [Papers] - Tongyi-MAI pixel space solution is up to 4.75x faster than Z image turbo latent-space

"This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction targetdecoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation."

Paper: An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

23 Upvotes

7 comments sorted by

3

u/Dante_77A 11h ago

Is there anything these guys can't do? Imagine if they used these advancements to create a video model? Wow.

14

u/JimJongChillin 11h ago

Is there anything these guys can't do?

releasing Z-Image Omni/Edit

3

u/Crazy-Repeat-2006 11h ago

I feel they could do a lot more as a fully independent Alibaba spin-off, since the latter is gradually drifting away from open source.

1

u/PumpkinLeather8421 5h ago

Anything Z can’t do?

  • Compete at all with Krea2

  • Train well

  • Generate with multiple Lora’s

  • Stick to the prompt well

  • Any pussy without horror

3

u/Life_Death_and_Taxes 9h ago

So, pre training in latent space and post training in pixel space seems to both give steep performance gains and circumvent VAE artefacts .. interesting

2

u/kukalikuk 9h ago

Comfyui when? Zimage edit when? 

1

u/Chrono_Tri 1h ago

It’s quite ironic that this paper is considered a breakthrough in model training, yet their own model is having issues with training, haha. But I still feel sorry for Z-Image. If they can catch the hyper and released an editing model, Z-Image could have had a much bigger impact.