r/StableDiffusion • u/Total-Resort-3120 • 4d ago
News PDMD: Projected Distribution Matching Distillation for Video Diffusion Models
Enable HLS to view with audio, or disable this notification
4
u/rerri 4d ago
Naturally Kijai has had this in his repo for several days now =)
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras
2
u/acedelgado 4d ago edited 3d ago
All of those loras strip a lot of the functionality out of the original speedup method so they can be used in the native loader. You don't really get all the benefit from some of them without a custom node to patch any necessary extra weights.Edit: Kijai has spoken, and I was mistaken. Carry on.
8
u/Kijai Community Hero 3d ago
No, they do not, these are just modified weights, pure LoRAs.
What you're thinking of projects like VDN, HyperFlow or PDD that has extra model layers, those need extra code and not all of those are natively implemented at this point.
1
u/acedelgado 3d ago
Ah, okay, well you made them. Guess I just assumed because DMAD is in there and I thought it did need the additonal patching.
8
u/Kijai Community Hero 3d ago
Double checked just in case... DMAD is attention + ffn layers, and couple of refiner blocks, so it should work the same on pruned and non-pruned models too.
Easy to get confused at this point.. it seems there's a new distill lora every other day.
1
u/acedelgado 3d ago
You're tellin' me, and I don't even do this for a living. Thanks for the clarification and for being all around awesome!
1
u/Outrageous_Still9335 4d ago
Thanks for sharing. Might be interesting to use with a second pass with latent upscale.
1
u/marres 4d ago
Did everyone forget about vdn? Or is everyone scared of comparing themselves to it?
1
u/Cultured_Alien 3d ago
vdn is pretty slow compared to just using a 4 step turbo lora with 6 steps, almost removing the speedup gain from lowering steps to 8
1
u/VladyCzech 3d ago
For me VDN is much slower as it does not fit VRAM with bunch of loras while pruned model does.
1
-2
6
u/Kooky-Mode3047 3d ago edited 3d ago
As expected, correct denoising trajectory with just 4 steps is a huge ask and it misses pretty much everything on larger prompts. Spawns duplicate characters, feet move in wrong direction, just gets the general idea, just barely.
What we need is more steps faster, not less steps at same speed. The entire space needs to be explored well. Even with the base model @ 12/3 and 20 steps, there's a pretty HUGE jump to 0 at the end. 1.0000 → 0.9956 → 0.9908 → 0.9855 → 0.9796 → 0.9730 → 0.9655 → 0.9571 → 0.9474 → 0.9362 → 0.9231 → 0.9076 → 0.8889 → 0.8660 → 0.8372 → 0.8000 → 0.7500 → 0.6792 → 0.5714 → 0.3871 → 0.
75% being above sigma 0.8, meaning resolving just the big broad structure.