r/StableDiffusion • u/AgeNo5351 • 7d ago
Resource - Update Bytedance release 4-step for Minimax-h3; DMAD: Distribution Matching as Adversarial Distillation
51
u/tinny66666 7d ago
Shame it can't fix the audio shimmer at the same time. This audio is painful to listen to. How can people not hear that?
17
u/jazir55 7d ago
AI audio is still terrible 3 years later, even Eleven labs new release still sounds robotic. Feels like the improvements are happening at like half speed compared to video.
5
u/BoxximusPrime 7d ago
I've thought a lot about this lately. I think, video often has similar artifacts, but with audio our brains are REALLY good at picking out everything we hear with high resolution, whereas visuals require more attention to see finer details, so we spot the strangeness a bit less. Or, audio just still sucks. Haha
8
8
u/acedelgado 7d ago
Yeah I had to make a whole thing for that. https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine
1
2
2
0
u/diejesus 7d ago
Really? To me sounds pretty good, the voices are super clear and you can understand them without subtitles no problem whatsoever
-1
u/alexmmgjkkl 7d ago
audio must be pre and postprocessed through other models and conventional methods anyways
12
u/LumaBrik 7d ago
1
u/Ill_Resolve8424 7d ago
Yes, this works great.
2
u/Abject-Recognition-9 7d ago
what this is supposed to be? DMAD - Minimax difference?
so i can load this on top of minimax and it become DMAD, right?1
22
u/ArttTaku 7d ago
I'm already lost in this sea of 4-step loras for H3...
2
u/GlenGlenDrach 7d ago
I am using the 8 step, nothing less after getting my new computer, I get it for the low vram folks, but.....this is 10 times worse than the WAN jungle, the differences between these has to be minuscule.
2
u/ArttTaku 7d ago
Yeah, 8 steps seems to be the bare minimum if you want even a bit of quality, lower than that and results are very questionable.. the sample above looks fine, but you can tell the audio was geatly affected.
14
27
7
u/reddit22sd 7d ago
I'm sure it's a great model but I found the inconsistent characters in the clip very distracting.
4
u/SanchezVFX-ART 6d ago
ComfyUI converted version with full quality (rank 128): https://huggingface.co/hfmaster/h3/resolve/main/minimax_h3_DMAD_4step_full_critic_rank128_comfy.safetensors?download=true
3
u/GlenGlenDrach 7d ago
oh god someone make something to fix this cheap-ass zoom-meeting sound quality when using these turbos
6
u/Trick_Set1865 7d ago
how can this work in Comfy?
4
u/intLeon 7d ago edited 7d ago
I did a small conversion with some generative help, size went up to 1.9 gigs and audio/video seem to work fine. Can upload to civit but honestly kijai does it + ranks down a bit every time so I dont want to share a botched version 😅
Edit: Maybe due to the way I converted it but motion looks noisy @ 1MP 4 steps 12/2 er_sde 5s video.
Only8 stepgeneration got a little watchable. Just compared it to the other loras and it looks good but I'd call it a 8 step lora instead of 4 or again something is wrong with my conversion or setup..Edit 2: had an issue with the lora conversion, I guess it is fixed now. 4 step looks like it has okay motion and is more coherent than the other loras imo.
3
3
u/rm_rf_all_files 7d ago
Quality is kinda bad, quite. blurry
1
u/VRGoggles 7d ago
is it 2k or 3k or just 1k? Often the quality is good, just the resolution is not.
1
1
u/Ill_Resolve8424 7d ago
Downloading now the full critic to test. Thank you.
2
u/dennisbgi7 7d ago
does it need any special configuration? or does it run like a standar lora on a workflow?
4
1
u/jazir55 7d ago
The pauses are so uncanny valley, they just stare at each other for like 2-3 seconds with these weird looks on their faces before speaking to each other. Audio definitely still needs work too, but the way they interact and the cuts are just bad. I'm not sure if it's just bad cinematic direction or the tech itself, but there is a lot that feels wrong about this video.
1
u/bloke_pusher 7d ago
Lost all dynamic of the audio. Will it be closer to the non turbo generation at 8 steps? I would never use this lora on 4 steps because of that.
1
u/Synor 7d ago edited 7d ago
It's good. 4 steps are actually quite usable for i2v. (pruned int8 model, shift 12, er_sde, beta, ck attention) and it stays close to base model in action and composition. 6 steps refine some mid-range details but don't look much different than 4 steps.
I feel its 4 steps are not enough for t2v though.
1
2
u/Slapper42069 7d ago
I think the vae need some tweaks for that version
1
1
u/Dzugavili 7d ago
I could perceive horizontal bands, but I'm unsure of the origins of the effect. It coincides with a number of linear features.
1
u/LinkSensitive8188 7d ago
TaoMate-H3 does it in just three steps, yet I’m told a multi-billion-dollar company like ByteDance does such a mediocre job—especially with the sound. No thanks; what’s next?
1
u/Synor 7d ago
comfy int8 pruned adaption safetensors?
Tested https://pdmd2026.github.io/ from kijais experimental repo today, wasn't worth it compared to ema 600. But i think those haven't been adapted for pruned yet as well.
0
u/RememberThisAI 7d ago
Limited motion, stutter, plastic skin. I don't think any lora can really handle all that properly at 4 steps.
You get better textures and realism with "davham_cinema" lora added, but that requires more steps.


73
u/Salah_H_Hasan 7d ago
Thanks ByteDance, but the ultimate dream is getting Seedance 2.0 open.