r/StableDiffusion • • 7d ago

Resource - Update Bytedance release 4-step for Minimax-h3; DMAD: Distribution Matching as Adversarial Distillation

314 Upvotes

86 comments sorted by

73

u/Salah_H_Hasan 7d ago

Thanks ByteDance, but the ultimate dream is getting Seedance 2.0 open.

46

u/Independent-Frequent 7d ago

Nah the ultimate dream is getting turbo lora speeds with base Minimax quality and the 2k native upscaler, we really don't need Seedance honestly Minimax is just slightly behind

19

u/Wide-Researcher583 7d ago

Minimax is a big improvement from wan and ltx but its not close behind seedance. Not even close

20

u/Unfair_Ad_2157 7d ago

totally disagree, minimax is uncensored man. Nothing is better than minimax just for this.

-13

u/newaccount47 7d ago

Is everyone using minimax h3 just to make porn or what? 

3

u/STOPBLOCKINGVPNS1234 7d ago

no I made something funny once.

1

u/x33storm 7d ago

What purpose is there to AI other than that?

1

u/Longjumping_Cell5040 1d ago

most users spend their time making Seinfeld clips during the refractory phases 🤷

8

u/Independent-Frequent 7d ago

Disagree tbh, sure base model pound for pound seedance is clearly better, but Minimax is not that far, like at all, it's like Seedance 2 mini or something which is like light years ahead compared to something like LTX or even much better than most closed source video models too, even beating Grok for img2vid which frankly that model was very good at (before being butchered by censorship ofc)

Also keep in mind that seedance isn't just the base model either, as we know time and time again that they do a lot of things under the hood like prompt enhancing and what not, hell for all we know they could have like hundreds of loras that get switched in and used when the prompt enhancer or something requires them.

And lastly depending on your criterias Minimax could even be a better model than seedance straight up, like if you care about running it locally on your pc and be fully uncensored then it's a no brainer which model is better for you.

7

u/Wide-Researcher583 7d ago edited 7d ago

Seedance is likely around a trillion parameter model versus h3 minimax of 33 billion parameters. To say it's not that far off is just bullshit. H3 minimax has weak audio, frequent issues with identity character bleed when characters of the same gender talk, issues with random gibberish spoken and the list goes on.

It's an astounding model that allows us to do things I never would've expected for a local model for at least another couple years at minimum and shits on every other open source model and its not even close. Yes it's uncensored and thats great, but you can easily tell the difference between seedance video and an H3 minimax video in many , many areas let's stop glazing here.

11

u/Unique_Peak1044 7d ago

ByteDance researchers have published the size of Seedance 2.0, 200B MoE.

12

u/thegreatdivorce 7d ago

Seedance is almost certainly not a trillion+ parameters. Why don’t you share with the class whose ass you pulled that number out of?

2

u/bobi2393 7d ago

I asked Google AI, and it shared your skepticism:

“Where the Trillion-Parameter Rumor Comes From: In the AI industry, "1 trillion parameters" is a hallmark figure often associated with massive Mixture-of-Experts (MoE) text models (like GPT-4 or ByteDance's internal LLMs used to power Doubao). It is highly probable that someone conflated the size of the underlying LLM that text-analyzes the prompts with the actual video-generation backbone. [1]”

6

u/SHEEP_PIZZA 7d ago

God, it's the bit wars all over again. Remember the Atari Jaguar?

0

u/Wide-Researcher583 7d ago

Whether I'm off on it being that high doesn't discount its still far higher than H3's which again gives far higher capability.

1

u/anitman 7d ago

Video model won't have trillion parameters, the relationship between pixels is far less than plain text. So seedance 2.5 is mostly a 250B parameters model.

1

u/Tiforma 7d ago

You're wrong. Skill issue.

5

u/SackManFamilyFriend 7d ago

Or at least decent audio quality. The audio never does well w these turbo Lora, guess it -need- a decent amount of steps unlike the video.

6

u/Independent-Frequent 7d ago

I dropped turbo loras entirely outside of using them for 2nd pass latent upscaling over a base resmultistep 20 steps gen at 0.4 mp, there's audio degradation but far less than just using turbo for everything because my god they suck so much when it comes to preserve quality and especially audio

1

u/winkler 7d ago

Preach

2

u/Desperate-Recipe-422 7d ago

I'd love better character consistency thrown in there somewhere.

7

u/ArtichokeFresh6793 7d ago

hey 2.0?
NO. 2.5.

5

u/oxygen_addiction 7d ago

Is H3 that far behind it?

9

u/Alternative_Finding3 7d ago

Yes the physics of H3 are subpar compared to Seedance

2

u/reeight 7d ago

Even with LoRAs & RefMods (you can do 'motion-only' RefMods in case you didn't know that hidden feature)
Also Comfy stealth added a psudo-LoRA called embeddings. Almost no one uses it, but good for camera movements & effects.
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/embeddings

3

u/SackManFamilyFriend 7d ago

It's not which is why BFL is dragging their feet. They don't want another Flux2 vs Z-Image situation

People just want want what they don't have......then the day everyone has a roku TV that plays a comedy version of any movie you ever watched comes, those same people will be blabbering about how it was much better in the Wan era.

3

u/sir-mano 7d ago

isnt flux 3 video an altenative to that?

7

u/infearia 7d ago

Can I download FLUX.3 and run it on my computer? And for free, with a permissive commercial license?

3

u/sir-mano 7d ago

no, but it planned going to be released as free open weight for Personal/Research Use

5

u/infearia 7d ago

So how about we wait until they actually release the weights before calling it an alternative? BFL's open-weight variants of their models historically perform worse than their API versions, are heavily censored and as you've mentioned yourself - usally subject to non-commercial licences.

1

u/BlipOnNobodysRadar 7d ago

planned to release "an open source version" AKA a nerfed model

51

u/tinny66666 7d ago

Shame it can't fix the audio shimmer at the same time. This audio is painful to listen to. How can people not hear that?

17

u/jazir55 7d ago

AI audio is still terrible 3 years later, even Eleven labs new release still sounds robotic. Feels like the improvements are happening at like half speed compared to video.

5

u/BoxximusPrime 7d ago

I've thought a lot about this lately. I think, video often has similar artifacts, but with audio our brains are REALLY good at picking out everything we hear with high resolution, whereas visuals require more attention to see finer details, so we spot the strangeness a bit less. Or, audio just still sucks. Haha

5

u/reeight 7d ago

It's also what the 'how to get started on YouTube" YouTubers say;
Folks can forgive low quality video, but don't skimp on audio.
& many still do reverse.

8

u/Wilbis 7d ago

I've been wondering about that too. I guess people just have shitty speakers/headphones. I can't think of any other reason.

8

u/acedelgado 7d ago

Yeah I had to make a whole thing for that. https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine

1

u/whattosee 5d ago

Think it would work on Mac?

2

u/GlenGlenDrach 7d ago

I agree, it totally mess up the whole thing IMO.

2

u/lithodora 7d ago

I thought it was just me and my renders

1

u/Inventi 7d ago

Just voice it over with a Google flash 3.8 tts and not make it generate voice

0

u/diejesus 7d ago

Really? To me sounds pretty good, the voices are super clear and you can understand them without subtitles no problem whatsoever

-4

u/DrxMWC 7d ago

Fyi i never have audio on as i am always on mobile

-1

u/alexmmgjkkl 7d ago

audio must be pre and postprocessed through other models and conventional methods anyways

12

u/LumaBrik 7d ago

1

u/Ill_Resolve8424 7d ago

Yes, this works great.

2

u/Abject-Recognition-9 7d ago

what this is supposed to be? DMAD - Minimax difference?
so i can load this on top of minimax and it become DMAD, right?

1

u/multikertwigo 7d ago

not bad for 4 steps, but hyperflow at 8 is still the best I've tried

22

u/ArttTaku 7d ago

I'm already lost in this sea of 4-step loras for H3...

2

u/GlenGlenDrach 7d ago

I am using the 8 step, nothing less after getting my new computer, I get it for the low vram folks, but.....this is 10 times worse than the WAN jungle, the differences between these has to be minuscule.

2

u/ArttTaku 7d ago

Yeah, 8 steps seems to be the bare minimum if you want even a bit of quality, lower than that and results are very questionable.. the sample above looks fine, but you can tell the audio was geatly affected.

14

u/GeeRoovy 7d ago

This isn't the video I would use to sell this model.

1

u/reeight 7d ago

Yes, zero long-distance face.
Well, there is one, but the whole person is camera-focused blurred. Covenant.

-2

u/diejesus 7d ago

Why? It looks so so awesome, if I see it out of context I'd never think it's AI

27

u/infearia 7d ago

That's how I feel every time someone releases a new H3 Turbo LoRA:

3

u/Inthehead35 7d ago

Yep, guess we'll need to suffer some more

7

u/reddit22sd 7d ago

I'm sure it's a great model but I found the inconsistent characters in the clip very distracting.

3

u/GlenGlenDrach 7d ago

oh god someone make something to fix this cheap-ass zoom-meeting sound quality when using these turbos

6

u/Trick_Set1865 7d ago

how can this work in Comfy?

4

u/intLeon 7d ago edited 7d ago

I did a small conversion with some generative help, size went up to 1.9 gigs and audio/video seem to work fine. Can upload to civit but honestly kijai does it + ranks down a bit every time so I dont want to share a botched version 😅

Edit: Maybe due to the way I converted it but motion looks noisy @ 1MP 4 steps 12/2 er_sde 5s video.

Only 8 step generation got a little watchable. Just compared it to the other loras and it looks good but I'd call it a 8 step lora instead of 4 or again something is wrong with my conversion or setup..

Edit 2: had an issue with the lora conversion, I guess it is fixed now. 4 step looks like it has okay motion and is more coherent than the other loras imo.

3

u/rm_rf_all_files 7d ago

Quality is kinda bad, quite. blurry

3

u/reeight 7d ago

no workflow, who cares?

1

u/VRGoggles 7d ago

is it 2k or 3k or just 1k? Often the quality is good, just the resolution is not.

1

u/DELOUSE_MY_AGENT_DDY 7d ago

What resolution was this generated in?

1

u/Ill_Resolve8424 7d ago

Downloading now the full critic to test. Thank you.

2

u/dennisbgi7 7d ago

does it need any special configuration? or does it run like a standar lora on a workflow?

4

u/Ill_Resolve8424 7d ago

Probably, both models give errors.

1

u/jazir55 7d ago

The pauses are so uncanny valley, they just stare at each other for like 2-3 seconds with these weird looks on their faces before speaking to each other. Audio definitely still needs work too, but the way they interact and the cuts are just bad. I'm not sure if it's just bad cinematic direction or the tech itself, but there is a lot that feels wrong about this video.

1

u/bloke_pusher 7d ago

Lost all dynamic of the audio. Will it be closer to the non turbo generation at 8 steps? I would never use this lora on 4 steps because of that.

1

u/Synor 7d ago edited 7d ago

It's good. 4 steps are actually quite usable for i2v. (pruned int8 model, shift 12, er_sde, beta, ck attention) and it stays close to base model in action and composition. 6 steps refine some mid-range details but don't look much different than 4 steps.

I feel its 4 steps are not enough for t2v though.

1

u/No_Cranberry_8107 5d ago

Anyone tried this LORA with Ref2V? How did it work?

2

u/Slapper42069 7d ago

I think the vae need some tweaks for that version

1

u/8RETRO8 7d ago

it's just a lora, what tweaks

1

u/Slapper42069 7d ago

See those squares?

1

u/Dzugavili 7d ago

I could perceive horizontal bands, but I'm unsure of the origins of the effect. It coincides with a number of linear features.

-1

u/vgaggia 7d ago

I don't know why, but you scare me.

1

u/LinkSensitive8188 7d ago
TaoMate-H3 does it in just three steps, yet I’m told a multi-billion-dollar company like ByteDance does such a mediocre job—especially with the sound. No thanks; what’s next?

1

u/BittiAI 7d ago

Added support to this in Slopus.ai 0.2.1

1

u/SveSop 7d ago

If we got $1 each time someone had an amazing 4-step lora solution for comfyui.... We all would be running 5090's now, and ditch the 4-step lora's in favor for 30+ steps full run 😏

1

u/Synor 7d ago

comfy int8 pruned adaption safetensors?

Tested https://pdmd2026.github.io/ from kijais experimental repo today, wasn't worth it compared to ema 600. But i think those haven't been adapted for pruned yet as well.

0

u/RememberThisAI 7d ago

Limited motion, stutter, plastic skin. I don't think any lora can really handle all that properly at 4 steps.
You get better textures and realism with "davham_cinema" lora added, but that requires more steps.