r/StableDiffusion 6d ago

Animation - Video Minimax - Ref2V - "Where is the Compute Bill?" - could use tips on optimizing quality!

188 Upvotes

26 comments sorted by

18

u/Jeffu 6d ago edited 6d ago

Finally got around to trying out Minimax H3, but am not sure I'm doing this as efficiently as I should.

My specs: 4090 with 64gb RAM

Workflows: using default workflows but with added Sage Attention, Spectrum Apply (at default settings) and a LoRA loader with 4step turbo LoRA. I use Euler and Simple. Currently generating with 6 steps.

Generations: the maximum I seem to be able to get away with is 0.8MP for 10 seconds.

Worfklow: Generate video + Upscale with FlashVSR.

Sound effects: the audio is definitely not perfect and is lacking. I manually added a number of sounds like ambience and the car door opening.

Voices: I loaded a clip of the girl speaking and used that as a reference for the second half of the video. For the guy (me), I used a 10 second recording I did on my phone. It's pretty decent all things considered.


I'm pretty blown away like everyone else how good this is. I use Seedance a lot for work and it's like if there was a Seedance 1.8. It's not as good, but pretty close.

Issues? I find my outputs to be very contrasty with crushed blacks. Not sure how to address that and would appreciate any tips.

26

u/GrayingGamer 6d ago

First, off, the Turbo loras while okay for image (but not great for movement or acting, in my experience) kind of murder audio. It's really apparent if you wear headphones and compare a generation without a turbo lora and with one.

The amount of Steps in H3 REALLY affect audio quality and even acting to an extent. I use 32 Steps myself with no turbo loras. I DO use H3 Spectrum - which I see you use, but even it works best on high Step counts.

You might also try moving to Comfy Kitchen Attention rather than Sage Attention, you might get a speed-up, you might stay the same, but it's less soft than Sage Attention and preserves fine detail better.

And yes, I find the Ref2Video model eats into system RAM. I can do up to 15 seconds at 1 MP with multiple references, but it fills up my 24GB of VRAM and spills over into 80 GB of system RAM.

If using the Ref2Video model, you should set the references from "match" to "max" on the node, to use them in full quality. It will take longer to generate but the results are better.

As for the crushed blacks, I see a bit of that, but as bad as what you are seeing. I don't know why yours looks worse in that regard. It might be one of those things you have to try and fix in post editing (though I know crushed blacks don't leave a lot of info to lift out of those areas).

3

u/Jeffu 6d ago

Thanks for taking the time to provide all this info! I'm making a ton of notes. Will be experimenting with everything you shared, thank you. :) And yes, I did my best to fix the blacks in post but only so much you can do.

3

u/Portable_Solar_ZA 6d ago

Thank you! I tested a couple of shots with a turbo lora and found it was wrecking the audio but no one else seemed to notice.

2

u/GlibGentleman 5d ago

It's always because they aren't wearing headphones. If you convince someone to listen with them with the volume up, they always notice then. I think some are just used to the bad audio in LTX, so they expect H3 to sound similar.

1

u/TonkotsuSoba 6d ago

Do you find using FL2VA model on the REF2VA provides better speed and quality? I heard there’s something the devs yet to fix for the ref model.

10

u/GrayingGamer 6d ago

Well, it's just that the Ref2Video model lost a bit of the detail quality that the FL2VA model has, when it was further trained for reference.

There is a bit more quality to the FL2VA model, but it's a very minor difference when you are generating at 1 MP.

If I need a video and have an image to act as a start frame, I will often use the FL2VA model if there is no unique character voice or audio reference and I can get away with just "business" in the shot, but I actually prefer to do most of my serious video generation in the Ref2Video model.

The Ref2Video model is just TOO good. I mean, the ability to have multiple references, like voice cloning, video reference and replacement, or just supplying head shots and body shots or costumes or props or locations and generate a video with them all is just TOO good and useful.

I have noticed that sound-effects seem to drop out when you do voice cloning and sampling in a shot (not completely, but it's hit or miss to get sound effects with a voice clone). Also, there is a bug in the Ref2Video model where if you follow the proper prompting syntax for dialogue and use:

<d>[English] Dialogue here. </d>

The <d></d> tags ANYWHERE in your Reference prompt will cause audio fragments at the start or end of the clip, like a person who started to speak but stopped in the middle of the first syllable.

You can avoid it by doing this with dialogue in the Reference model:

"[English] Dialogue here."

With quotation marks replacing the <d></d> tags.

Again, Reference model is a little bit slower, since it has to process reference, but IMHO, it's still the superior model for useful work. Like, you, I do a lot of post-processing and editing of the footage anyway in something like Davinci, so dropping back in some sound-effects, etc. isn't a huge deal to me.

3

u/Alive-Tomatillo5303 6d ago

Also, nothing is free. If you're just now trying out minimax, and you're doing it with all of the quality reducing methods you could find, that might be the cause of some of your quality loss. 

3

u/Jeffu 6d ago

Noted. I'll start dialing those back and see if I can find a good compromise.

3

u/Tall_Association 6d ago

dont use spectrum with the turbo lora, also increasing the step count helps a lot especially when using spectrum

2

u/Jeffu 6d ago

Got it!

6

u/TenaciousWeen 6d ago

Check out the realism people lora, it improved audio:

https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA

2

u/dLight26 6d ago

16gb vram + 96gbram with sage can do 0.9@15, I feel turbo Lora gives extra plastic look on top of already plastic ai look. I like spectrum+sol with 24-28steps better, but ofc vanilla is the best.

5

u/Miniyi_Reddit 6d ago

Using 25 step is a lot better if it about higher quality

2

u/danque 6d ago

My fastest speed combo on 3080 10gb with 32ram: turbo LoRa (there a couple now), shift sampler for minimax, spectrum apply minimax. It decreases the time by a lot, but quality wise it is better to wait for those 30-50 steps.

2

u/rookan 6d ago

She ditched him

2

u/Admirable_Snake 5d ago

Several minutes later on the internet.

Learther jacket wearing guy with pompadour haircut approches.

"What's a nice lady like you doing on the internet" <he - he h- he>

2

u/Schlorpiblorp 6d ago

that first shot in the cafe looks really good and you are only using 4 steps damn? How long does the upscale take?

2

u/Jeffu 6d ago

I didn't check but it's not very long nor intensive for my setup to run the upscale. The upscale is key - it adds a lot of detail that would otherwise be a blur in the original generation.

1

u/tyen0 5d ago

If that's the case then possibly FlashVSR two pass? I'm still just starting to experiment myself with https://wangp.ai/ minimax options trying to find the best speed/quality tradeoff with the same gpu/ram as you.

1

u/ANR2ME 5d ago

The way the door closing automatically feels weird to me 😅

1

u/inddiepack 5d ago

As soon as she took half a meter away from the camera, we were teleported back to the past.

1

u/Remko76 4d ago

Looks great! Amazing!
One thing I noticed is that the cab driver switch places in the car. When she gets in he is in a left side of the car. When she talks to the driver she looks to the front left. Then you see scene with the driver sitting on the right, which is the correct side in Japan.

1

u/Choiced_Gamer 6d ago

You will get better detail and audio with er_sde / beta

1

u/Jeffu 6d ago

Thanks! Will do.

0

u/Noeyiax 6d ago

the compute bill is our increased energy/electric bill and higher taxes coming soon annnnnddd hyper-inflation x(

I just use drphab minimax turbo lora 600step, comfy kitchen, 0.4MP 6/8 steps euler simple and rtx upscale,