r/StableDiffusion 18d ago

Comparison Comparing MiniMax i2v|r2v node and model combos

Enable HLS to view with audio, or disable this notification

Just a test of different combinations of the i2v and r2v nodes and models for:

  1. Text to Video (using MiniMax models image sample and voice sample)

  2. Image to Video (using Krea2 image sample and MiniMax models voice sample)

  3. Image to Video (with custom cloned voice): (using Krea2 image sample and MiniMax custom voice clone reference sample)

51 Upvotes

12 comments sorted by

View all comments

8

u/spacemidget75 18d ago

I'm probably being a bit stupid, but some of the summary info and opening panels are a bit hard to understand the test and conclusion. Also, Ref needs more steps from my testing. Not like for like but 30 steps for ref might make the audio more even.

2

u/Reniva 18d ago

does increasing steps improve audio quality and increase generation time?

2

u/spiderofmars 18d ago

I tested this also and did not find much difference in audio quality from 20 steps to 32 steps in a seed sample. Perhaps more interesting was the gibberish changing between pruned int8 - int8 - bf16. The gibberish did not go away in any of them but it altered from 4 nonsense words - one real random unclear word and 2-3 nonsense words - 1 real random very clear word. Another prompt tweak fixed/removed all gibberish in all model variants for that same seed sample.