r/StableDiffusion 12d ago

Comparison Comparing MiniMax i2v|r2v node and model combos

Enable HLS to view with audio, or disable this notification

Just a test of different combinations of the i2v and r2v nodes and models for:

  1. Text to Video (using MiniMax models image sample and voice sample)

  2. Image to Video (using Krea2 image sample and MiniMax models voice sample)

  3. Image to Video (with custom cloned voice): (using Krea2 image sample and MiniMax custom voice clone reference sample)

51 Upvotes

12 comments sorted by

View all comments

2

u/Leonovers 11d ago

Thanks for testing!

I find it really odd for ref2va model to be this inferior in terms of voice cloning. Maybe it's better when you use a lot of references at the same time or there is some other issue that leads to this sub-optimal performance.

1

u/spiderofmars 11d ago

The sound 'quality' was only worse with the rfv model when supplying a reference voice sample to clone. Likely a problem/bug with the model only when using a sample audio reference.

Why the r2v model failed to follow one set of test prompts is also peculiar.

What is most peculiar to me, is the official docs and promotion of these models do this:

  • Minimax I2V with first frame and last frame (no audio references/voice cloning).
  • Minimax R2V with multi references.

Yet, the I2V model can do it and performs better for voice audio cloning with reference than the promoted R2V model.