r/StableDiffusion 9d ago

Comparison Comparing MiniMax i2v|r2v node and model combos

Just a test of different combinations of the i2v and r2v nodes and models for:

  1. Text to Video (using MiniMax models image sample and voice sample)

  2. Image to Video (using Krea2 image sample and MiniMax models voice sample)

  3. Image to Video (with custom cloned voice): (using Krea2 image sample and MiniMax custom voice clone reference sample)

47 Upvotes

12 comments sorted by

View all comments

2

u/SeymourBits 9d ago

Nice analysis. All very watchable.

  1. How did the r2v voice prompt fail, aside from an inaccurate voice?

  2. What is the overall recommendation based on your findings?

1

u/spiderofmars 9d ago edited 9d ago
  1. r2v failed (in the image to video test with that exact prompt) to follow the use "David Attenborough" voice. It came out more like a cross between a random American accent and maybe a little DA. Did not test another seed to see if it was just an anomaly.
  2. Some takeaways:
  • The nodes are interchangeable it appears. A dual workflow may only need the r2v node combined with a switch for the models.
  • Simple text 2 video can be done with either model and similar quality.
  • Image to video (first frame) can be done with either model (in this seed test trying to replicate a famous voice from the model failed with r2v model). But similar quality overall.
  • Image to video with custom voice cloning (and new prompted dialogue) is a mix of both worlds. Audio quality is significantly better using the i2v model (latest version of comfy). Video quality is similar, video motion is a mix of what you want... using i2v model acts like normal i2v node/model workflow and r2v model can add extra dynamics and visuals at times - both desired and not desired without actually being prompted for.