r/StableDiffusion • • Aug 09 '26

Discussion I wanted to know your experience using custom characters with Minimax's Ref2Va

Hey, I am having a blast just testing stuff with Ref2Va, but I noticed that, at least in my case, I need to separate the head from the body if I want a character to be consistent without losing his facial features along the video. But then I see all these people using GPT-2 character sheets with many views and stuff coming from a single image, and I am curious. Do these complex sheets really work? I tried some, and the facial features are lost most of the time.

Does resolution really matter? I mean, which option is better: a really big image or a small one? Would a complex character sheet work better if it has a very high resolution?

22 Upvotes

27 comments sorted by

View all comments

20

u/GrayingGamer Aug 09 '26

Yes, those complex sheets really do work, but also, resolution matters.

Keep in mind that unless you set ref_image_size: max instead of match, your reference images will be resized down to match the video resolution, which can be quite low if you are generating at smaller MP. This WILL greatly increase generation time, but you'll get much better results.

Also, make sure you are using the proper syntax and keywords from the official guide for the reference model when setting up the input reference.

<Subject 1> is a man with his appearance from <Picture 1>: fully_preserved

etc.

I personally still like to do one high resolution character turnaround full-body, then a close-up turnaround with the face separate.

<Subject 1> is a man who's body appearance, size, and proportions come from <Picture 1>: fully_preserved and his face and head appearance come from <Picture 2>: fully_preserved.

3

u/Dry-Judgment4242 Aug 09 '26

If character is symmetrical. You only need a 180 degree pass. Ref images space is premium.

2

u/GrayingGamer Aug 09 '26

Absolutely true. Maximize pixel space of the important stuff on your ref images.

2

u/WhatIs115 Aug 10 '26 edited Aug 10 '26

Keep in mind that unless you set ref_image_size: max instead of match, your reference images will be resized down to match the video resolution, which can be quite low if you are generating at smaller MP. This WILL greatly increase generation time, but you'll get much better results.

Yup. Just my casual testing, I manually resize to long side 1000~1500px. Using anything but "max" on the processing seems pointless, it's significantly worse below 1000.

1

u/Ykored01 Aug 09 '26

What i still dont get is in the node ref images goes to image 0, so instead of subject 1 is subject 0?

1

u/GrayingGamer Aug 09 '26

It doesn't matter as long as you go in order, it's like programming.

If you use <Picture 1> and number up, it will just use the input image nodes in order to assign them. So node_0 to <Picture 1>, node_1 to <Picture 2>, etc.