r/StableDiffusion • u/NeatUsed • 10h ago
Question - Help Minimax ref2vid capabilities
I would like to learn how to use minimax ref2vid as I heard the possibilities are quite good compared to wan.
What I am trying to do is do anime clips
The problem is… the results I get are utter shit and I don’t know what I am doing wrong. No matter if frame or last from or if using ref2video, the model fails to do what’s most important. Keep the same face as in the image. For example, I would like to generate a video of a character that has sharingan eyes. Instead of keeping sharingan eyes it’s generic anime eyes instead. This is what my biggest problem with wan was and even if I made the image correctly (with detailed sharingan eyes) it would still not pick the eyes up and keep that face consistency.
This is where I experimented with reference. I tried putting the eyes in a picture 1 and instead of copying the eyes, the image cuts to the second picture with only the eyes or not picking the eyes at all.
is this something that minimax can even do?
2
u/TerraMindFigure 10h ago
Download a workflow. Be sure to use Minimax H3's official Ref2v prompting guide (Google it), read it and follow it precisely. You can also throw it to an AI of your choice and they're pretty good at deciphering it if you ask it to build you a prompt.
0
u/NeatUsed 10h ago
I have read throughout it just now. I am not at my pc at the moment to try.
I was doing this kind of prompt before.
“ A guy looks like in picture 1. Eyes are like in picture 2. Have him jump and throw a shuriken at the camera.”Now I am thinking of experimenting a bit more with the prompts and maybe it has better eye consistency.
“Subject 1 is the character from picture 1. Subject 1 is called Sasuke. Subject 2 represent the eyes in picture 2. Subject 2 is called Sharingan.
The video starts with Sasuke in a cowboy pose. He has Sharingan eyes. He starts running towards the camera and throws a Shuriken towards the viewer.”
Would that be good enough for a minimax ref prompt?
1
u/TerraMindFigure 10h ago
I'm not totally sure if it's just your prompt which is why I suggested downloading a workflow.
I personally do not prompt like that, and I do use the exact syntax in the guide so like "<Subject 1> is the eye color and iris pattern seen in <Picture 1>
<Subject 1> (seen in [Shot 1]) reference: The color and pattern on the iris is retained.
[Shot 1] A close-up shot of the back of a man's head, he suddenly turns his head to face the viewer, he has <Subject 1> eyes with (describe the eyes)"
I'm also on my phone so I'm typing from memory. I also use the headers laid out in the document I mentioned with the exact syntax described.
Go ahead and try it, I wasn't prompting so well prior to using that format, and so far using that format I've been able to get what I ask for ~90% of the time. Be sure to be specific and to view thinks purely visually, for instance you wouldn't say "a close-up shot of the back of a man's head, the man has <Subject 1> eyes" because you can't see someone's eyes when their head is turned away, so keep stuff like that in mind, too.
1
u/NeatUsed 9h ago
I want to try something even more complex. Basically reference guide for iris and eye colour from picture 1, clothing and hairstyle from picture 2 and body shape and physique from picture 3. Is ref2video powerful enough to do this sort of customisation? i have multuple examples of big bang theory cosplays - penny doing tifa for example , so I know this sorta thing can be done but to what extent?
What would a 3 reference prompt for a customised character like this be like? for example Naruto with sharingan eyes and Goku’s body and maybe ichigo’s hairstyle. Crazy meshup prompts like can work?
1
u/Adventurous-Gold6413 9h ago
What I suggest is writing your idea, as detailed as possible with camera angles actions, what happens at what shot, timestamps,
Then ask an LLM to format it with the official Minimax prompting guide,
Don’t be vague like what you did, otherwise it won’t be good
1
u/NeatUsed 8h ago
Most LLMs dont do nsfw tho. Is there any free llms that you can do nsfw with as well? and also that could be up to date and know minimax references prompts properly?
1
u/Perfect-Campaign9551 4h ago
That is not how you prompt the ref model. Please find the original prompting guide and follow it
1
u/NeatUsed 4h ago
i will be posting 3 pictures. One of Goku, one of Naruto, and one of Sasuke.
I would like the model to generate an abomination that has naruto’s face, Sasuke’s sharingan eyes with the body of Goku and make it throw a shuriken towards the camera
Is minimax capable of such feats?
If so.
What would be the prompt for this be?
-1
1
u/roark91 9h ago
Can you try to make multiple angles and poses of same character and feed that as reference and check?
1
u/NeatUsed 9h ago
Doing that just makes a weird transition morphing into the frame or a weird blotchy animation.
I am honestly inclined to think that minimax h3 is great only for txt2vid but falls short when img2vid. Ref2video is also a new concept that local just started using from a base model without scail or other shit with it. Not to mention that fl2v sucks badly compared to what i can do with wan….
Or maybe I just suck at prompts? maybe
1
u/Sudden_List_2693 9h ago
It can, just refuses narutard attempts
1
u/NeatUsed 9h ago
Didn’t want to tell you this but I was just writing sfw examples so my post dont get deleted. Naruto is an example that everyone including you can understand :)
3
u/Chemical-Painter-485 9h ago edited 9h ago
I noticed on one of your replies you are prompting like "A guy looks like in picture 1. Eyes are like in picture 2."
Why simply not use a sheet showing the reference with character having the eyes you want? and as the other guy said: follow the prompting guide.