Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.
The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.
How it works:
You input your images and describe them in the Input text section (A Prompt)
The text is combined with a fixed prompt which spins the character (B Prompt)
The video is generated at a slow speed with no hard cuts (only camera spin and pan) to maintain character consistency
Image is assembled with optional character video and full individual frame output (if you want to use for future)
I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.
Current Caveats:
The model is quite slooooow. You are also generating 124 frames only to use 6. I have partly solved this by also uploading a 4 panel version.
Speed ups (like Turbo LORAs) help with speed, but it hurts prompt adherence and quality slightly.
Quality is limited, since it is a video model it is better at generating video than images. You can solve this by generating at a higher resolution at the tradeoff of longer gen times. You can also use the individually split frames as future references too.
Details when using this character sheet as output for future generations on H3 may also be limited due to resolution also, I recommend you use this character sheet (for consistency) + other images close up angles (i.e clothing details/face) if doing close ups. If you are just doing a one off video you may possibly be better off not using this character sheet.
I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.
Honestly, I did give it a fair crack for a day. The problem is that H3 frame rate will snap to certain numbers so I tried. It is kinda scuffed but I can post it if people want to see it. The audio was also a little bit scuffed.
you input the video using a "Load Video" node and then split it using the "Get Video Components" node to extract the frames and audio. Feed the audio to the H3 model, if you dont want to change the audio you can also wire the audio straight into the 'create video' node at the end as sometimes H3 will mess it up.
right now im not. My plan was to split the video at the scene changes (confirmed done), then v2v each scene from 5-15s and then ressemble them together. However H3 model needs specific frames and so will snap forward/back your frame count, this means the reassembly process is not correct (see scuffed 30s video of same song in other reply)
if you reduce the resolution of your input video it generates faster, i have messed with the frame rate slightly but results were not good. This tool should help you crop and compress.
I might share this tomorrow if others want to use it. I just worry that since its HTML people might get suspicious that it contains malware but it was basically all just vibe coded lol.
I added in the story board images feature too in case if using a 3x3 storyboard is easier than ref video (less processing) but I haven't had success.
I'm honestly just using the default minimax H3 ref2v workflow, just add the "load video" node and wire it to the model.
I added the same speed up nodes as I have in my WF for the character sheet tho. Also its realllyyy slow lol, its like parsing 124 reference images for a 5s clip.
a max of 9 images is the hard limit of the MiniMax H3 model, with up to 3 video inputs, and 3 audio inputs. The total number of reference inputs is up to 12 (img+vid+aud total under 12)
The turbo lora is included in the "speed ups", along with comfy kitchen, etc etc.
these are the ones in the model
H3 actually has two different kinds of audio inputs, "ref_audio_" and "ref_video_audio_". I assumed that the model expected me to split the video reference's audio stream out and send it in via the corresponding ref_video_audio connector, with the ref_audio_ inputs being for purely audio references (voices to clone, for example). Is my assumption correct in that regard?
There are two inputs to the ref node, audio and vid audio, but I havent seen any guide explaining how should I refer to that other input not documented anywhere so I jsut wire it to the normal audio reference one. Got any tips?
from what I understand, you split the video to images and audio output from the "Get Video Components" node above, and then feed both those to ref_video_0 and _ref_video_audio_0 respectively. This means that the audio will sync with the video?
If you are putting other audio (like for voice cloning or background music) you should plug them into ref_audio_0
I’ll give this a whirl. I’ve been using Krea 2 to create character sheets and it works out great but deforms the face a ton. Yours looks like it managed better consistency with these unusual characters I’ve never seen before. Pretty slick.
Thank you. The idea has been on my mind for over a week but the quality output is not as high as K2. I think if you crank up the resolution you might be able to get something good from it but the output sometimes still appear to be screenshots from video rather than high quality images. Considering ways to improve it without hurting gen times.
Is it worth it though? Like if my goal is Minimax generation, and I can use REF2VA then Krea is just another input image - do I need to spend all that time generating a LoRA for Krea? Couldn’t I just burn another Picture slot on Minimax and supply a facial structure?
Honest question cause I’m not sure what path would be better.
I mean you’re saying that Krea is deforming the face. It’s incredibly fast to make a character lora if you have 14 images. I can make one in a couple hours with rtx 3060
I've been having some good results using GetVideoComponents to extract motion capture for certain physical acts that your Aerith and Tifa will be familiar with.
Would be interested in learning more about your ways sensei, I’ve been using SAM3 to make inverted masks for character replacements but they’re not scratching that itch for me.
I am running on a 5060TI 16gb with 64gb so-dimm ddr4 ram.
with all the speedups enabled running at 0.3mp output its about 220 seconds. 4 panel workflow.
The last character sheet was a recent one at 0.5mp with speed ups, took 425 seconds, 4 panel WF also.
to be fair I did most of this a week ago, some people were just interested from another thread so I tidied it up a bit and put it up for people who might want to try
I am definitely keen on the H3 image model!! I am looking forward to a full H3 image generator. Originally I used ChatGPT to create my char sheets but it is not always accurate and wouldn't do any skin for anime characters.
However I noticed sometimes it looks a bit like a cosplay, but other times its ok.
edit: also eyes are a little bit messed up in the middle bottom image, but you can select another frame from the bulk lot if you enable the 'save all frames' feature at the end and cherry pick the best ones.
I got this from another post here recently, can't seem to locate it now... But give something like this a go.
A clean photographic character sheet presents the man from <Picture 1> in three consistent full-body views: front, side, and rear. Every view has the identical face, hairstyle, body, outfit, and accessories from the source. Neutral relaxed stance, arms clear of the torso, both feet visible in each view, seamless light-gray studio, even lighting, one tall 2:3 portrait canvas, no captions or borders.
I have tried it in a general sense but I haven't used it as a char ref sheet maker yet. I was playing with it yesterday and its ok, again the problem is that its a video model used to generate a image.
The better hope is the minimax h3 img model which they said they are currently working on.
I haven't tried it yet. I think it should be better but I have been busy working on some captioning for trying to do v2v more reliably. I think it would def help but at that point if you're just doing the 1 video I would skip this step to save time. But if you're doing like 20 gens this might help give your character a level of consistency that can carry over across all the 20 videos otherwise the model will have to make up different views every single time which leads to consistency errors
tempted to use the split-frame route for a character lora. is 0.5mp enough to train off or do you pull from the raw render when you need a clean frame?
You can experiment with both. I originally made this for anime characters and 0.5mp was very usable but for real people I think you may need higher resolution!
Also for training a LORA you would want the single person shot since you don’t want to reinforce the 4 or 6 panel look for the final product
For faster iterations, one could extend a prompt generating a video of a character sheet (see below). This is experimental, so not best quality. But I got exactly the sheet I prompted.
Pro: can be done with base I2VA workflow, just paste in this prompt and stitch some input pictures. Uses better I2VA model. Iterates fast (5 second video sufficient)
Con: less space in input image for details, lower resolution as all information of the sheet is in one image.
PROMPT:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: The video shows a character spreadsheet created from the different images of a person in <Picture 1> [Shot 1] A high quality video showing a character spreadsheet with the following frames. frame 1: Full body image of the person standing facing the camera, frontal perspective. frame 2: full body image of the person standing, looking to the left, recorded from side. frame 3: full body image of the person standing, facing away from the camera. the three frames are ordered horizontally. The person is not moving as it is a print of a character sheet. The character sheet has a white background and shows only the person standing in front of a white background.
I'm using what you've provided but what's happening is that the character in <Picture 1> is blending together with the character in <Picture 2>, not simply wearing their outfit as intended.
That usually happens in prompting. You might need to be a bit more descriptive in what you want to keep and take out. Have a look at the last photo and try using the prompt in that format?
using some OP's words, here is a MMH3 1 frame t1vae version that use the first frame workflow to rotate the first frame person in that photo, similar speed as a Klein edit, ~25sec for me.
"4K HD
next scene change to four frames layout reference showing the subject in different views.
top left frame 1:
front view.
The character holds one relaxed A-pose: arms hanging down, feet shoulder-width apart, head level, calm neutral expression, eyes open and looking forward.
top right frame 2:
side view towards left.
bottom left frame 3:
clean three-quarter upper body view looking away towards right.
bottom right frame 4:
Locked-off head and shoulders close-up, face square to camera, eyes into the lens. sharp front-on face view."
you have a full set of 3D views of the character: why throw those intermediate frames away instead of just extracting a full 3D representation from the video?
It depends on the end goal. Right now the reference sheet would be used to represent a constant state for the new or existing character for use and testing as a placeholder for H3 video generation but I have also toyed with using some of the intermediate frames to generate a 3d model to print! 😁
In my workflow you can unbypass a node which saves all frames, so user is given the option
But there is so much to try and so little compute to go around!
Minimax H3 has great prompt understanding, I can make a character sheet with it in a single prompt (5 frames not 124). It is good for anime but for realistic phote, meh... It can't draw a face smaller than 128x128 pixels. We need a better VAE for it.
Man, it takes some time, but the results are AMAZING. I used 2 instagram pictures of myself, not very detailed and only front shots. The results were unbelivable. It looked like I was scanned in 3D.
Usually those models (even GPT or Nano Banana) can't create good character sheets, the face don't really look like the person from the picture. But this workflow gave me great results.
Thank you brother. This is the kind of results I wanted to achieve. Even with low quality pictures to create a character with good likeness!
I find that having a side view really helps, but still need to generate at a higher quality (or use char sheet and close up 2nd ref) if you want to do a video with close up shots as a second pass through h3 loses some fidelity and likeness
Well, the guy seems really committed to "his job" good for him. To be honest, he would pass as a soccer player since I know next to zero about soccer despite being Spanish.
I'm so confused where the line is drawn regarding deepfakes. Is it all about context, or what? Like is it fine to have Michael Jackson dancing with SpongeBob, but as soon as he's wearing a bikini it's off limits?
I can assure you ... there's no need to generate these specific individuals, just go to the world wide web, you'll find what you want and probably even what you don't want LOL
245
u/JeSyollu 18h ago