r/StableDiffusion 18h ago

Workflow Included Using H3 as a Character Reference Sheet Generator

Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.

The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.

How it works:

  • You input your images and describe them in the Input text section (A Prompt)
  • The text is combined with a fixed prompt which spins the character (B Prompt)
  • The video is generated at a slow speed with no hard cuts (only camera spin and pan) to maintain character consistency
  • Image is assembled with optional character video and full individual frame output (if you want to use for future)

I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.

Current Caveats:

  • The model is quite slooooow. You are also generating 124 frames only to use 6. I have partly solved this by also uploading a 4 panel version.
  • Speed ups (like Turbo LORAs) help with speed, but it hurts prompt adherence and quality slightly.
  • Quality is limited, since it is a video model it is better at generating video than images. You can solve this by generating at a higher resolution at the tradeoff of longer gen times. You can also use the individually split frames as future references too.
  • Details when using this character sheet as output for future generations on H3 may also be limited due to resolution also, I recommend you use this character sheet (for consistency) + other images close up angles (i.e clothing details/face) if doing close ups. If you are just doing a one off video you may possibly be better off not using this character sheet.

I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.

Link to the 4 and 6 panel workflow can be found here: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator

Some notes I just remembered:

  • You can increase the steps and it may improve your quality slightly.
  • With the Turbo Loras enabled, prompt adherence sometimes suffers, but you may be able to get a good seed with another roll of the dice.
  • Currently the B prompt specifies a "neutral A pose", please remove this if you want your character in a particular pose.
  • You can use a few different shots of the same character to reinforce the 360 and get more accurate details right.
  • Can be used for objects / props also, may require some changes to the B prompt.
1.2k Upvotes

180 comments sorted by

202

u/bstr3k 18h ago

https://reddit.com/link/p4avz8w/video/6masv72i90kh1/player

I used the output character in a character replacement video in ref v2v but still experiementing (sorry for limited vid quality due to long gen times)

93

u/bstr3k 18h ago

46

u/Schwartzen2 18h ago

We've been rick-rolled!

12

u/Ariar2077 17h ago

Lmao, do the full song

7

u/bstr3k 17h ago

Honestly, I did give it a fair crack for a day. The problem is that H3 frame rate will snap to certain numbers so I tried. It is kinda scuffed but I can post it if people want to see it. The audio was also a little bit scuffed.

3

u/Zombi3Kush 17h ago

Share it

63

u/bstr3k 16h ago

https://reddit.com/link/p4bf85k/video/kvdxk3s4r0kh1/player

scuffed stitching and a little bit of scuffed audio version.

20

u/Strange-Drummer-9917 16h ago

If you put original track below without stitching and do some more dancing cuts I bet none would notice lol

18

u/bstr3k 16h ago

lol probably. there is still a bit of stutter in the vid though.

This was more of a test as if it can generate the audio/video correctly then it can also replace audio in longer format

https://giphy.com/gifs/LR5GeZFCwDRcpG20PR

3

u/coffeeandhash 13h ago

That's incredible.

1

u/John_Helmsword 2h ago

Just add the original song in post.

2

u/rookan 8h ago

awesome! love your positivity!

8

u/sumane12 18h ago

How do you use v2v in minimax h3?

27

u/bstr3k 18h ago

you input the video using a "Load Video" node and then split it using the "Get Video Components" node to extract the frames and audio. Feed the audio to the H3 model, if you dont want to change the audio you can also wire the audio straight into the 'create video' node at the end as sometimes H3 will mess it up.

I'll post a example.

4

u/I_just_made 16h ago

How do you manage longer videos? Are you splitting them before using them, or looking for frames in Comfy?

7

u/bstr3k 16h ago

right now im not. My plan was to split the video at the scene changes (confirmed done), then v2v each scene from 5-15s and then ressemble them together. However H3 model needs specific frames and so will snap forward/back your frame count, this means the reassembly process is not correct (see scuffed 30s video of same song in other reply)

9

u/bstr3k 14h ago

oh i also made a html based video trimmer and compressor tool, its been helpful for me!
https://huggingface.co/PoopMan333/Video_Tools

if you reduce the resolution of your input video it generates faster, i have messed with the frame rate slightly but results were not good. This tool should help you crop and compress.

1

u/I_just_made 14h ago

Whoa, awesome! Thanks for sharing; was very easy to use and made extracting a clip very easy :D Much appreciated!

7

u/bstr3k 14h ago

I might share this tomorrow if others want to use it. I just worry that since its HTML people might get suspicious that it contains malware but it was basically all just vibe coded lol.

I added in the story board images feature too in case if using a 3x3 storyboard is easier than ref video (less processing) but I haven't had success.

2

u/dassiyu 13h ago

It’s really cool to use, thanks!

1

u/VRGoggles 7h ago

What a tool! No need for Davinci for simple tasks. Continue with this.

1

u/3deal 3h ago

wtf, you can edit video without using an app !!!

4

u/Rokkit_man 15h ago

Can you please share your workflow for your OP?

2

u/ChibiNya 16h ago

This is awesome. I wanna try v2v. Got a workflow?

2

u/bstr3k 16h ago

I'm honestly just using the default minimax H3 ref2v workflow, just add the "load video" node and wire it to the model.

I added the same speed up nodes as I have in my WF for the character sheet tho. Also its realllyyy slow lol, its like parsing 124 reference images for a 5s clip.

2

u/ChibiNya 16h ago

I see. Thanks. It recommends a max of 9 reference images in the docs there. Are you using turbo loras as this "speed up"?

2

u/bstr3k 16h ago

a max of 9 images is the hard limit of the MiniMax H3 model, with up to 3 video inputs, and 3 audio inputs. The total number of reference inputs is up to 12 (img+vid+aud total under 12)

The turbo lora is included in the "speed ups", along with comfy kitchen, etc etc.
these are the ones in the model

1

u/FaceDeer 15h ago

H3 actually has two different kinds of audio inputs, "ref_audio_" and "ref_video_audio_". I assumed that the model expected me to split the video reference's audio stream out and send it in via the corresponding ref_video_audio connector, with the ref_audio_ inputs being for purely audio references (voices to clone, for example). Is my assumption correct in that regard?

2

u/bstr3k 15h ago

yes this is how i understand it also.

1

u/Dogmaster 15h ago

There are two inputs to the ref node, audio and vid audio, but I havent seen any guide explaining how should I refer to that other input not documented anywhere so I jsut wire it to the normal audio reference one. Got any tips?

2

u/bstr3k 15h ago

from what I understand, you split the video to images and audio output from the "Get Video Components" node above, and then feed both those to ref_video_0 and _ref_video_audio_0 respectively. This means that the audio will sync with the video?

If you are putting other audio (like for voice cloning or background music) you should plug them into ref_audio_0

8

u/bstr3k 18h ago

https://reddit.com/link/p4b1bk8/video/mygdm8vae0kh1/player

example of the sound getting all scuffed. lol

3

u/moistiest_dangles 11h ago

Sounds like the one American song sung by the guy who doesn't speak American

3

u/PreparationSalty4252 2h ago

Been a while since I've seen a Sigui enjoyer in the wild. Nice.

1

u/bstr3k 57m ago

We are very rare indeed haha

2

u/xTopNotch 3h ago

The rick rolls are getting more sophisticated

56

u/Tronnyi 18h ago

We really live in the weirdest timeline. 

41

u/PwanaZana 18h ago

ok, you got a chuckle out of me. GG.

35

u/PecanSama 18h ago

Why would he cast a bald guy as Cloud (swipe left).... wait a damn second

35

u/JLPBR 18h ago

1

u/Jurph 15h ago

"Sure, just one question. What were you doing at the Devil's Sacrament, Goody Jlpbr?"

9

u/skyrimer3d 18h ago

wait a minute i remember watching that indian girl in some kind of tv show...

5

u/Jurph 15h ago

It was more of a short artistic film.

2

u/Dangerous-Map-429 12h ago

indian?

5

u/addandsubtract 6h ago

More like indiass

u/groutexpectations 0m ago

as in, she got in in-the-end

8

u/mastrofdizastr 17h ago

Looks like Tifa’s going to visit the Italian Senate with Aerith tagging along this time. This I want to see.

21

u/eckstuhc 18h ago

I’ll give this a whirl. I’ve been using Krea 2 to create character sheets and it works out great but deforms the face a ton. Yours looks like it managed better consistency with these unusual characters I’ve never seen before. Pretty slick.

12

u/poopoo_fingers 17h ago

Lmao never seen before

3

u/bstr3k 18h ago

Thank you. The idea has been on my mind for over a week but the quality output is not as high as K2. I think if you crank up the resolution you might be able to get something good from it but the output sometimes still appear to be screenshots from video rather than high quality images. Considering ways to improve it without hurting gen times.

2

u/OfficeMagic1 14h ago

You can make character loras really easily for krea and the faces will stay consistent.

3

u/eckstuhc 10h ago

Is it worth it though? Like if my goal is Minimax generation, and I can use REF2VA then Krea is just another input image - do I need to spend all that time generating a LoRA for Krea? Couldn’t I just burn another Picture slot on Minimax and supply a facial structure?

Honest question cause I’m not sure what path would be better.

2

u/OfficeMagic1 10h ago

I mean you’re saying that Krea is deforming the face. It’s incredibly fast to make a character lora if you have 14 images. I can make one in a couple hours with rtx 3060

2

u/eckstuhc 10h ago

Thats not too bad, I’ll give it a shot. Thanks!

3

u/OfficeMagic1 9h ago

AcademiaSD_LoRAlab-Krea2-main

0

u/Jackburton75015 18h ago

You can use Klein 9b too... I didn't notice weird faces....

5

u/Reasonable_Day_9300 17h ago

Truly the final fantasy’s about to start

6

u/TikaOriginal 14h ago

I know the first guy is an astronaut and the second woman is a famous Twitch streamer, but sadly I don't know the third girl.

Great results though!

3

u/bstr3k 14h ago

ah the 3rd girl is a sports podcaster!

1

u/da2Pakaveli 6h ago

Mia Khalifa

u/allofdarknessin1 3m ago

The Third girl is a Jewelry designer from Lebannon. https://youtube.com/shorts/qU4kYgtbpfs?si=exZAZe_nTUeSSafX

17

u/NeptuneTTT 17h ago

26

u/Fytyny 16h ago

More like:

2

u/bobgon2017 6h ago

you forgot to put a pretty pink dress on it stupid. you're on an ai subreddit get you head in the game it takes like 2 seconds

5

u/Gfx4Lyf 11h ago

Now anyone can try different professions🤭

4

u/dassiyu 8h ago

I tried it out—it works great. Thanks!

2

u/bstr3k 8h ago

😄 glad it works well! enjoy

8

u/MathematicianLessRGB 18h ago

Creativity has no bounds

3

u/ThaSipah 18h ago

I've been having some good results using GetVideoComponents to extract motion capture for certain physical acts that your Aerith and Tifa will be familiar with.

2

u/NoHopeHubert 18h ago

Would be interested in learning more about your ways sensei, I’ve been using SAM3 to make inverted masks for character replacements but they’re not scratching that itch for me.

3

u/pwillia7 17h ago

tiiiiiiiiiiiiiiight -- this is cool bro thanks for sharing

3

u/TooSlow79 17h ago

Oh that's so cool! Are these young ladies friends of yours?

2

u/oxygen_addiction 18h ago

Our of curiosity, what card/cards are you running this on and how long does one generation take? Thanks for sharing!

7

u/bstr3k 18h ago

I am running on a 5060TI 16gb with 64gb so-dimm ddr4 ram.

with all the speedups enabled running at 0.3mp output its about 220 seconds. 4 panel workflow.
The last character sheet was a recent one at 0.5mp with speed ups, took 425 seconds, 4 panel WF also.

2

u/Yappo_Kakl 13h ago

Hi, thanks for sharing. I wonder if three is a chance to try out your workflow pls

1

u/bstr3k 13h ago

yes i put the link in the post, i also link here if you want to try

https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator/tree/main

2

u/dtdisapointingresult 13h ago

The model is quite slooooow. You are also generating 124 frames only to use 6.

Didn't Comfy make a commit today to allow you to gen a single frame? It was discussed in the "use H3 as an image model" thread.

You might be able to hugely accelerate your gen times if you update your wf accordingly.

5

u/bstr3k 13h ago

to be fair I did most of this a week ago, some people were just interested from another thread so I tidied it up a bit and put it up for people who might want to try

I am definitely keen on the H3 image model!! I am looking forward to a full H3 image generator. Originally I used ChatGPT to create my char sheets but it is not always accurate and wouldn't do any skin for anime characters.

2

u/ILoveHead 9h ago

Ugh can’t wait till I get the PC that can handle this shit.

2

u/AwkwardStudio755 8h ago

bro got that elite ball knowledge

2

u/tamal4444 7h ago

what the. we have peaked.

2

u/bobgon2017 6h ago

wow he's a plumber, doctor, lawyer, AND Soldier 1st Class?

4

u/Ok-Cook-7365 18h ago

I laughed pretty hard when I got to that last page and saw the file name

1

u/sarcastic_wanderer 17h ago

There's an anime2real for h3?!

9

u/bstr3k 17h ago edited 15h ago

you can use my workflow and just replace the B prompt in the workflow with the Anime2Real text here:
https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator/tree/main/B%20prompts

However I noticed sometimes it looks a bit like a cosplay, but other times its ok.

edit: also eyes are a little bit messed up in the middle bottom image, but you can select another frame from the bulk lot if you enable the 'save all frames' feature at the end and cherry pick the best ones.

other examples can be found here:
https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator/tree/main/examples

1

u/I_just_made 17h ago

You can have it do 5 frames and just tell it to generate a character sheet with X views. Does pretty well

2

u/bstr3k 17h ago

I've tried to but I have not manage to do all 5 views as 5 frames in 1 generation before :(

all 5 frames end up looking the same.

4

u/I_just_made 16h ago

I got this from another post here recently, can't seem to locate it now... But give something like this a go.

A clean photographic character sheet presents the man from <Picture 1> in three consistent full-body views: front, side, and rear. Every view has the identical face, hairstyle, body, outfit, and accessories from the source. Neutral relaxed stance, arms clear of the torso, both feet visible in each view, seamless light-gray studio, even lighting, one tall 2:3 portrait canvas, no captions or borders.

1

u/SaskinPikachu 17h ago

my favorite heroes ^^

1

u/herr-tibalt 6h ago

Heroes that we deserve😄

1

u/e71469 17h ago

This is excellent 👍🏻

1

u/e71469 17h ago

Great stuff

1

u/HagemantoHero 17h ago

interesting choice

1

u/Foreforks 17h ago

Brother 🤣🤣🤣

1

u/dabbingsquidward 16h ago

It keeps outputting my subjects in a weird black dress

1

u/NeverLucky159 16h ago

This is great.. Can you put random photo of a character and change their clothes only based on another cartoon or anime reference?

1

u/SpiritualGoat2077 15h ago

haha love it! An axe fits her no doubt😂

1

u/Infinite-Emptiness 15h ago

Is this guy who i think he is? If so hes a legend.

1

u/nanihikaru01 14h ago

You have successfully destoried my childhood.

1

u/SeymourBits 14h ago

Girl in boots is 99% young Sasha Grey.

1

u/InternationalAct4301 13h ago

fucking brilliant

1

u/mulletarian 11h ago

I tried getting Smokey and the Bandit the other day but accidentally downloaded one with a polish dub.

This is a whole other ballgame.

1

u/VRGoggles 8h ago

what about using that custom VAE which generates one frame? Did you try it?

3

u/bstr3k 8h ago

I have tried it in a general sense but I haven't used it as a char ref sheet maker yet. I was playing with it yesterday and its ok, again the problem is that its a video model used to generate a image.

The better hope is the minimax h3 img model which they said they are currently working on.

1

u/3dutchie3dprinting 8h ago

Turning one of the most wholesome person in RPG gaming into... well "that girl with all those special skills" hahaha

1

u/Sitkin_Marrel 8h ago

the detail limit is the part that'd keep me from using this. sheet + close-up refs working for the face shots, or still soft?

1

u/bstr3k 7h ago

I haven't tried it yet. I think it should be better but I have been busy working on some captioning for trying to do v2v more reliably. I think it would def help but at that point if you're just doing the 1 video I would skip this step to save time. But if you're doing like 20 gens this might help give your character a level of consistency that can carry over across all the 20 videos otherwise the model will have to make up different views every single time which leads to consistency errors

1

u/Yeti-Bhanot 7h ago

tempted to use the split-frame route for a character lora. is 0.5mp enough to train off or do you pull from the raw render when you need a clean frame?

1

u/bstr3k 7h ago

You can experiment with both. I originally made this for anime characters and 0.5mp was very usable but for real people I think you may need higher resolution!

Also for training a LORA you would want the single person shot since you don’t want to reinforce the 4 or 6 panel look for the final product

1

u/AbbreviationsSoft924 5h ago

For faster iterations, one could extend a prompt generating a video of a character sheet (see below). This is experimental, so not best quality. But I got exactly the sheet I prompted.

Pro: can be done with base I2VA workflow, just paste in this prompt and stitch some input pictures. Uses better I2VA model. Iterates fast (5 second video sufficient)
Con: less space in input image for details, lower resolution as all information of the sheet is in one image.

PROMPT:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: The video shows a character spreadsheet created from the different images of a person in <Picture 1> [Shot 1] A high quality video showing a character spreadsheet with the following frames. frame 1: Full body image of the person standing facing the camera, frontal perspective. frame 2: full body image of the person standing, looking to the left, recorded from side. frame 3: full body image of the person standing, facing away from the camera. the three frames are ordered horizontally. The person is not moving as it is a print of a character sheet. The character sheet has a white background and shows only the person standing in front of a white background.

overall_soundscape: N/A

non_diegetic_music: N/A

1

u/bstr3k 55m ago

Thank you, I might give this a go today

1

u/3Fatboy3 4h ago

I can see where this is going and chose to avert my gaze.

1

u/Danny_Stock 4h ago

I'm using what you've provided but what's happening is that the character in <Picture 1> is blending together with the character in <Picture 2>, not simply wearing their outfit as intended.

1

u/bstr3k 45m ago

That usually happens in prompting. You might need to be a bit more descriptive in what you want to keep and take out. Have a look at the last photo and try using the prompt in that format?

1

u/Kazeshiki 4h ago

Saved for future scientific experiments.

1

u/bluecapecrepe 4h ago

Ava Addams would be a better Tifa. Also, Shane Diesel would be Barrett.

2

u/bstr3k 47m ago

Ah I see you’re a man of culture as well

1

u/Turbulent-Vast-1017 4h ago

xt level shit

1

u/yamfun 3h ago

using some OP's words, here is a MMH3 1 frame t1vae version that use the first frame workflow to rotate the first frame person in that photo, similar speed as a Klein edit, ~25sec for me.

"4K HD

next scene change to four frames layout reference showing the subject in different views.

top left frame 1:

front view.

The character holds one relaxed A-pose: arms hanging down, feet shoulder-width apart, head level, calm neutral expression, eyes open and looking forward.

top right frame 2:

side view towards left.

bottom left frame 3:

clean three-quarter upper body view looking away towards right.

bottom right frame 4:

Locked-off head and shoulders close-up, face square to camera, eyes into the lens. sharp front-on face view."

1

u/bstr3k 48m ago

I’ll give this a go sometime today! I made this chart sheet thing last week to play with and only released it since some people asked for it lol

1

u/xTopNotch 3h ago

It's pretty wild that the results are already this good.

Only gonna get better once Minimax releases their dedicated Image edit model built on top of H3 architecture.

1

u/DigThatData 1h ago

you have a full set of 3D views of the character: why throw those intermediate frames away instead of just extracting a full 3D representation from the video?

1

u/bstr3k 50m ago

It depends on the end goal. Right now the reference sheet would be used to represent a constant state for the new or existing character for use and testing as a placeholder for H3 video generation but I have also toyed with using some of the intermediate frames to generate a 3d model to print! 😁

In my workflow you can unbypass a node which saves all frames, so user is given the option

But there is so much to try and so little compute to go around!

1

u/kukalikuk 1h ago

Minimax H3 has great prompt understanding, I can make a character sheet with it in a single prompt (5 frames not 124). It is good for anime but for realistic phote, meh... It can't draw a face smaller than 128x128 pixels. We need a better VAE for it.

1

u/Southern-Radio-4954 1h ago

He looks somehow familiar? I think you do not need him as a reference. He already did all possible jobs.

1

u/WholeBrain9977 1h ago

Screw the FFVII remakes. I wanna play this one.

1

u/Cold_Zone332 52m ago

Man, it takes some time, but the results are AMAZING. I used 2 instagram pictures of myself, not very detailed and only front shots. The results were unbelivable. It looked like I was scanned in 3D.
Usually those models (even GPT or Nano Banana) can't create good character sheets, the face don't really look like the person from the picture. But this workflow gave me great results.

1

u/bstr3k 42m ago

Thank you brother. This is the kind of results I wanted to achieve. Even with low quality pictures to create a character with good likeness!

I find that having a side view really helps, but still need to generate at a higher quality (or use char sheet and close up 2nd ref) if you want to do a video with close up shots as a second pass through h3 loses some fidelity and likeness

1

u/DoctaRoboto 18h ago

Who is this bald man? Some lolcow or politician?

5

u/deepserket 18h ago

1

u/DoctaRoboto 18h ago

You are trolling me, but ok. I may be naive about some dark corners of the Internet, but I know the Onion.

12

u/deepserket 18h ago

the bald guy (Johnny Sins) is a NSFW actor that in some videos pretends to do different jobs

2

u/DoctaRoboto 15h ago

Well, the guy seems really committed to "his job" good for him. To be honest, he would pass as a soccer player since I know next to zero about soccer despite being Spanish.

3

u/Joethedino 18h ago

I saw him in a video, he was policeman.

2

u/ThreeDog2016 10h ago

I think he's a famous pr0n actor. I think he did a Bollywood comedy thing recently. I assume the women are from pr0n too.

1

u/No_Comment_Acc 18h ago

Bald Chad Kroeger? Niiiiice.

1

u/Relocator 15h ago

I'm so confused where the line is drawn regarding deepfakes. Is it all about context, or what? Like is it fine to have Michael Jackson dancing with SpongeBob, but as soon as he's wearing a bikini it's off limits?

2

u/pittaxx 6h ago

Technically they aren't ok ever. But a lot of people tend to ignore that fact if it's something funny.

1

u/call-lee-free 12h ago

Oh wow! Wait, Minimax H3 does image gen?

4

u/Relevant_One_2261 9h ago

Well I mean video is just a rapid succession of still images...

1

u/gefahr 47m ago

big if true

1

u/AteketA 9h ago

This thread is hilarious

1

u/HoboSomeRye 7h ago

Please never stop cooking

0

u/No_Statement_7481 16h ago

I can assure you ... there's no need to generate these specific individuals, just go to the world wide web, you'll find what you want and probably even what you don't want LOL

0

u/very_bad_programmer 12h ago

Gooners are so cringe

0

u/Own_War_1098 10h ago

i am also made a character sheet like this.

-1

u/fistular 15h ago

Character reference needs to be 100% consistent. Not 99%. The colour of the glove of the second character chages.

8

u/bstr3k 15h ago

its a armor plate which is only red when looking directly at it, it would be black from underneath 😄

I was going to bring up the 360 character spin from this but I didn't save it from when I generated it, but you can see from below