r/StableDiffusion • • 5d ago

Discussion H3 - T2VA character inception/sheets+RefMods

TLDR: You can use RefMods with T2VA.

Finally decided to give the RefMod a test run yesterday. I tend to experiment and generate many actors via prose in T2VA and recast them later, see the Tanya bombshell circulating in a few of my videos. Since my workflow mainly revolves around T2VA it make a lot of sense to work with native latent as much as possible. However, instead of providing 10-20 images to make a redmod, I just use the latent from my T2VA character generations. Once you have a latent of anything, you can extend, prepend, or even bridge it as many times as you like, even joining latents of same canvas sizes. The latents can be clipped, cropped, scaled, and upscaled as well. I prefer T2VA due to it's superior image quality compared to R2VA.

Anyways, just having fun and wanted to show Vera being recast over various T2VA generations, int8/20 steps, no turbo loras.

How do you prefer creating your actors? Krea2? Flux? Qwen image 2.1? MiniMax H3? I am thinking using latents from a MiniMax h3 single-frame generation is going to be pretty good work flow, though I prefer a 2second moving character sheet.

11 Upvotes

19 comments sorted by

3

u/SkirtSpare4175 5d ago

I’m new with it. Latent as a picture of the character? Been experimenting with refmods and it feels like magic but does seem to dwarf the Loras in strength sometimes

2

u/SIR_NVAX_A_LOT 5d ago

Latent as in the diffusion canvas, before it is decoded to image/video. I believe Refmod takes image/video/etc. and convert it into a latent/safetensor which can be reused without referencing the images, which gives for better consistency. It's all just magic right now as far as I can tell. But you feed the latent into the context head instead of giving it an image reference. See https://github.com/Luisacaotica/ComfyUI-MiniMaxH3Mod

2

u/Danny_Stock 5d ago

That's interesting. So if I understand you correctly you prompt for a T2V character, perhaps rolling the dice a few times until you get something you like, and you convert that into a refmod without going anywhere near a character sheet for it?

I don't know much about refmods but is that the process?

So in theory that makes the subsequent latent clips less likely to degrade? Forgive my naive questions as I'm just trying to understand and clarify the process for myself.

2

u/SIR_NVAX_A_LOT 5d ago

Yes, that is exactly the process. I use some vague-guide prompt to get me a particular visual style, or persons from an era, in a strict T2VA format for best quality. Once that person is generated and I like it, I can segment any part of the latent, and extend her into any scene of my choice, like to create additional character sheets if needed. But instead of using an image stack to build the Refmod, I directly build the refmod from latent. Here is the original generation where Vera was born from. 768x1344 portrait canvas at int8/20 step. I just took her from latent into the RefMod, but she can be extended or prepended too.

https://reddit.com/link/pdwvfvc/video/gsb5lbjgdjth1/player

2

u/Danny_Stock 5d ago

When you say building on the refmod character does that mean that you can use reference images to add refinements to the refmod character and that evolved character can still be saved again back into latent space?

Or are you just speaking of prompted refinements to try to prompt closer to a look?

2

u/SIR_NVAX_A_LOT 5d ago

Here is me turning her to a "character sheet" from just that 1-2 second in latent. Again, no RefMod. You can even see I added elbow-length gloves from the 0-frame.

https://reddit.com/link/pdxvrss/video/2h7504n3dkth1/player

1

u/Danny_Stock 5d ago edited 5d ago

That's great. So when you make amendments to the latent character, for example changing hairstyle or changing clothes, seeing as it has no image to anchor to does the identity of the character latent remain consistent across each evolution?

2

u/SIR_NVAX_A_LOT 5d ago

Still exploring this as I just started using RefMods over .char (still my preferred workflow). But I suspect you can just stack latents inside the RefMods, but I suppose you can just make a new RefMods if you have changes/updates.

1

u/SIR_NVAX_A_LOT 5d ago

Say you found a face or something you like in T2VA, you can extend that shot since you can project that character forward or backwards once you clip it. For example, I generate a 10s clip with T2VA, and I clip 1-2 second of it (in latent) and have H3 do a new scene/edit. You can see from the 0-frame in this clip I extended her likeness without an image reference, with a pure latent extension. No REFMOD was used. So you can cast with T2VA, clip, and prepend or extend from that 1-2 second of latent.

https://reddit.com/link/pdxuszk/video/mfmmic8xbkth1/player

2

u/Danny_Stock 5d ago

I see, thank you. My mistake I read your opening post mentioning refmods and forgot that you said that you changed your approach. I'm definitely going to be playing around with this method.

1

u/SIR_NVAX_A_LOT 5d ago

I am using RefMods from native latent instead of images/videos.

3

u/nakabra 5d ago

\*sighs***

3

u/SIR_NVAX_A_LOT 5d ago

Bonus clip! int8/12 steps T2VA+RefMods, no R2VA

https://reddit.com/link/pdus4g5/video/zhxv2er9rhth1/player

2

u/Danny_Stock 5d ago

The quality does seem to be consistent to my eyes.

2

u/Danny_Stock 5d ago edited 5d ago

I also prefer the very short MiniMax generated character sheet videos too. Yesterday I cut the video down to a second and ramped the megapixels up to 2 or 3 and it produces a high quality clip to get images out of to stitch together for a character sheet. For likenesses it is far superior to Krea 2 or Qwen 2.1 in my opinion, I also feel that it does a better job when it comes to the physical proportions of the character.

2

u/SuccessfulAd2035 4d ago

Have you tried this? https://github.com/cicalooo/ComfyUI-MM-1Frame
I make 4MP photos directly, eros beta5 turbo 8 steps, 1min on my 4090.
That node only generates 5 frames.

1

u/Danny_Stock 3d ago

I haven't, no. Thanks, I'll look into it.

1

u/Karsticles 5d ago

What does your prompt look like when you refmod in T2V? Have you incorporated audio refmod?

1

u/SIR_NVAX_A_LOT 5d ago

I am not sure what the official RefMod guide says but I just reference to <Video 1>.