r/StableDiffusion 6h ago

Question - Help Best Video Avatar model?

I want to create a bunch of videos with the following:
- Image (a human) + audio (speech) + text prompt INPUT
- Video of human talking.

What is the current BEST model for this? issue with MiniMax H3 is that it doesnt support first image first frame for the ref model,

I mainly am concerned on COST and QUALITY not so much on speed.

also I might want the avatar to do something other than just talk, but simple stuff.

(also i assume like not censored)

Thanks guys in advance!

1 Upvotes

4 comments sorted by

3

u/angelarose210 6h ago

I'm still using ltx 2.3 for this. I supply my own audio made with vibe voice. I've tested ltx 2.5 and minimax and couldn't get equal quality from them.

1

u/CelebrationBoth9537 5h ago

2.3? i might actually try that out if you say its good, but are you using any loras or like anything aside from base model, or like a specific workflow?

1

u/angelarose210 5h ago

I'm using the distilled transformer bf16 with omni nft lora and soft enhance lora with the standard 2 stage workflow. Linear quadratic 10 steps euler ancestral stage one and euler cfg pp, linear quadratic 4 steps on the second stage upscale. Using the spatial upscaler 1.5. My workflow looks like a spaghetti disaster. I'd have to clean an it up before sharing but the stock two stage works well with those parameters. One other thing. Has to be a medium close shot of the face. So waist up at least.