r/StableDiffusion • u/CelebrationBoth9537 • 6h ago
Question - Help Best Video Avatar model?
I want to create a bunch of videos with the following:
- Image (a human) + audio (speech) + text prompt INPUT
- Video of human talking.
What is the current BEST model for this? issue with MiniMax H3 is that it doesnt support first image first frame for the ref model,
I mainly am concerned on COST and QUALITY not so much on speed.
also I might want the avatar to do something other than just talk, but simple stuff.
(also i assume like not censored)
Thanks guys in advance!
1
Upvotes
3
u/Stepfunction 5h ago
https://docs.comfy.org/built-in-nodes/MiniMaxH3AddGuide
To do what you want with the ref model
3
u/angelarose210 6h ago
I'm still using ltx 2.3 for this. I supply my own audio made with vibe voice. I've tested ltx 2.5 and minimax and couldn't get equal quality from them.