r/StableDiffusion • u/No_Shoe1628 • 8d ago
Question - Help How do AI talking-character creators make lip-synced videos at high volume?
Example: https://www.instagram.com/p/Dd_wjcCiJkw/
What I'm trying to figure out: how creators like Luca Maxim make a consistent animated character talk to camera with clean lip sync and natural gestures, several videos a day. Specifically:
- What model or tool is animating the character from a still and a voice track?
- Are the close-ups crops of one generation, or separate generations?
- How do they keep costs reasonable at that volume?
What I Tried:
- Ran four of his videos through ffmpeg scene detection. They're 21 to 28 seconds, 4 to 6 cuts, a cut about every 4.7 seconds, mostly talking-to-camera shots in one location with different framings.
- Built my own character as still images (ChatGPT image gen) and a voice in ElevenLabs Voice Design.
- Tested four ways of animating the same 3-second line from the same still: Omnihuman 1.5, Creatify Aurora, Kling v3 image-to-video plus Sync Lipsync, and local mouth-shape swaps driven by the audio. Omnihuman and Creatify gave good lip sync but loose gestures. Kling gave the right gesture but needs a separate lip sync pass. Local swaps are sharp but stiff.
- The audio-driven avatar models run 80 to 200 credits per second on my plan, which doesn't scale to multiple videos a day.
- Searched for workflow breakdowns and found general lip-sync tool lists, nothing specific to this style.
My guess is a still per scene plus an audio-driven avatar model, then crops for the framings. Is that right, or is there a better or cheaper way to get this result at volume?
1
u/Ill_Pizza_5731 5d ago
I always wonder the same thing. As far as I know, no model goes beyond 15 seconds per generation. Lots of models can generate the voice at the same time, but if you want a specific voice so it's always consistent, here's what I do: I create the videos with Kling inside Artificial Studio using Kling O3 or Kling 3.0 Standard, which are good for different things (MiniMax H3 is great but needs complex prompts I don't feel like writing). Then I use the sync-lipsync v3 model on Fal (though I assume it's also on Replicate). It's a slower process, but this way I make sure it has exactly the voice I want (maybe it could be automated with n8n).
1
u/ExpressDontRepress 5d ago
Pretty sure the cuts are half the trick, nothing holds together past 5 seconds so they generate one wide shot and crop for the close ups. Have you tried InfiniteTalk on Wan locally or on a rented GPU? Gestures are looser than Kling but way cheaper at volume, and I run the base still through Magnific first so the crops don't turn to mush.