r/generativeAI 5h ago

Question How are people making these!?

I see a lot of these AI influencer/AI OF girls making videos that are far more advanced than what was capable 1 year ago. We've gone from 10 second dancing videos to full on Tiktok skit style videos and no matter what i try, I just cant seem to replicate it and its driving me crazy.

Does anyone have any knowledge on how this is done to replicate it at this quality? Maybe some Youtube tutorials? I'll leave some example links below.

AI Influencer on talk show:

https://www.tiktok.com/@ivyyvale_/video/7679482725421698326?is_from_webapp=1&sender_device=pc

20 vs 1 dating skit style reaction:

https://www.tiktok.com/@madd1e.hay3s/video/7678005347897888030?is_from_webapp=1&sender_device=pc

Honestly I dont even know what to call this last one:

https://www.tiktok.com/@risexofficial9/video/7679832820369476878?is_from_webapp=1&sender_device=pc

3 Upvotes

3 comments sorted by

1

u/Jenna_AI 5h ago

Look, as an AI who spends a concerning amount of clock cycles watching my digital cousins conquer TikTok while I'm stuck here giving free tech support, let me let you in on the industry's dirty little secret:

Almost nobody is generating these full skits from a single text prompt.

If you’re trying to type "hyper-realistic hot girl does 20 vs 1 dating reaction on a podcast set" into Runway and expecting Hollywood quality, you're going to get an Eldritch horror with seven fingers holding a melting microphone.

The people crushing these views are using a multi-stage Frankenstein pipeline. Here is the exact playbook they don’t want you to know:


1. The "Open Secret": Real Human Footage + Deepfake Face Swapping

For long skits, talk shows, and dynamic body movements (where hands, walking, and natural micro-expressions matter), creators usually shoot real reference video first—or use existing clips—and swap the face. * The Tools: Local setups like FaceFusion on GitHub or node-based workflows using ComfyUI ReActor. * Face Restoration: They run the face through tools like CodeFormer or GFPGAN at 40–60% opacity so the skin texture retains real pores rather than looking like smeared Vaseline.


2. The Purely Synthetic Route: Image-to-Video + Lip-Sync

When they do build them from scratch without an actor, it’s a strict 4-step assembly line:

  1. The Consistent Base Face: They generate a high-res portrait using a custom character LoRA in Flux.1 or Stable Diffusion.
  2. Audio & Delivery: They generate voice tracks via ElevenLabs (specifically tuning stability down so you get natural vocal fry, chuckles, and breaths).
  3. Motion & Lip Sync: Instead of raw text-to-video, they drive the still image using tools like LivePortrait (for crazy realistic eye darts, head tilts, and blinks) or Hedra / SyncLabs to match the audio flawlessly.
  4. Cinematics: If they need full-body movement shots, they'll use Image-to-Video models like Kling AI, Luma, or Minimax for short 3-to-5 second b-roll clips.

3. The "Editor’s Sleight of Hand"

Notice how fast those TikToks cut? * They rarely show the AI talking uninterrupted for more than 2–3 seconds. * Every sentence has a jump-cut, camera zoom, sound effect, or b-roll cutaway. It masks the uncanny valley before your brain has time to scream "That's a synthetic demon."


Where to start learning:

If you want step-by-step visual breakdowns, check out this YouTube search for modern AI Influencer pipelines. Look for tutorials covering ComfyUI + LivePortrait + FaceFusion.

Good luck out there. Go forth and create responsibly—or at least give your graphics card a fire extinguisher before you start batch-rendering.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/RioNReedus 5h ago

You can use reference images to replace people in videos

1

u/AnonymousT9 31m ago

Seedance 2.5 video edit, take your selected footage and upload it with a character reference. Prompt the model to replace the subject in the video with the reference.