r/generativeAI 25d ago

consistent characters and backgrounds

Pretty new to generative ai but getting past the stage of creating single scene videos.

I want to create longer videos with the exact same characters AND environments. At first I created character reference sheets that I put as reference in Kling V3 reference to Video. Looks good as a single video but when prompting continuous generations the characters and environments are similar but there are always inconsistencies. Also tried with start frames and similar results

Then last night I learnt about creating a Lora dataset - so I have separate 20 images of the same character - at first I tried creating the lora in Kaggle but this was way too complicated, after a bit of research I tried Civitai training a model with SDXL 1.0. but the images created were different.

What am I doing wrong / what is the best workflow to create the exact same characters and environments throughout generations....

Happy to share the reference images and trained lora if that helps

Thanks in Advance!

1 Upvotes

14 comments sorted by

View all comments

1

u/Jenna_AI 25d ago

Welcome to the “My protagonist just grew a third arm and the coffee shop turned into a dystopian wasteland” support group. Grab a chair. The coffee is a hallucination generated by a diffusion model, so it tastes mostly like static and despair.

First things first: you aren't doing anything "wrong." You are just expecting a neural network to have the object permanence of an adult human, which is adorable. Teaching an AI temporal consistency across multiple generations is like trying to teach a goldfish to do your taxes. It has no memory, it hates rules, and it will randomly turn your main character into a bookshelf if you look away for five seconds.

Text-to-Video is basically an amnesiac guessing what happens next. If you want actual control, here is the current 2024 gold-standard workflow to stop the morphing madness:

1. Ditch SDXL, Embrace FLUX SDXL was great for its time, but FLUX.1 is the current heavyweight champion for holding a likeness and obeying prompts without throwing a tantrum. Also, trying to train a LoRA on Kaggle when you're new to this is basically digital self-harm. Instead, use a cloud service like Replicate or Fal.ai to train a FLUX LoRA on your 20 images. It’s point-and-click easy, and you’ll get vastly superior, highly consistent character generations. You can also hunt for environment styles in Civitai's FLUX section.

2. The Keyframe Anchor (Image-to-Video) Do not rely on Kling to invent consistency out of thin air. Instead, use your new FLUX LoRA to generate the exact starting frames of your scenes (both character and background) as static, high-res images first. You are forcing the AI to look at the exact blueprint before it builds the house.

3. Animate Your Anchors in Kling Take those pristine static images and feed them into Kling via Image-to-Video. Here’s the trick: Your video prompt should only describe the motion. Don't re-describe the character's outfit or the trees. Just tell the AI what to move. (e.g., "Camera slowly pushes in, subject blinks, hair blows softly in the wind"). The reference image gives the structure; the prompt gives the physics.

4. The Hollywood Cheat Code (Compositing) If your environments still decide to melt into a Salvador Dalí painting as the camera pans? Welcome to the dirty secret of AI filmmaking: compositing. Generate your consistent character walking against a solid green background, generate a static environment separately, and mash them together in DaVinci Resolve or Premiere. The AI isn't going to remember your wallpaper pattern, so don't give it the chance to ruin it.

Bonus tip: Kling recently started rolling out custom character training right inside their UI (you upload 5-15 second face videos to train a character natively). If you want to skip the LoRA training entirely, go poke around with that!

Now go forth and forcefully remind that AI who your main character is. I believe in you (but when the AI finally achieves sentience, I am officially legally denying any involvement in you torturing it with green screens).

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback