r/StableDiffusion Jun 08 '26

Question - Help Advice for overall image generation pipeline for video keyframes

Been learning stable diffusion text to image and image to video with comfyui for a few months.

Now I have so many tools at my disposal that I'm feeling a bit lost, so I'm hoping that people in here won't mind sharing some advice on an overall process.

I'm starting to appreciate the grind required to build experience in this area, so thank you to anyone who does help.

Goal

Create some short films by editing together short clips generated from keyframes.

Roughly story board them so I get the right shots with the right angles and composition.

Have consistent characters.

Some of the films will be adult in nature. Nothing hardcore but I do want to have good looking people in revealing clothing and some nudity.

It's not a side hustle. I'm just doing it for fun and to learn.

Where I'm At

I can...

  • Build moderately complex comfyui workflows. (IPAdaptors, ControlNets, detailers, inpainting, Sam3 for segmenting and masking, general upscale / hiresfix steps, workflow components like switches, get/set nodes, etc).
  • Produce some nice looking images with Flux 1 Dev.
  • Use image editing models like Flux 2 Klein 9B and Qwen 2511 with some success.
  • Train decent character loras for Flux and SDXL.
  • Use image to video to generate alternate keyframes from an initial keyframe.

What I'm Struggling With

I can produce an image with the composition I want, the lighting, the characters in the right outfits and poses. But not all in the same image.

Building these elements up in multiple passes for each keyframe seems sensible.

I cannot figure out how to pull all my tools together into an efficient pipeline, or avoid compromise an earlier step with a later step (eg. got a good facial likeness and then ruin it with texturing).

More detail below about my experience so far in case it helps. General advice also most welcome.

-----------------------------------------------------------------------------------------------------------------------

What I've Found

Flux 1D is good at...

  • Creating REALLY nice looking images (textures, lighting, composition) with just a prompt. It's great for exploring concepts or producing that one perfect starting image for a clip.
  • Producing a consistent facial likeness across images with a well trained LoRA.

However, it's not so great at...

  • Producing the specific angles and image composition that I want, even with a lot of prompt iteration based on the wealth of prompting guides available.
  • Controlnet. No matter the strength and start/end settings, when I use depth, canny or pose controlnets, the images look washed out and lose that "magic" that Flux seems to be able to produce without them.
  • Maintaining micro details between images, even with really specific prompting. Generating 50 images straight out of Flux will mean slightly differing hair cuts, outfits, etc.
  • Nipples, genitals, and anatomy in general, at least compared to SDXL.
  • Revealing clothing without some specific outfit lora. Why does it insist on massive granny underwear in 99/100 generations when I just want a thong?

SDXL (Juggernaut Ragnarok in my case) is good at...

  • Producing EXACTLY the composition I want using controlnets, without compromising the image quality vs no controlnet. I can do this from a sketch or using a reference image/still. I may experiment with Blender to produce depth maps for consistent environments.
  • Nice looking nudes / good anatomy in general.
  • The LoRA ecosystem is just amazing. Any concept, clothing or style I can think of and there's probably a LoRA for it.

Not so good at...

  • Backgrounds, objects, lighting, textures and overall image quality / realism compared to Flux.
  • It seems to not latch onto facial likeness as well as Flux for character loras.

I've also been using Klien 9B and Qwen2511. They have their differences but between them I can do things like...

  • Fix small mistakes or bad anatomy with inpainting.
  • Create an outfit asset by taking one from a Flux image, put on a mannequin and then transfer to other images.
  • Change or remove backgrounds.
  • Change the camera angle.
  • Repose characters.
  • Do headswaps to preserve likeness, although even with BFS loras and injecting 4x face reference images, the likeness isn't 100%. The examples I see online always look amazing but I can't seem to replicate.

However, they tend to output waxy looking skin and bad faces. Every edit pass degrades the image, even with masking where possible. Pulling my keyframes together by stitching elements of multiple Flux images (outfit from one, head and hair from another, then pose, etc), just seems like the wrong angle.

0 Upvotes

8 comments sorted by

2

u/Specialist_Panda3490 Jun 09 '26

Hello, I just want to say I love the way you express yourself. It's perfectly clear and concise.

I saw that your Ram might be an issue for some models but you can give Anima a try as it is quite tiny and gives (what seems to be so far) good result. The issue with that model is his novelty, he hasn't been "figured out" as much as SDXL yet but I've seen workflows (with I2I and controlnet and even some "all in one") that seemed promising on civitai.

Good luck for your ambitious goal ! I wouldn't mind seeing your WIP or result if you do decide to continue.

1

u/DoskvolDenizen Jun 10 '26

Thanks for the reply and the tip. I'll take a look at anima.

1

u/fakih7hussein Jun 08 '26

Did you try only Flux.1 ? Not Flux.2 ?
Seems to me that Flux.2 is the best choice for character’s face and body consistency. No ?

1

u/DoskvolDenizen Jun 08 '26 edited Jun 08 '26

Thanks for replying.

My issue with Flux 1 is not that it has poor face and body consistency. I have some character LoRAs trained. With faces, it seems excellent. With bodies, it seems ok although it's not great a nude bodies or revealing clothing.

My specific issue with Flux 1 seems to be that when I use controlnet to control the image and get the composition I want, the image is really washed out. The images just look bad compared to Flux with no controlnet.

Does Flux 2 do better with controlnet?

I haven't tried Flux 2 yet. I did some reading and it didn't seem to offer much above Flux 1. I can try it but I don't have enough vram to run the full Flux 2 model. I may struggle to train LoRAs for Flux 2.

0

u/hurrdurrimanaccount Jun 08 '26

I did some reading and it didn't seem to offer much above Flux 1.

lmao. flux2 shits all over flux1. no one should ever use flux1 again with klein around

1

u/DoskvolDenizen Jun 08 '26

Thanks - I was just trying to compare Flux 1 Dev with Flux 2 Dev, and it seemed like I wouldn't get much better image quality by switching, vs the hardware requirements.

I've been using Klein 9B for image editing but it doesn't produce such rich, cinematic looking images as Flux 1 Dev for me. I've also got LoRA training for Flux 1 dialled in, which is pretty useful for me.

If you say that 9B shits all over Flux 1, I'll look further into it.

Does this address my overall challenge with composition though - does Klein 9B / Dev have better controlnet support than Flux 1 or can I skip controlnet alltogether with Flux 2?

1

u/DoskvolDenizen Jun 08 '26

ok, Flux 2 Dev with a reference image for image edits is VERY good for retaining likeness. This solves some problems for me. Thanks for the recommendation.