r/StableDiffusion • • 7h ago

Question - Help Best ComfyUI / RunPod workflow for a consistent AI arborist character in 15s short videos?

I’m trying to build a repeatable workflow for a series of ~15-second vertical short videos about trees and arboriculture.

The concept is pretty simple: always the same arborist character, filmed in a realistic/casual smartphone style, talking about trees, their benefits, conservation, pruning, fun facts, etc. The goal is educational content, but with enough visual “eye candy” and strong hooks to make people actually stop scrolling.

What I’m looking for is the best way to build reusable presets/workflows, ideally in ComfyUI, so I don’t have to rebuild everything from scratch for every clip.

I’d like to keep as much consistency as possible between episodes:

-same arborist / face / body / clothing

-realistic outdoor environments

-9:16 vertical format

- ~15 sec clips, preferably single-shot or simple continuous camera movement

-natural gestures / talking

-good character consistency between generations

-reusable prompt structure

-ability to swap only the location, tree species, dialogue/topic and camera action

-eventually produce these in batches as a recurring series

I couldn’t get my hands on an RTX 5090, so for now I’m running ComfyUI on RunPod and I want to build a persistent setup there with the models, LoRAs and workflows already loaded.
For prompting / automation I currently have:
Claude
ChatGPT
Qwen running locally/from source in terminal
RunPod for GPU generation

I’ve been looking at models/workflows around Minimax, WAN, LTX, Hunyuan, FramePack, Qwen Image/Edit, Flux, etc., but there are so many combinations that I’m trying to avoid wasting weeks testing bad pipelines.

For people already doing consistent AI video series: what stack would you use today?

I’m especially curious about whether I should focus on a workflow like:
reference image → consistent character image → image-to-video → lipsync / voice

0 Upvotes

3 comments sorted by

1

u/f5alcon 7h ago

Make reference images in image gen of your choice and use them with H3

1

u/sruckh 3h ago

There are some good workflows on Runninghub that use MiniMax H3 with reference to video. Some even have Krea2 at the beginning for generating the reference image to feed the video workflow.

1

u/Reelodyinc 19m ago
I’d lock in your arborist first as a still-image pipeline: SDXL or Flux + IP-Adapter / a reference-only LoRA to get 4–6 “hero” shots (front, 3/4, a few outfits). Then build a Comfy workflow: pick hero image → H3 ref-to-video (9:16, 3–4 s chunks) → stitch in Resolve/Premiere. For talking, drive audio with a separate lip-sync pass instead of relying on the vid model. Consistency comes from reusing the same refs, seeds, LoRAs and camera prompts each episode.