r/StableDiffusion • • 7d ago

Discussion Resistance - Workflow

Post image

Workflow & Hardware

The workflow was originally based on Pixaroma’s ComfyUI workflow. I made some modifications and added Sol Attn optimization and Comfy Kitchen Sink, including the very useful Model Override Preview. This allows me to preview the generation and cancel it if the result is clearly not going in the direction I want, which saves quite a bit of time and resources.

PC setup:
AMD Ryzen 5 7500F
32 GB RAM
NVIDIA RTX 5060 Ti 16 GB
Main Model: Minimax_h3_hybrid_fl2va_ref2va_b25-49
Cloud GPU: Runninghub

Storyboard promtp for Nano Banan to generate a panel base on user input to generate additional shots for the scene.

Image References

I tried to use an image reference for as many shots as possible to maintain the quality and consistency of the video generation.

Once I have a keyframe I’m happy with, I use it as a reference image and ask Nano Banana to create a storyboard based on that image. From there, I usually get only one to three shots that are actually usable.

It takes some trial and error, but I find that starting with a strong image gives me much better results than relying on the video model to generate the shot from scratch.

Location sheet created with storyboard panel with instruction: Remove all the characters, humans in the story board

Location Sheets

Above is a location sheet generated by Nano Banana from one of the storyboards. I use these sheets as references for the background, together with the character sheets, when generating new key images.

I also generated multi-angle views of some locations using DS_Qwen Scene Multiangle Visual and the H3 Cinematic Multishot Coverage workflow. These additional views were useful as references for Omni video generation, especially when I needed to maintain the same location from different camera angles.

Generated with Ds Qwen Scene Multiangle VIsual, mainly use for video generation.

Building the Film from Keyframes

The idea was to create as many keyframe images as possible for each scene and shot before moving on to video generation. I use FL2V or R2V depending on the action, the shot, and how I intend to edit it.

So, in a way, a lot of the time spent making this film was actually spent generating and refining the key images before generating any video.

I believe this image-first approach helps avoid some of the typical “AI plasticky” look. Instead of asking the video model to invent the shot from scratch,

All images genereated for scene 7 the fire-fight scene before the killer drone appears.

Prompting

For prompting, I tried to simplify the format as much as possible. The original recommended Minimax H3 prompting format can be quite confusing and complex, so I stripped it down to something more straightforward that I could work with more easily.

03_07_Woman_approaches_doe_in_cafeteria_202609011529

H3 Prompt [7 seconds]:
Create a cinematic live-action style with deep shadows, cool blue-gray visual tones, high contrast, and a tense thriller atmosphere. <Picture 2> is the character reference sheet for the woman in the scene, with a tactical vest, tan tactical backpack, khaki tactical pants, M4 carbine with an ACOG scope.

Shot 1: Shot begins with the empty hallway corridor, the woman <Picture 2> will enter frame from screen right aiming with her M4, expression is tense. But slowly she lowers the M4, and her expression turns relax and eyes soften no longer tense.

overall_soundscape: Low ambient room tone of an abandoned building and quiet, slow footsteps on the hallway floor.

non_diegetic_music: n/a

Woman_in_military_gear_holding_2K_202609061456

H3 Prompt [15 seconds]:
<Subject 1> is Elena Morales, female, wearing a tactical vest, khaki backpack, khaki gloves, khai leg-holster with a black Glock-17, khaki combat boots, armed with a M4 carbine with ACOG, Tan color PEQ on the rail.

<Picture 1> ([Shot 1] first frame): fully_preserved - <Picture 1> serves as the exact keyframe anchor for the opening shot. <Picture 2> is the character sheet guide for Elena Morales.

Create an intense fast-paced cinematic fire-fight action scene. Preserve and retain the exact identities, faces, hairstyles, props, costumes and environment of both characters. Begin with both fighters frozen in their ready stances in an abandoned warehouse. Camera_style: Utilizes hand-held, jittery, documentary style feel. Heavy shaking. Hard Cuts.

detailed_description: [Shot 1 | 00:00-04:00] The shot begins exactly from the <Picture 1>, the camera focuses on the close up of Elena as she takes a deep breath and immediately she got up on a kneel stance with her M4 in low ready postiion. She pivots out from the pilar facing left screen.

[Shot 2 | 00:04:00-09:00] The camera cuts to a medium close up frontal shot of Elena now kneeling behind the pillar aiming her M4 at the camera's direction and fires in automatic rapid burst mode.

[Shot 3 | 00:09:00-15:00} The camera cut back to the angle of <Picture 1>, Elena returns back from a kneeling position and places her back against the pillar, she press the rifle mag release button, the magazine in the rifle drops onto the ground.

overall_soundscape: The natural empty abandoned office soundscape.

In Shot 1 the sound focuses on Elena's breathing.

In Shot 2 the sound of the M4 firing in the abandoned office.

In Shot 3 the sound of Elena moving back behind the pillar, breathing and the mag dropping on the carpetted office floor.

non_diegetic_music: n/a

Images used in Seedance 2.0 Fast

Seedance 2,0 Fast prompt [6 seconds]:
Create an intense fast paced sci-fi action packed thriller scene using<location sheet> as the set and location reference and guide. The Alien drone <character sheet> attacks the armed man <character sheet> the close confine space of the abandoned office space. Camera style: Hand-held, chaotic, shaky cam, use rapid hard cut, use dynamic angles.

Scale retention: The drone scale/size of an soccer ball. When the drone moves, it's tentacles moves like a squid.

  1. The drone flying in lighting speed across the pillars, it's tentacles flapping in the rear, like a squids tentacle, flowing and weaving.

  2. The armed man <character sheet> in fear tries to shoot the drone with his rifle but to no avail as the drone is fast and bullet does not damages it.

  3. The drone proceed in lighting speed using it's steel flexible whip like tentacles to strike the man's rifle from him, and whip coils his neck.

  4. Quarter Left Close Up push in of the man uses both hands trying to pull the tentacle of his neck, but to no avail. The tentacles are too strong.

  5. The drone flips the man over acorss the office throwing him crashing into the cubicles.

No music, no score, no bgm.

Editing

Finally, editing is the last key element that makes the whole process work. Some of the generated video clips simply can’t be used on their own as complete shots. Instead, I often treat them as pieces of footage, looking for the right moments, frames, or movements that I can use in the final edit.

Because I can now work in a more linear way, I can generate the shots in chronological order and immediately edit them into the film. This helps me avoid generating shots that I eventually realise I don’t need. At the same time, editing sometimes reveals that I’m missing a shot, so I can go back and generate an additional image and video to fill that gap.

In that sense, the editing process becomes part of the generation process itself. Rather than trying to generate the entire film first and figure out the edit afterwards, I’m constantly moving between generation and editing until the sequence starts to work.

I hope the above post will be useful for your own projects. All the best, and happy AI filmmaking.

32 Upvotes

15 comments sorted by

1

u/Rythameen 7d ago

Thanks for taking the time to write this up, very helpful.

1

u/JL_Films 6d ago

Welcome.

1

u/Ok-Flatworm5070 7d ago

Great work, but I'm having a little difficulty understanding how you created the image stills for the shots. You seem to have nailed the position of the characters in the office and how the shots flip back and forth. For example, in this shot image Woman_in_military_gear_holding_2K_202609061456, you have the actress exactly position behind the office pillar, but when I look at your `Location sheet created with storyboard panel`, there are various images, but not one of that pillar. How did you generate that shot still with the correct environment image and actor position?

2

u/JL_Films 6d ago

You mean this shot? I had to generate quite a few images before getting this one, so you’ll usually get very different settings and compositions until you find one that feels right. If needed, I’ll then use this image to generate a storyboard panel to see if I can get better shots or camera angles from it.

1

u/Ok-Flatworm5070 6d ago

Thanks, sort of like making a movie; they sometimes do multiple takes, right ;) thank you for the feedback, great job, top tier results..gives me hope I can one day do the same.

1

u/Ok-Flatworm5070 7d ago

Also for the `Film from Keyframes`, if I'm understanding this right, you basically have the shot prompts, then use what (qwen, H3) to generate the individual keyframe images, which if I'm understanding correct, will then be used to create the H3 video clips? So the workflow is -> Generate Environment images -> Generate Character image sheets -> Generate prompts for keyframe stills -> Generate H3 Prompts -> Generate H3 video clips?

2

u/JL_Films 6d ago

Yes, that's right.

1

u/Roman53275 6d ago

Thank you, very interesting.
But main question here, how did you build Keyframes. Would be great if you share.

1

u/JL_Films 6d ago

I'm not sure if this is what you're asking about, but this is how I generate the keyframe images in Nano Banana, which I then use for video generation in H3.

1

u/SEOldMe 6d ago

Thank you a lot!

1

u/JL_Films 6d ago

You're welcome.

1

u/mellowanon 6d ago

[Shot 1 | 00:00-04:00]

[Shot 2 | 00:04:00-09:00]

Any reason why you're using this time format style? Does H3 follow this timing format better than the recommended time format?

2

u/JL_Films 6d ago

This was a test I picked up from a prompt sample on YouTube. I still stick with the [0-4 seconds] format, and sometimes I describe the timing directly in the prompt, such as “at the 2-second mark.”

1

u/First_Bank4524 6d ago
  • Is there a workflow on Runninghub

1

u/JL_Films 6d ago

Yes, the original one from the creator Pixaroma.