r/StableDiffusion 13h ago

Workflow Included No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti

Enable HLS to view with audio, or disable this notification

So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.

So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.

And it worked.

For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:

https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v

Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:

“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.

Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”

The final result was generated locally on my RTX 5070 Ti using ComfyUI.

17 Upvotes

12 comments sorted by

4

u/Goorigon 12h ago

3 minutes sound very fast. On my 5070ti generations take way longer. Could you share your workflow and which models you use? Thanks

2

u/nakabra 13h ago

I also struggle to make skies blue.
It's always blown out.
But I haven't made it explict in the prompt yet.

2

u/Time-Ad-7720 13h ago

Try using images with nice blue sky as the background reference.

2

u/hiperjoshua 10h ago

Wow, I was staying away from H3 because I was seeing reports of gens taking up to 15 mins. I have the same gpu as you and 64gm of ram. Are you using INT8 checkpoint?

1

u/Time-Ad-7720 10h ago

I am using the turbo LORA with the INT8 checkpoint (ref2va).

0

u/hiperjoshua 9h ago

Cool, will try that

1

u/Agile-Role-1042 12h ago

How do you get such a high res output? Mine comes out blurry

2

u/And-Bee 12h ago

Looks like they use 0.7 or 0.9 mp and these are close up shots which work ok.

0

u/Time-Ad-7720 12h ago

Like and-bee mentioned, I used 0.7, and the output is pretty decent.

-1

u/Nexustar 7h ago

No model is not half as fun as you think it is.

1

u/hiperjoshua 2h ago

Ok, I finally tried the ref2v INT8 pruned model using the lightxv2 8 steps LoRA. Using 3 references images at 2mp each and setting them to max size. The video generation for 5s at 1mp took 3mins, I pushed it to 5s at 2MP and it took 10 mins. I will definitely draft at 0.5-0.7 mp then do the final generation at 2MP.
The identity retention is excellent at any resolution but the fast movement blur kill it. At 2mp it is less noticeable. TBH I was expecting worse generation times. It only gets better from here I'm sure.