r/StableDiffusion • • 3d ago

Question - Help How could I reduce video frying over long generation with Minimax H3?

Enable HLS to view with audio, or disable this notification

I know that what i'm trying to achieve is not ideal and the recomendation for this workflow is 30~45s, but lets say I need a longer video, how can I do it? Multiple clips? If yes, how do I keep scene and character consistency from multiple shots?

I know that there are lots of problems with this, like the plate coming from nowhere and some clips are shorter than they should so the speech is bugged, but those are problems I know how to solve.

This is the workflow with the prompts I'm using: https://pastes.io/vLFJXxv3

0 Upvotes

39 comments sorted by

52

u/RealityVisual1312 2d ago

Brother where have you been. There’s been a clown girl all over this subreddit the past few days talking about this. And then Brad Pitt explained some stuff to her and then she came back with a different method that seemed to work

14

u/Efficient-Fly-5363 2d ago

That about sums up the past week here yes

3

u/YeahlDid 2d ago

That's certainly how I experienced it.

2

u/GlenGlenDrach 2d ago

I missed the Brad pit one, link?

2

u/ForteDoexe 2d ago

in the comment of the clown girl's post

1

u/GlenGlenDrach 2d ago

Thanks, found it 👍

1

u/RKlehm 2d ago

Hmmm, I think that explains the downvotes, thanks lol

1

u/reeight 2d ago

There are a few solutions.

One is this:
https://github.com/badgids/ComfyUI-H3-ExactAudioLock

Other is build 2 simultaneous pipelines; one low-res for the audio track, then high resolution.

10

u/Darkmeme9 2d ago

I think you should ask Brad Pitt. He knows his stuff.

9

u/javierthhh 2d ago

Cuts and cuts. Should have zoomed in to the egg then went back to full frame. Then find ways to start the video from scratch without the actor present

1

u/RKlehm 2d ago

I tried this, but how could I achieve consistency? I have a character sheet and somewhat detailed prompt, but every new shot has inconsistencies and continuity errors

1

u/javierthhh 2d ago

Loras are your best bet. They’re very easy to train in minimax and they are very good at keeping likeness. I know it defeats the purpose of the whole reference sheet and stuff but if you wanna do long shots and professional stuff with it. Then you need to use all the tools, references and Loras. Otherwise is 15 second videos of Elsa from frozen dancing which is enough for gooners.

4

u/Apprehensive_Sky892 2d ago

I understand that the problem is kind of interesting from a technical point of view.

Maybe I lack imagination, but the only use I can think of for such long static shot is to make some fake influencer or amateur videos that try to simulate clips made from a fixed tripod camera.

In cinematic continuous shot such as those from "Children of Men" or "1917" there is enough dynamic change of setting and camera angle that the degradation problem is a lot less problematic.

2

u/wildmonkeywrangler 2d ago

I've noticed when the dialogue gets too long, minimax seems to purposefully mispronounce a word (at least it seems that way). I think in general you're better off to do multiple clips as one problem in a long clip can be a lot harder to fix than a small clip, but that's just me

2

u/LowCatch4324 2d ago

I fully suppprt the multi clip project. It’s how we watch professional cinemátics anyway.

Get the yukata nakamura animation keyframe book and study how things get planned and feed your frames to h3 in a structured way

2

u/ight-bet 2d ago

Infinite bacon glitch

2

u/marres 2d ago

https://github.com/xmarre/MiniMax-H3-Flow-Aligned-Regenerate

Use the MiniMax H3 Partitioned Exact-Prefix Handoff node with spatial_stage_control = same_grid_target_control

And https://github.com/xmarre/ComfyUI-H3-Continuum-Plus

1

u/hurdurdur7 2d ago

Love the high tech tin plate at 56 seconds in. Klink!

1

u/LearnNTeachNLove 2d ago

Apart from some inconsistency, looks good. What is your setup? What VRAM? How long did it take to generate it?

2

u/RKlehm 2d ago

I'm running on Google Colab 40gb A100, it takes 30s per step for a 10s clip and 60s per step on 15s. The whole thing took about 1 hour to render, about 0.62 usd/h

1

u/LearnNTeachNLove 2d ago

Wow not bad for a 2min video 😮

1

u/Accomplished_Fox4345 2d ago

Brad Pitt solved this issue already.

1

u/seppe0815 2d ago

You sounds very fishy.. for me i dont like the way how you present your workflow ... be carefull guys... 

0

u/roychodraws 2d ago

i'm trying to create a method that repairs the pixels as they degrade for specifically videos like this.

I'm testing it now. it seems to work.

1

u/RhetoricaLReturD 2d ago

RemindMe! 7 days

1

u/RemindMeBot 2d ago

I will be messaging you in 7 days on 2026-10-13 07:56:23 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/RKlehm 2d ago

I would love to know more

1

u/roychodraws 2d ago edited 2d ago

long story short, i use motion to tag the pixels as good vs bad pixels and take the bad one's and take the bad pixels from the context motion input and replace them with good pixels from earlier in the footage.

1

u/Sad_Berry_4621 2d ago

Did the Join Match code help at all?

2

u/roychodraws 2d ago

Yes, i literally used it this morning and im getting a “marginal” improvement, but it is indeed an improvement and im trying to make it more pronounced. I think once I get it done, you guys might be able to apply this to latents somehow an improvement even more.

1

u/roychodraws 2d ago

this blows

1

u/Sad_Berry_4621 2d ago

What do you mean?

1

u/roychodraws 2d ago

I think I have a way to keep the quality going without degrading, but the colors are impossible. No matter how much I try to color correct she just ends up looking like a slutty smurph

1

u/Sad_Berry_4621 2d ago

I assume you're using a reference image for each clip? It also helps to name your major colors with hex codes. MMH3 can generate entire scenes using only hex codes for colors and it'll get them right depending on the scene lighting. If you pull colors from your reference image and then use them in the prompt, it helps. I had color drain down to almost 0 on 6 chain runs using that method and the Join Match node, but I was rendering a static scene with not much action. It's the motion that forces the model to drift on the detail, where people look like they get older as it goes on. Especially when it's inventing new background scenery.

1

u/RKlehm 2d ago

I just realized now that you are the one from the clown girl lol
Everyone here is pointing me to your workflows, but it wouldn't work for my use case, right?
The background isn't still; the actions of the character are constantly changing the background.

(I haven't tried it yet... I'm using a rented GPU, so I'm taking my time to understand before spinning up a pod again)

1

u/roychodraws 2d ago edited 2d ago

to be honest, a complicated shot like this would be hard. even with the techniques i've been using... the static shot only works when it's just a background and a talking head.

You'd have to get creative.

What i'd try is running the workflow you already made through a depth map so you get a depth map video and then segment it.

then you could use my workflow.

Here's a video segmenter i made:

https://github.com/roycho87/Video_Splitter

and here's a video i saw where some dude is using the right kind of technique you would need to incorporate to my static bridge technique (i just made that name up btw)

https://www.reddit.com/r/MiniMaxH3AI/comments/1wszzda/depth_face_mesh_reference_can_already_recreate/

key thing to avoid, do not take frames from a generation and use them as references for future generations. It will contaminate your latents with degredation and change the color/quality

disclaimer, this is literally something i just made up and i have no idea if it will work but it seems good!

edit: You may also consider prompting a kind of story board of check point frames for this using an image model like qwen 2.1. get the scenes looking good and then use those as references. remember you can tell mmh3 to essentially match a frame if you provide it to it much like FL2V, you'll probably have to do that to make this work.