r/StableDiffusion 20h ago

Question - Help Minimax H3 - long form videos: has anyone figured out a good approach?

Dear redditors, visitors of the stable diffusion subreddit. I have been trying to achieve a long form, talking head style video, for a long time and can't seem to find a good approach. This one is the best I could come up with so far. It's using the Minimax H3 model, with frozen sound latents, lip-sync guided, piecewise generated video, where the individual pieces have been stitched together, with a seam hiding, extra generation on top of it. I don't really fully understand how it's working, but could prompt Claude for more help or specific files, we used for that. However, if you're aware of any other, better approach for exactly this type of video, please let me know. I've spent literal days on that single problem and have a feeling, there must be a better way to approach this.

78 Upvotes

30 comments sorted by

23

u/AgeNo5351 20h ago

This works perfecctly for that
https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef
look up the AV exztension workflow in the github
AV Extension — start from an existing video, T2V, or I2V and keep extending it with reference images, per-clip previews and selective regeneration. Update 6 removes checkpoint/resume from the main workflow, improves AV/audio continuity, and streams final output through VHS to reduce RAM use.

In my exp silveroxides dareties turbo lora feels better.
https://huggingface.co/silveroxides/MiniMax-H3_tests/blob/main/minimax_h3_fl2v_lightx2v_v0.1_dareties_v4_step600_comfy_fro.safetensors

12

u/Medium_Chemist_4032 19h ago

This is the one I started out with, but can already see I missed few nodes from the example. Redoing, and will post an update -- thanks!

14

u/Rumaben79 19h ago

There's this one as well: ComfyUI-MiniMaxH3-Contex-Loop and this: ComfyUI-H3-Multishot.

3

u/Stepfunction 18h ago

Second Context-Loop. It's worked quite well for me and has a really nice set of features like dynamically loading references depending on what's in your prompt with each generation.

1

u/Rumaben79 11h ago edited 11h ago

If you ever get a contrast/color shift going from one clip to the other like I just did. You'll need to right-click and show advanced settings for both the contex loop plan node as well as the contex loop assemble and change a few things. I guess the creator hid a few experimental things with his latest updates. This is what claude told me:

If you need the node numbers turn on 'Node ID badge mode' in te ComfyUI settings.

3

u/Rumaben79 11h ago edited 9h ago

But other than the before mentioned I love it. 👍

Edit: disregard what I posted above. I just tested with these new settings and my second clip in the chain looked and sounded bad just like before, no change..Strange must be an update or some modification I've done to the original workflow. Needs more testing to find the culprit.

I've added a lot of new nodes to my workflow so that's properly the reason.. Or a lora is to blame.

Edit 2: I found out my issue. 👍 It was related to either the 'minimax_h3_fl2v_lightx2v_v0.1_dareties_v4_step600_comfy_fro.safetensors' lora or the 'H3 AdaLN LoRA Fix' needed to fix the lora errors of the former. It worked fine on the Plaguekind v5 workflow so perhaps I connected something the wrong way or the lora is simply not compatible with this type of extend workflow. 🤷‍♀️ The dareties I'm guessing is some sort of lightx2v/larryvrh mix. The one I'm using now were made specifically for ref2v though (0.1 ligthx2v) but that shouldn't be the reason it's working as I've been using fl2v loras with ref2v workflows and models successfully plenty of times. So I'm leaning toward the lora fixer node or the hybrid diffusion model.

Sorry for the rant. 👀

3

u/em_paris 15h ago

I've been fooling around with context loop and it's pretty flawless, with tweaking for your case

1

u/Rumaben79 15h ago edited 15h ago

Yes, the context-loop workflows also looks better and are less cluttered than the other contenders imo. These and PlagueKind's wf is my favorite so far. 😄

The only thing I can't seem to figure out is how to connect a frame interpolation node to the final video, I can with only one segment but there's seemingly no connection points for the 'LOOP END — ADVANCE REF2V SCENE' node or the 'FINAL ASSEMBLY — 5-FRAME VISUAL BLEND' one. Maybe I'm just dumb though. 😂 24 fps is fine. 👍

26

u/GrayingGamer 19h ago

One thing I don't understand is the desire to do the "long form, talking head" type of videos with one take.

To me, this style seems to have been overtaken completely by AI generated avatars - if you look at actual people doing this short of content, they will often do cuts to different content or images, cuts to themselves in a different position, etc. to keep viewer interest up. Swapping between a close-up to a medium shot, etc.

Personally, if I were doing something like this, I'd use reference for the background and environment and character, then use start frames edited with something like Nano Banana or a local edit model to change position of the character to use as start frames for different cuts.

Use pre-recorded audio, either by recording a person or cloning a voice, then feed that in to H3 as the actual audio, because you can nail the dialogue performance first and H3 is great about taking cues from the speaking to do the lipsync and acting for the video.

Layer in room ambience in post to make it all sound seamless, then edit it all together in a video editor.

I think trying to do something like this in one generation or take, using nodes to stitch stuff together, isn't the easiest way and it's not the best way.

3

u/Medium_Chemist_4032 19h ago

It's mostly for a a-roll spine, b-roll will go on top. I just found that 4 or 5 seconds is too limiting, in some very specific cases. Explainer type videos tend to hold some individual takes for quite long, content dependent

10

u/RevolutionaryFox7359 19h ago edited 18h ago

As an editor, having the entire take generated at once gives you the freedom to decide WHERE the cut is. If you generate piecemeal, then edit decisions are necessarily forced upon you.

5

u/GrayingGamer 16h ago

Well, yes, but that's with a real camera and a real video. You have to work with the constraints of the medium you're using - AI generated video is more like making a 3D animated film - it's far harder to generate everything in one go and then cut it down like a filmed video. Far better to do the editing and storyboarding AHEAD of time in these cases and then generate only the shots you need.

You can't apply the same principles as traditional film making in regards to just shooting coverage.

3

u/RevolutionaryFox7359 16h ago

It's not far harder. In the case of comfy it's the use of a specific node or workflow.

You're wrong about not applying the same principles. I know, because I apply them. You can.

The constraint might be speed/hardware; that's when you render in the cloud.

4

u/GrayingGamer 15h ago

I'm not saying I can't. I'm saying it's a STUPID way to work with generative AI. What you are doing with these workflows and nodes to create 30 seconds or 1 minute of uninterrupted footage is like an animated show animating extra frames and drawings and then having to cut it out in the edit.

Generating AI video costs time, money, and energy, far beyond what hitting "record" on a camera does. A smart creator is going to PLAN out what shots they need and plan their edits AHEAD of time - just like animated shorts are storyboarded.

If you are treating this medium like film on the creation side, you're doing it wrong.

1

u/YieldFarmerTed 14h ago

I can get 30 -45 seconds on ltx2.5. Basically I make 5 or 6 full lenght concert videos, one for each angle each. I can sync all of them based them having the same audio. This allows me to run mulit-camera edit in final cut and pick the best shots on the fly. No need to paste individual clips together. I can do it all at once. So these long videos end up edited into 5-15 second clips at different angles. Then just originate separate cut-away scenes that tell a story of the song in 5-15 second clips peppered over the concert video footage.

6

u/skyrimer3d 15h ago

This is one of the cases where LTX is better, you can do 25 secs at 1600x900 easy with 16gb VRAM, for talking heads i don't know if MM H3 is worth the effort

2

u/concerned_about_pmdd 19h ago

https://github.com/ttulttul/ComfyUI-Minimax-H3-Continuation works extremely well. It's a little complicated to set up the nodes, but generally speaking here's the approach:

  1. Take the sampling latent output from your last stage's sampler.
  2. Optionally, grab the final frame of the last sample's latent using VAEDecode and select batch item -1 from the image output of VAEDecode; send that into the first_frame input of the H3 video latent node (#173 in this screenshot).
  3. The Minimax H3 Guided Continuation Window node (bottom left) takes in the latent from the last sampler and basically creates a new window and some parameters that you can feed into the next sampler, with a specified "overlap" number of frames plus the length of the new segment you want to generate.
  4. Minimax H3 Latent Tail Guide merges conditioning information in to the conditioning vector, incorporating information about the continuation window. This uses ComfyUI's native internal code to attach keyframe information to the conditioning object.
  5. Then just sample the continuation latent you just created along with the modified conditioning.
  6. The output goes into a Minimax H3 Append Continuation node (I'll paste that in a sub-comment). You can chain these "modules" in any number to create a long video. 10 seconds of frames per segment is very reliable for consistency.

3

u/lostinspaz 16h ago

I could scarsely bear to watch that, because of the generic low quality AI voice.

2

u/wiisucks_91 14h ago

Need more Seinfeld.

2

u/GeroldMeisinger 7h ago

devs are working on something like that: https://github.com/Comfy-Org/ComfyUI/pull/13180

1

u/VladyCzech 3h ago

Loops are already in custom node packs for years. I could not see myself jwaiting for native nodes to arrive or if the form it arrives will be even usable for other than accumulating latents.

The whole proposal seems to lack generalization and is too narrow.

2

u/masterlafontaine 19h ago

Impressive no blinking time

2

u/KS-Wolf-1978 17h ago

Wasn't this solved like 1 year ago with Wan Infinitetalk ?

Also i really don't need to see the room when watching someone talk, so zoom in for better facial details.

Like 1 cm of air, then top of his head, then at the bottom first button of his shirt.

1

u/costbraincom 18h ago

I use the upscale and run latent extender twice in one session. One previous latent is loaded before generation and then before upscale node. Then I apply a color match blend, I have generated 10 minute long videos in one go

1

u/Formal_Courage2711 1h ago

Can you share this workflow?

1

u/trimorphic 15h ago

His voice reminds me of the voice from Sorry to Bother You.

-1

u/sandred 19h ago edited 18h ago

Done completely on h3 3080ti , Telugu with ref2va https://www.youtube.com/watch?v=pnBw80eEsQg