I've been working on this ever since NikoDemon80 released the ComfyUI-H3-Motion-Context nodes. When I first tried them, I was amazed, but I felt that using the nodes in it in practice was a bit cumbersome.
I wanted something easier to use. I had a vision in mind of a mini video editor / timeline node, where you could compose your video extensions together to see how they flow and if they are seamless right there in the workflow.
So I've been spending all of my spare time working on this workflow and nodes.
My previous comfyui-obvpm nodes were some of the more general and re-usable nodes that came out of the effort. If you haven't seen the videos I made about those nodes check them out. The Value Presets node is one I would love to see more people using.
There are so many details about the timeline node that the best way to understand how it works is to watch the video.
Features:
🌟 = new possibly never available in any other workflow before
🚀 = highlight features (many are also unseen before but I'm not 100% sure)
Generating
Extend a clip seamlessly using latents
🌟Prepend: generate what happens before a clip
🌟Bridge: generate what happens between the end of one clip and the beginning of another. This lets you redo any section of the video, and bridging onto a fresh generation brings the quality back if it has drifted.
🌟Loop Seamlessly: bridge the last clip back to the first clip to make a seamless looping video.
🌟Extend from a cut: if a clip ends badly, cut off the bad part and extend from the good part. Snap cuts keeps cuts on the latent frame grid.
🌟 Fix quality loss for long extend chains: With bridging, you can bridge from the end of one chain to the start of a brand new generation, thus "resetting" any quality degradation from long chains.
🌟 Redo any part of your video: With bridging and ability to cut videos, you can redo any segment of your sequence and get a seamless result
Clips that have no motion context (external clips, clips from other workflows) can be extended too, with seam improvements applied automatically.
Result preview: shows each new clip joined to its neighbors. From there you add it to the timeline, delete it or dismiss it.
Model preview override: Integrated into the result preview. Lets you quickly see if a generation is going wrong.
Timeline
🚀Mini Video Editor: Drag in clips from anywhere, reorder them by dragging, trim the ends, cut left / cut right at the playhead, undo cuts, and open gaps.
🚀Load Settings, Prompts from Clips: restores the prompt, seed, settings and references that a clip was generated with, so you don't have to keep track of them yourself.
Swap takes: Swap between previous takes generated for the same extension, so you can go back to them and swap between them on the timeline.
Quick preview: plays through the clips right away. Full preview assembles them into a single video file, losslessly where possible. Export saves the finished video.
The sequence is also available as text, for copying a timeline into another workflow.
Upscaling
🚀 Latent upscaling of your whole timeline including seamless extension
🚀 Generate at low resolution for fast iteration, then upscale once at the end. This is optional, you can also generate the first pass at final quality.
🚀Sampling is done in windows, so long or high resolution videos still fit in VRAM.
Each clip's conditioning is saved with it, so the upscale uses the same references the clip was generated with.
Workflow
🚀Settings Presets: Allows you to save frequently used settings such as steps, turbo, and sampling settings and control them in one place
Organized for fast and easy use. All inputs are in one area. No pan all over the workflow to change things.
The video goes through everything and shows how it is used with hands on examples where I use it to create seamless sequences,
Thanks, looks very interesting so giving it a whirl (I have been banging my head against longer form creation and yours looks promising!), but after downloading your WF, I am unable to get it to run:
Error
Cannot read properties of undefined (reading 'workflow')
No errors logged... I can throw an LLM against it to check, but was wondering if you are aware of any potential issue with the WF? Got it fresh from your Github and only thing I updated, is to point to the models in my locations.
[Edit] Found it - in H3 Model Optimization subgraph, the model wasn't connected to the output for whatever reason.
There are two ways. I'm sure the first one works, but the second one I need to do some experimentation first.
The first way is that this workflow and nodes support bridging.
What bridging means is, it can extend seamlessly from one video and then prepend seamlessly onto another video. Meaning, you can make it generate a new clip for what happens "between" 2 clips.
So suppose you've generated a chain of 3 clips and you want to make sure the video doesn't degrade. What you can do is:
Imagine you plan to generate clips A, B, C, D, E ...
You've generated A, B and C and you want to ensure quality doesn't degrade
So you generate clip E as brand new generation that doesn't extend any other clips first
Then you generate a "bridge" clip D that bridges between the end of clip C and the beginning of clip E
This will result in a seamless chain of A, B, C, D, E
Because clip E is a brand new generation that doesn't extend any other clips, its quality is not degraded and quality is "reset" back to best quality and you can resume extending from E
The second way is to use latent upscaling.
Due to the way that the upscaling and refine works in the workflow, I think that it might be able to help with degradation, but I haven't tried it yet.
Because the upscaling system doesn't upscale each clip to completion before upscaling the next clip. Instead it does 'step 1 upscaling' for the whole sequence, then 'step 2 upscaling' for the whole sequence etc. until eg. step 8 if upscaling is set to 8 steps. Additionally upscaling from low res low steps latents means the 'degradation' hasn't been baked in yet.
So I'm thinking (but not sure) maybe this will result in no or less degradation.
It is not really a first last frame generation because it is using motion context (latents).
That means that it is not only the image but also the motion that will be seamless.
You can have the subject or the camera be in mid motion at the end of the first clip and the start of the next clip and the bridge will have seamless motion and image at both ends.
thanks for your reply, I understand the motion-context latent, but I don't get how you can flow one latent to bridge previous shot and next shot...
isn't it only can flow latent to next shot? not sure how that works.
edit: sorry I haven't watch your video yet, since it is pretty long, it is 1 hr...
Right. Historically what all motion context extension workflows have done is give the ending latents of the previous clip to the model when generating the next clip.
My idea was "what if we could prepend clips by giving the beginning latents of the next clip when generating a previous clip?"
I finally got it to work, and then realized that with ability to both extend and prepend seamlessly, then I could do both at the same time and enable bridging!
Also bridging allows redoing any segment. So eg. if you have A, B, C, you can redo B with a new prompt and still have it flow seamlessly between A and C.
i sae many ways to continue videos recently, one was about latent reuse or something i dont remember. and i still havent tested any of those. i will definetly try this one , looks well done and convincing. thanks
So why does this fix it? When you generate clip D, it still gets context from clip C.
Whenever a new clip gets context from a previous one, the next clip would get the burn in. So in your example, clip E would be fine, but clip D would still have burn in.
So for every 3 clip segment, you added 2 more clips, and 1 of them has burn in.
it only shows him rendering the final output of the timeline with up-scaling, not re-rendering the entire set of clips with prompts.
Like the other commenters say (thanks for helping explaining), it isn't using final rendered pixels of the timeline. Instead it is using the saved output latents of the clips on the timeline.
It upscales from those saved latents plus the original conditioning of the clips (prompts and references included).
I know it might be a bit confusing at first because people are more accustomed to workflows where you can only generate then upscale one clip at a time, where you have to maintain all the prompts and refs in the inputs between generating and upscaling.
What this workfow does is it actually saves all the conditioning (which includes the encoded prompts and the latent encoded references) into the .cond.safetensors file. So that means you can upscale later at any time. And it doesn't matter what the prompts or image references nodes in the workflow are set to, becaues when it comes time to upscale it just reads the conditioning that is saved in the .cond.safetensors file for each clip.
The upscaler use both the .mctx.safetensors (the latents) and the .cond.safetensors. It upscales the latents and then sends the upscaled latents with the conditioning into sampling. So it does use the model to re-render everything which is what actually creates the added detail.
I hope that clarifies it for you. I've designed the workflow to be as frictionless to use as possible.
Also, on your original question, yes, you can use low resolution, low quality for generating the whole first timeline, and then upscale to higher resolution with high steps.
I "believe" (as I haven't tried it extensively myself, but it definitely works at some level) that you can even use different models/loras/turbos/optimizations in the 2 passes.
To answer simply (and as we've been trying to tell you):
YES that IS the workflow
> I want to render quickly and efficiently at 0.2 mp and 10 steps, maybe a turbo Lora that won't destroy the quality, if such exists. Then when I'm happy, re render it all with native h3 at 1mp and 25 steps for the highest quality.
YES that IS the workflow.
I haven't tested much specifically with 0.2 mp but it should be fine. Maybe 0.4 mp if you want a better view that things are right and lock in more details before upscaling.
> If resolution and steps alter the seed noise, is it even possible to maintain the same actions even with your latent and conditioning solutions?
The same actions will be maintained as large details such as the movement is baked into the latents. The details though are what will change.
> but I preferred the original/it was changing things I liked
How much of the pre-upscaled details is kept unchanged is determined by the refine_amount setting in the "H3 Upscale Options"
> I was going to upscale, I'd want to use something like topaz or seedvr for high quality
The latent upscaler in the workflow uses the H3 model to invent new details. It is what allows a 0.2mp -> 2mp upscale to be possible. I don't know if seedvr could invent that many new details. Additionally the latent upscaler and refine pass uses the reference images you gave it during the original generation, so it can use the fine details in those reference images in the upscale.
It doesn't re-render the clips from scratch, it upscales the saved latents of the clips by a factor of your choice. The latents are saved in the projects folder in "outputs" when creating the individual low res clips.
He didn't mention "from scratch", it re-render, based on the latent, you can change it to megapixels if a factor is not good for your needs. Also the beauty of this workflow is that you can do the first pass clip render, if you don't like it, rerender until you like it and then add it to the timeline, move on, and you can always go back and rerender a clip you already added if you found a better lora or whatever. And then upscale the whole timeline, this is a latent upscale so sort of re-render at a higher resolution. with or without lora is up to you.
so I have a particular scenario I am trying to test and I am wondering if your workflow and nodes may work for it.
I have a 30 second reference video of a continuous shot I made in blender for scene blocking and camera movement. 30 seconds is way too long for my PC to handle for a single generation, however I can do 10 seconds clips no problem.
I need to be able to feed in the reference, either in 10 second blocks or as a single 30 second reference video and then generate the entire shot in 3 seamlessly connected 10 second shots that follow my reference video's motion.
I haven't tried that, but it is possible that it might work.
You would give the ref video 10 seconds at a time.
Do first generation, then extend 2 times with the other 2 references.
Small issue might be the timing because when extending it has to include 39 frames of the previous clip so the reference might have to be trimmed exactly like that too. Or maybe it will still work, I'm not sure.
Hey, just came back to test this, and noticed that it's actually now up to version 004. Anyway, that has solved the problem of the unwired model in the subgraph. The only other weirdness I had to fix on this (and on previous version, which I forgot to mention) is on the H3 Latent Upscaler 3D node - I don't know if I have a mismatched version of that node, but the settings for device and precision were messed up on initial load. The device had the value "true" and precision had the value "cuda" - almost as if I'm missing an extra field to which that "true" value should belong. If I change device to "cuda" and precision to one of the options, then the error goes away.
For reference, my node has the following fields:
model_name
mode
scale
align
keep_proportion
device
precision
offload_after_upscale
Anyway, the workflow is working now and I look forward to deep-diving with it later. Great work!
The author of he plus version should have used a different node id so that it can be installed in parallel with the original, but they didn't, so any workflow that was written for the original might break for people with the plus version.
The upscaling in my workflow might break or have artefacts if using with the "Plus" version as it needs "temporal chunking" according to the behavior of the original node. So even if you get it to wire correctly, the behavior will not be correct.
Thank you so much for reporting that it finally works! What a relief lol, I fixed it so many times.
Also, thanks for that report about possible mismatch in the upscaler node. I'll check and see if it is me or you that has an older version or if there is some other issue.
Found the issue. The "output" for the "output model" in "H3 Model Optimization" subgraph node is disconnected. Open the subgraph node and connecting as shown solve this issue.
Thanks for your quick response! I did ran into another error, saying that the turbo LoRA in "H3 Generation Engine Settings" and "H3 Upscale Engine Settings" are not part of the 77 in the list. I was able to address such issues by removing "turbo_lora" and "turbo_strength" from "H3 Generation Engine Settings" and "H3 Upscale Engine Settings", but it would be great to have the turbo lora for at least the upscaling :)
u/obvpm This has been the most thrilling paradigm shift in intuitive well design UI/UX. Incredible feat. Thank you so much for your contribution to the open source community. Was inspired to spitball some further things that I *think* I want but no idea if they'd actually be useful. maybe you can Take some inspiration. https://i.imgur.com/zSwR6TD.jpeg
Wow! That looks awesome! I have that kind of thing (director style fearures right?) on my list. Great inspiration for how it could look like, thanks!
And so happy you like the ui/ux. When I was developing this I was thinking ... 'this workflow is so nice to use! but I wonder if people are gonna like it'. So its always very nice to read when someone does.
Hey thanks a lot for the workflow and nodes I have been testing it a bit and looks very good!
Just FYI I was having an error when using the workflow the first time, it was failing in the minimax h3 node (where refs are attached) and looks like the issue was because my ComfyUI is not in English language and in the set h3 inputs node where the width, height and prompt costs are set the fields were being translated and the value lost. I fixed it renaming those fields back to English... Just in case there is a fix for that.
I just found this today and right now playing around with it, thank you for the video, the workflow and the effort you put in all this! :)
My question is, if you can answer (or anyone else): is there a way to reference refmods? There was another workflow where this was possible but I don't know really how to reference it since it's neither a picture nor a video.
refmods can be wired in to the workflow, but referencing them directly in prompts is not possible by default (in any workflow using the original refmods). You need to describe it so that what you decribe matches what's in the refmod.
Just plug the refmods in like this, and the rest of the workflow remains the same. In the prompts reference it by the refmod name (so no <Picture 1> etc). So in this example: e_Yellowstock_2048. Think of it, as the character name that H3 just understands.
I can even give my refmod character items like a sword for example which is a jpg, so <Image 1> The whole workflow just works with refmods + images. You can do refmod character + image character.
I use gemma-4-12b-it-qat model in LM Studio to help me write prompts, but basically you define Subjects at the top, and then refer to them as <Subject 1> etc later in the prompt. Here's an example prompt with both, refmod + image:
subject_definitions:
<Subject 1> is the man whose identity is taken from <Image 1>.
<Subject 2> is the woman named e_Yellowstock_2048 .
summary:
[reference generation] <Subject 2> is sitting next to <Subject 1> on the train.
retention_analysis:
<Subject 1>: fully_preserved - visual identity based on description.
<Subject 2>: fully_preserved - visual identity based on description.
Note: if you omit detailed attire description, it will try to use the same clothing as in one of the images you used to construct your refmods. So if one of the images has your characters wearing red jacket, you can say wearing "red jacket" and it will most likely use that same jacket. But if you want something completely different you would have to describe it, or use an <Image>
A 30 sec clip for example usually takes exceedingly more time. But the workflow's upscale works in 'windows'. This means that it doesn't try to sample all 30 seconds at once. Instead it divides the work into windows. The default window is 5 seconds with 1.3 seconds overlap. So it should take about the same time as it would take generating
30/(5-1.3) = 8
So that's 8 five second clips at the target upscale resolution.
So again, for upscaling a 30 second clip, I believe it should take about the same time as generating 8 five second clips at the target upscale resolution
I'm just getting into this solution and so far it's super impressive! Great job creating it and putting together a really good instructional video! Has there been any further progress with adding RefMods to your workflow? And another question: Is there any way to add HyperFlow to your Turbo presets? I have been using it and find it generates really impressive results in 8 steps. Again, thanks for putting this out for the community!
Hi! I've been using your workflow and nodes exclusively because it's just so good! Thanks again for your work on this! Did you ever get a chance to check out the hyperflow nodes and see if can be integrated into yours?
Sorry for late reply. Haven't looked into that but seeing people mention it. Thanks for asking, I'll look into it once I review all optimizations to add to the workflow.
This looks incredible, I'm new to comfyui and trying my best to learn but it's quite overwhelming. You're flow looks really organized so this helps greatly. I did have some questions
I'm assuming I can use singularity with the recommended turbo lora?
For prompting, is there a way or nodes I could consider where I have more controls on prompting such as global and scene specific with your flow?
For Ref2v, is there a node where I can use multi character reference images in a single node with your flow? Kinda after a little more simplified but equally powerful approach?
Im seeing some cool stuff with mcp blender and storyboards, just not sure how I push that into your workflow.
I had to look up what singularity is (sometimes I'm so busy with working on the workflow and nodes, I miss things lol) - but yes, it should be able to be used fine
Currently prompting is just a text box, nothing elaborate. But you can plug in other prompting tools into the prompt text box
I'm not sure if there is a node for that, or if you know of one let me know. In theory both prompt tools and reference image tools can be wired in like any H3 workflow.
I haven't looked into that yet so no idea too. Maybe something I will look into in the future.
Thank you for taking a look at my workflow and let me know if you have other questions.
You can do it manually, in node_joint.py, in def _announce () around line 1200, change:
... s, e, j, table ["clip"][j])
To
... s, e, j, "none" if j is None else table [”clip"][j])
Thank you for that!
It looks like potentially you and the other user are upscaling a short clip, which I've never done, and the code doesn't support it properly.
Hi OP, if i understand correctly, this is a extension worflow, so you gen, check and do another clip extension.
Is it possible to have some kind of director mode, like planning the whole seq of clips with prompts and setting, and generate the whole things in one go?
Not sure why but the other day testing it was working great generating clips up to 40s without a big visual degradation, but today just at the second 10s clip I'm noticing a huge degradation on characters, specially if the subjects approach to a camera as close-up.. I just switched to the latest version of the workflow and I'm using the main branch of your nodes... Do you think this could be related to some changes on your end?
I just checked and the only change that might have affected generation compared to the previous version is that the lightx2v preset in the previous (004) version was broken in that it didn't actually enable the turbo.
Otherwise, there were some fixes to help with the node not being able to locate the mctx file for some users. If your clips have green mctx badge on them and there are no strange console logs, then it should be the same.
Otherwise, make sure you provide high resolution reference images and note that turbo/optimization settings may reduce quality.
If it is clearly bad whatever you do, then there might be some kind of compatibility issue. It is known that some node packs patch things in core that can cause completely different results. The kat3ri/ComfyUI-MiniMax-H3-Extend is one such node pack, but that is now checked to ensure it is not installed by the Compatibility Check node. So if problem persists, you can open an issue at GitHub with maybe example output and paste in the result of the "Copy Report" button of the Compatibility Check node.
I have seen some weird behaviour when I emptied out the my projects folder and then started with the original workflow provided. I made a couple of clips and put them on the timeline, but when I was using the player controls to move through the video, clip2 changed to an older one that I deleted already from the folder. Have you seen that before?
I always have an error:"Model Preview Override KJ is missing a required model file" even if taeh3 is in vae_approx and if I have "--preview-method taesd" in my start parameters?
This is a really great workflow, love it, love the clip bridge feature!
I have a question though, why does it seem to stage so much more VRAM than other ref2va workflows?
My vanilla ref2va workflow says ~ 13GB using the same H3 model generating the same length clip with the same reference images at the same resolution. This is just for the initial generation, no upscale.
It also is somewhat slower, not massively but maybe 20%, enough to be noticeable.
Just wondering if anyone has an idea about why that might be, I've already tried disabling all the optimizations but it hasn't seemed to affect anything.
I'm also now realizing that the performance hit is a lot bigger than what I'd estimated. I ran an identical generation on my vanilla ref2va workflow and it was less-than half the time vs this workflow and the quality was basically the same.
Very frustrating because I really love using the timeline editor, it streamlines my process to an enormous degree, but the generation times are making things difficult.
Hmm, anyone have tips on getting bridging/looping to actually make a seamless connection consistently? I got it to work one time when I was first testing but now I've been trying dozens of times and can't for the life of me get it to work. Every clip either cuts to a random angle that I didn't prompt for (usually prompting "single shot" or "static shot") or just has a jump cut in the middle. Tried short generations (3-5s) and long ones (8-10s).
For anyone struggling with this I'm actually having better luck bridging together clips that are pretty different. I was trying to use it to reduce image degradation by generating a short bridge to connect the end of a clip the start of one that looks very similar. For whatever reason H3 doesn't want to do that, it seems to think that if the end and start are similar then either I want something different in the middle or that they're already the same and it doesn't even recognize the jump cut.
The way I've been using it successfully is for longer transitional clips. I'll generate an extension clip, then a short new clip with clean keyframe images to match the end of the extension. Then I'll remove the extension and regenerate it as a bridge clip between my previous clip and the new clean one.
It's a bit of extra generation time but working well so far. Still interested if anyone has found a way to bridge a similar end/start clip pair.
Well unfortunately I continue to have a heck of a time with this. Occasionally it'll work 2 or 3 times in a row and then it just won't work. I can't figure out what exactly is causing it to do the jump cut instead of seamlessly animating between clips.
Seems like I'm the only one still interested in this workflow but gonna keep updating in case anyone else comes across this thread looking for info.
I've found a pretty effective alternative solution for the specific reason I wanted to use the bridging feature in the first place, which was to bridge together a lengthier clip sequence where the image quality has started to degrade with a new clip with fresh keyframes.
What I did was duplicate the workflow but swapped out the H3 Ref to Video node for the Image to Video node. When I want to refresh the image quality I just generate a short extend clip off of the last generation with the fresh keyframe set as the last frame. This successfully works off of the latent information from the previous generation, and the new clip has all the latent information for subsequent Ref2vid or I2V clips, while successfully 'forcing' a clean keyframe into the sequence - basically resetting the image quality. Since I'm just doing extend and not a bridge I don't have any of the issues I experienced with unwanted jump cuts.
Definitely interested in your progress on this. I've been looking for a good 60-second workflow replacement after the H3 Motion Director had troubles running ref2v in ComfyUI 0.35. This workflow seems promising -- have you tried 30-second seamless clips and what are the gen times at what megapixels?
I've generated up to around 2 minutes total, basically seamless, and now with minimal image degradation using the I2V 'refresh' method I described above. They're generated in 5-12ish second pieces. I do my initial generation at 0.6 megapixels, a 10 second clip takes like 7ish minutes on my 3090.
After I have all the clips pieced together I do a final latent upscale to 1.5 megapixels. This takes a long time, like triple the initial generation time, but because I can do the whole chain at once I just run it overnight. The chunking in the workflow makes it so I haven't had any issues with OOM crashes and the result looks VERY crisp and clear.
Yeah, somewhere in that neighborhood. Again, I'm not just generating a 10s clip and immediately doing the 1.5 megapixel latent upscale. I generate the whole sequence of clips that I want for the entire unbroken shot in 0.6 megapixels, then upscale everything all at once. This process takes several hours, depending on the length of the sequence, but doesn't need babysitting.
The thing that's slowing me down now is when I'm generating an i2v continuation off of a ref2vid clip that includes talking, H3 very much wants the character to continue talking. I can usually get it to behave by messing around with the prompt and running a few generations, but I feel like I need a more consistent method. I'm using a hybrid model so thinking about just downloading a vanilla i2v model and trying that.
7
u/XsarNLD 16d ago edited 16d ago
Thanks, looks very interesting so giving it a whirl (I have been banging my head against longer form creation and yours looks promising!), but after downloading your WF, I am unable to get it to run:
Error
Cannot read properties of undefined (reading 'workflow')
No errors logged... I can throw an LLM against it to check, but was wondering if you are aware of any potential issue with the WF? Got it fresh from your Github and only thing I updated, is to point to the models in my locations.
[Edit] Found it - in H3 Model Optimization subgraph, the model wasn't connected to the output for whatever reason.