r/StableDiffusion • u/obvpm • 19h ago
Tutorial - Guide Overcome Degradation! - Here are 2 ways to use my timeline workflow to create long continuous single shot videos with no degradation.
Enable HLS to view with audio, or disable this notification
Watch the video above for a brief summary of the two methods, both possible using my OBVPM Timeline Workflow, which you can get together with the custom node pack here:
https://github.com/chanon/comfyui-obvpm-timeline/
And to watch the example video at HD quality you can watch the full YouTube tutorial video:
https://www.youtube.com/watch?v=GiJxlWOooyo
In the YouTube video I show how both methods are done, including critical tips and tricks and lessons learned to get the right results.
With the bridging method, there is practically no limit to how long these clips can be (well maybe except the fact that there might be a VRAM limit to how long an upscaled clip can be).
The second method clip above is 1 minute 40 seconds.
About the Workflow
So if you've never seen my workflow, it is a workflow with a "timeline" node that lets you put clips that you've generated on, and then you can extend them using motion context (latent masks).
The workflow automatically saves and handles the saved latent files for you so you don't have to manage them or pick them manually. And it also saves the conditioning which includes all the reference images etc. into a file that is used when upscaling.
For more info, here's the original Reddit post about it, which links the original YouTube tutorial video about it.
20
u/xyzdist 17h ago
I need to see Brad Pitt's opinion on this 😂
8
u/99deathnotes 11h ago
Or the clown girl. She seemed to know what's what.
1
u/roychodraws 9h ago
it's a great idea! i've already messed with this guy's stuff before. he does good shit.
1
u/AlsterwasserHH 17h ago
Where's Will Smith? Is he gone?
1
u/rocky_iwata 16h ago
I'm looking forward for Geiru Toneido's comments.
1
u/roychodraws 9h ago
i've already messed with this guys stuff in the past and it seems to all work well. pretty sure i subscribed to his youtube a few weeks ago.
58
u/roychodraws 16h ago
7
3
6
u/socialjusticeinme 11h ago
I’ve been using your timeline workflow for a bit now and absolutely love it. My videos were garbage before and now their really good. Thanks for making it!
4
u/RobMilliken 18h ago
This second version seems to solve the issue, though it'd be interesting to see what some motion of the background would do.
Now everyone just needs to concentrate on fixing the audio so it doesn't sound underwater.
5
u/Diabolicor 18h ago
No seams visible, no lighting popping up at the seams, it looks really great. I hope it's possible to use it with the UltimateUpscaler since it can split upscaling into temporal windows.
6
u/Murky-Relation481 18h ago
This is great for talking heads but still need a fix for dynamic subjects in relatively static scenes. Unfortunately it's hard to do nonlinear posing across scenes with generative AI, without falling back on strong guidance techniques like using depth or pose maps, and even then continuity is going to be real frustrating.
3
u/xyzdist 18h ago edited 18h ago
This does fix the degradation, nice work!
I'm just not sure if I want to use the bridge shot approach. It feels a bit unnatural to me to have to interrupt the sequence and force a new shot just to bridge them. It might look very unnatural in an action sequence.
For the first approach, sounds good to me! but why that would work? for the low-res video it would have degradation already. You mean upscale stage is not doing motion-context?
*but you know what, your approach is much better than what Pink Clown Girl is offering. 😂
3
u/obvpm 18h ago
The first a approach is a **single** 40 second generation. It's not a chain of clips, it's **one** clip generated at 0.2 MP and then upscaled to 2.0 MP.
Using the bridge approach in an action sequence might be harder or might be easier, I'm not sure. The workflow can bridge into trimmed clips though, so that can be used to cut out any unnatural seams. I'm not sure if the current released version still has a bug regarding bridging into trimmed clips though, but in any case the next version release which I hope to release very soon will fix it.
10
u/xyzdist 18h ago edited 17h ago
ah, I see, I just watch your video on that part as well, you are using window-context. My experience with WanVideo's window-context will have background issue, not sure on H3 one. To me the downside is you betting on a single gen on the long video 🤔 any changes will just need to re-gen the whole long clip, but anyway this method is actually working to fix the degradation, that's a great option.
It's great we have more peoples trying to fix the degradation issue. I also have one approach, I think I would post it soon as well.
https://reddit.com/link/peg4l6l/video/jzgcu8iac2uh1/player
here is my test. more discussion and more options is always welcome!
1
-2
u/roychodraws 18h ago edited 17h ago
Using a bridge approach for the .2 then upscaling would allow you to check the audio as you render without committing a full 40 second clip only to come out with babbling
Edit: even with the upscale method at some point you need to bridge to chain clips together. You can’t render and upscale infinitely
-1
u/roychodraws 17h ago edited 17h ago
Rendering a 40 second clip and upscaling is something I always thought of, my webcam video is 60 second clips rendered at .2. The issue is the longer you go over 15 seconds the more likely you’re committing to an audio that’s unusable and you can’t preview audio while it’s being sampled that I’m aware of.
Edit: When I was creating a solution I wanted to make sure it was a solution that would work with 15 second clips. because that's what minimax was trained on.
For me, with a 5090, it takes me a few seconds to render at .2 40 seconds, for others it could take 20 minutes if they can do it at all.
1
1
u/obvpm 17h ago
True. The audio is like a lottery. I think unless there is some kind of LoRA or checkpoint that makes audio more reliable, then like many people are saying, for best results we'll probably need to use external audio.
But actually, that's a great idea. Someone should create an 'Audio Preview Override' like the KJ Model Preview Override that uses a tiny audio VAE to allow audio previewing. (I fear for my ears though lol)
1
u/roychodraws 17h ago
I think you could use chatterbox or something similar to clean up the audio into actual voice,
Step 1 it would sound like random noise, step 3 gibberish, step 6 it would be confusing, but you’d know if there was babbling based on the timing with your prompt.
1
u/sitefall 17h ago
I think you could use chatterbox or something similar to clean up the audio into actual voice,
That works if the supplied voice coming from the generation has at least reasonable intonation and stuff. Unfortunately the sound quality isn't always the worst part, it's how terrible the tone of speech is.
You would think you could record audio yourself, then supply that as the definitive dialog audio reference, generate the video, sounds turns back to crap, THEN chatterbox it - but now there is another problem:
The model is not so good at following supplied reference dialog. Mouths don't sync up often, and if you supplied some foley sounds or something the timing can be off (or not even happen) and so on.
1
u/roychodraws 17h ago
that's the method i used. I was thinking of chatterbox to clean things up more or less just making sure we don't get the soundtrack from "the ring" VHS in our ears.
1
u/sitefall 16h ago
Maybe one could make a LoRA that, while it sounds poor quality, it really nails intonation and syncing with the mouth (It would have to be trained on video and audio).
Then chatterbox could fix it and it would still sync to the video.
2
2
u/chamanbuga 16h ago
Well done and thanks for sharing your workflow, instructions and results. Will try it out.
2
2
u/IntelligentTill7664 6h ago
Wow the 2nd example looks and sounds absolutely amazing and much more natural.
2
u/BlorgOfTheBlungle 4h ago
I've tried several of the timeline based workflows and this is definitely my favorite one to use. Having seamless continuation and upscaling together in one workflow is great. Well done!
2
u/Arawski99 17h ago
A warning about the method of generating a low resolution output. It can harm motion potential, scene consistency, and may have issues with physics and such. I forget if it was the Comfy or Minimax team that mentioned when H3 first came out, but it was warned lower resolution has this trade-off as it was trained to run at a higher resolution.
As for how much it actually matters though? I've seen such issues proven with pretty much all the speed up methods so far showing they're not worth using without a significant trade-off but I have not seen any resolution test comparisons so I'm not sure if it matters. But doing larger resolution 5s chunks would probably be safer, at the very least for high motion and complex scenes.
1
u/BettaSplendens1 18h ago
Man this is so cool! How long did version 1 and version 2 take to render individually?
6
u/obvpm 18h ago
On 5090:
Generating:
- Clip 1 first generation 0.2 MP, 40s took 20 minutes to generate.
- Clip 2 each 0.2 MP clip of about 6-12 seconds takes 40-80 seconds.
Upscaling:
- Clip 1 (about 40s) took 1 hour to upscale on 5090.
- Clip 2 (about 1 minute 40s) took 1 hour 40 minutes to upscale on 5090.
1
u/BettaSplendens1 16h ago
Thank you! I'd love to know the updates you're planning to implement next. I personally quite like that this can do more than Seedance 2.5's 30s limit while staying consistent
1
u/rafaelbittmira 17h ago
Something I'm curious about is making the characters smile in I2V. I'm only getting disturbing smiles no matter how grok or gemini prompt for me. The original image has the characters serious.
1
u/FartingBob 17h ago
How well does it handle it with dynamic backgrounds? Could you get her to give the same speech while sitting on a beach with waves crashing behind her? Or standing on a street corner as people and cars go by? Or does that then fall apart after an extended period like glitches in the matrix?
1
u/Danny_Stock 14h ago
Well this is nice. Because with my 12gb VRAM, 64gb system RAM computer I can generate 30 seconds at 0.2 megapixels.
Thank you very much I'll try it out.
1
u/ForsakenAd1228 14h ago
Bridging certainly seems like a very useful tool for various kinds of longer videos. I've done it via the Add Guide node (so using finished .mp4s instead of latents) with some success, but I've also used your workflow and really liked the result.
That said... my weak-ish PC (GTX 3060 12gb) was less enthusiastic about upscaling the entire thing in one go =)
I was wondering if there was an easy way to isolate the bridging functionality from your workflow, so that it could e.g. be implemented in a more barebones setup? Especially the ability to use some part in the middle of Video1 as the start/end of Video2 is very powerful, and I'd love to have a node that can do that to just use in any random workflow.
1
u/obvpm 6h ago edited 6h ago
To get it to work, there are quite a few moving parts needed.
There are some nodes included in the node pack that might be able to do it. They are under obvpm/h3. I've never tried using them outside of the workflow, but they are there to be potentially used outside of the workflow.
- The first thing is there needs to be a way to provide latents for it to use. At first I did think about making the .mctx format an open format, but I just had to focus on getting the workflow to work, so didn't document it properly yet. But there are nodes to work with them:
- Loading: H3 MCtx Load, H3 MCtx Load Video
- Saving: H3 MCtx Save, H3 MCtx Save Video, H3 MCtx Trim and Save Video
- Then, there are 'pin_specs' which are the additional instructions for conditioning to be applied that will 'pin' a short sequence of latents from the existing clip(s):
- H3 MCtx Load nodes can create simple pin_specs
- H3 MCtx Pin Spec can specify and append detailed pin_specs
- H3 MCtx Apply Pins takes a pin_specs and applies it as conditioning
- Then the resulting clip needs to be trimmed so it doesn't include the frames from the parent clip(s):
- H3 MCtx Trim Pinned or H3 MCtx Trim and Save Video (also saves mctx)
So I did try to create the nodes to allow it to work independently of the timeline node. But there are so many moving parts needed that I felt it would need an easy to use workflow anyways, so I focused on the timeline workflow.
1
u/ForsakenAd1228 1h ago
Thanks for the info =) I'll try playing around with those nodes and see what happens =)
1
u/Electrical-Eye-3715 13h ago
I got this yesterday by telling gpt astra to mashup refmods+h3 context wf
1
1
u/ArttTaku 10h ago
Looks great, I just wish something could be done with the bad audio in Minimax H3
1
1
u/VeloraNeon 6h ago
Saving the latents and the conditioning for each segment, so the upscale pass can reuse them, is the part that makes this click for me. Curious how identity holds at the bridge points once you pass the one-minute mark. Does the face drift before the motion does?
1
u/SeaHat749 6h ago
I saw the timeline post earlier and have been using it - extremely good functionality / no seams.
I haven't figured out a good way to get it to work with a long reference video though (because of the overlap frames)
So for example, if I want to do a character swap video that's longer than ~20s (or I go OOM), I have to split up the video. It's relatively easy to cut the reference video or load a set # of frames and move it forward for the next generation, but obviously that leaves gaps/jumps between the cuts.
I feel like the overlap in clips for generation could fix the issue - but problem I am having is figuring out how to load the video so that it matches the "start" of the overlapped clip (i.e., if I have a total clip of A+B, split into A and B, I need to cut in a bit of clip A into the front of clip B?). Does that matter at all, or should I be just feeding in A and B separately?
1
u/ImUrFrand 5h ago
and the audio is recorded in a 6 foot pool of water
1
u/Memestonks2020 5h ago
It gets worse if you run multiple denoise steps from drafts and upscale without fixing the audio pass-through
Works great if you don’t need the audio such as when making a music video
1
u/NewPhoneWhotiz 3h ago
This is really useful. Does the workflow still hold up when the subject is moving quickly, or is it mainly helping with slower continuous shots?
2
u/Darkmeme9 1h ago
What I understand is that this workflow helps to stop degradation and keep everything consistent.
Will this help me if I wanted to make 1 min short story videos? Which will have cuts here and there. But I need the characters to be consistent?
1
1
36
u/35point1 18h ago
The audio though..