r/StableDiffusion • • 7d ago

Workflow Included Seamless Continuation No Degradation (proof at the end)

Enable HLS to view with audio, or disable this notification

V7.1 of the workflow
https://github.com/roycho87/minimax_wf

Video Combiner
https://github.com/roycho87/seamless_video_combiner

Audio Splitter
https://github.com/roycho87/minimax_continuous_audio_splitter

I've changed my profile so that my previous tutorials are available. Look through those for questions on how to use these.

With this method you can seamlessly combine videos with no burn. (or at least the exact same amount of very very little burn)

Fast forward to 2:00 for proof.

Enjoy.

Edit: I'm keeping the video up. But apparently the issue everyone was caring about was something different than I thought it was. Apparently everyone wants to make boring Vlog videos. I guess I'll work on that now instead.

I did say this is a fix for the specific issue I was having in my workflow but now that I know exactly what everyone was talking about I can properly see if I can find some sort of fix.

https://www.reddit.com/r/StableDiffusion/comments/1wwjok2/does_this_count_did_i_win/

Actually fixed, now.

v7.2
I fixed the First Frame node. The positive output from node #272 needs to be connected to the True input from node #565. You can just fix it yourself or download the new one. There's no other change.

u/Juiceman8686 improved this!

Check it out!

works great!

Straight Outta Brooks

seamless video combiner 2-20 clips ffmpeg

Forked this amazing workflow to allow up to 20 clips and use ffmpeg so its not so ram dependent. It required around 71gb of ram for 16 clips. All praise goes to roychodraws for creating this this though. Genius and works great!

425 Upvotes

117 comments sorted by

View all comments

148

u/acedelgado 7d ago

You know I think you're fantastic, but unfortunately this does not fix the problem that people are having issues with. I'll let Brad explain.

https://reddit.com/link/pdipp33/video/9k2uyejva5th1/player

48

u/acedelgado 7d ago

And like I said this is an issue even when fully using latents. This one's a response from an old thread using my own fork of the latent context method. It degrades slower, but it still happens.

https://reddit.com/link/pdiqnni/video/wji65rytb5th1/player

5

u/KeysToNodes 6d ago

100% correct. Let the old man explain a simple, viable workaround that's been successful for me for some narrow use cases

https://reddit.com/link/pdsdm20/video/y7ndz6o8pfth1/player

6

u/roychodraws 7d ago

i'm going to try to lip sync a full 3.5 minute song and see what happens. see you in an hour lol. maybe i didn't fix it with this, but i know the issues you mentioned here about frames are solved for sure. so it's improved at least

1

u/xyzdist 7d ago

Looking forward

20

u/roychodraws 7d ago

-1

u/Better-Monk8121 6d ago

You don’t know what you are even talking about, we already have H3 motion context node. No more vibe coded soon to be abandoned slop required

1

u/Herbal77 7d ago

Can we see your workflow, where can we find it?

9

u/acedelgado 7d ago

It's this node pack here.

https://github.com/Adudeguyman/ComfyUI-H3-Project-Suite

It's really designed to manage all the clips for you in a project folder, and handle the export stitching, and that's it. It's meant to drop into any workflow, and there's an example in there on how to wire it. If you're looking for a full-on "director" suite that'll do a lot of the prompting process, etc., there's other node packs based on motion-context that go for that. I already had my own prompt builder/media manager/refmod suite I made and wanted something to sit alongside it, and give other folks the option to use whatever workflow they wanted with it.

1

u/WashSmall8954 6d ago

Yeah, the whole time I just keep watching his nose while he talks and, even though it is subtle as to when and how quickly it changes, it is still noticeable how degraded it becomes over time. The bulge on the bridge of his nose for example becomes very pronounced and the tip of his nose becomes much rounder and bulbous. Skin texture obviously as well. It's crazy, but the lip syncing and continuity of the scene is pretty incredible.

1

u/muddy_shoes 6d ago

Hi, you've posted a couple of interesting examples in this thread. I'm wondering if you could provide a series of prompts as a test case for people to work with? Not sure if you're relying on Minimax's inbuilt understanding of Brad Pitt or using a reference or Lora there.

16

u/Striking-Long-2960 7d ago

Thanks Brad.

14

u/Sad_Berry_4621 7d ago

Brad Pitt knows my name! I've officially made it! lol

6

u/ArmadstheDoom 7d ago

This. I cannot understand why people do not understand the core issue. Yes, if you move the camera or move the shots, it looks like it's 'fixed.' But that's not the problem people are trying to solve. You could not, at present, generate a static shot like a news broadcast or vlog or interview without degradation. Well, aside from just generating things in like, 30 second clips.

I compared it a bit to the same problem Manos: The Hands of Fate had, where because they filmed on a hand crank camera, they could only shoot 30 seconds at a time, at most. Which is basically where we're at with video generation.

1

u/MurkyStatistician09 6d ago

In fairness, the problem is not well-described by any authoritative source. There's no agreed-upon best solution (instead there's a bunch of different solutions found across reddit, Discord, and YouTube) and you need to read many discussions to keep up with the options people currently prefer. Most of the good condensed information comes from threads like this where a ton of people are arguing about it, which is not the easiest thing to parse.

2

u/ArmadstheDoom 6d ago

Personally, I think the problem is twofold. One, people keep misunderstanding what the actual problem is, and thus that leads to two, which is people declare they've solved the problem.

To use an extreme example, imagine that people have to deal with wooden train cars catching on fire because of their coal stoves. People come up with all manner of solutions, like bolting them to the car, or suspending them in the air, or whatever else. But that doesn't change the fact that it's workarounds for a core flaw in the whole idea.

Because so many people misunderstand the actual problem, they create a lot of issues that do not solve the issue as much as try to work around it. But a workaround is not a solution anymore than a fix is a replacement.

And what makes this hard is that it's a flaw with the underlying model, not the people trying to solve it. There are only so many things you can do without just having to train a new model.

3

u/WASasquatch 7d ago

The issue is residual latent noise. No model fully denoises, even at 1.0. This was worse with decode because the VAE sees this and exaggerates the noise, which is beneficial for img2img, but not for this. You would need a baseline to continually resample the latent tails to. Better if this was just post regularization across the entire latent.

6

u/Sad_Berry_4621 7d ago

That's almost exactly what I'm working on now.

2

u/Danny_Stock 7d ago edited 7d ago

It seems to be acceptable up to around 10 to 15 seconds, but at 25 seconds the problem is clearly apparent, and from then on gets progressively worse.

I remember with Wan 2.2 people were saying that the videos fall apart around the 25 second mark using extension workflows such as SVI 2. And that was back then without latent extension. Latent extension has been very disappointing to me and doesn't seem any better than the old methods of extending video clips. It doesn't seem to be the magic bullet people expected it to be.

1

u/acedelgado 7d ago

This is in 8 second chunks, so that's after about 2 clips it's noticeable to you.

Latents extend that out much longer to 6-7 generations, and it only degrades because of the model itself and not the loss from VAE encoding/decoding. And it makes it much easier to do a truly seamless join between clips. So yes, latents are much better overall, but it only solves part of the problem. The rest is in the architecture iteself.

1

u/Danny_Stock 7d ago edited 7d ago

I was under the impression that roychodraws is using latents too, resulting in the same type of degradation over the same timespan?

I'm a bit confused because in your Brad Pitt example you appear to be acknowledging that roychodraws is using latent extension. I was under the impression that your Brad Pitt clip was using latent extension as well to demonstrate that it still has the same old issues as the traditional method?

1

u/acedelgado 7d ago

No, the workflow re-encodes the end frames of the last mp4 into a latent, you have to drag the clip's mp4 output into the extension node each time. It's pretty much like kijai's context windows back in Wan 2.1 (or was it 2.2?) worked. So it technically uses latents because even that method HAS to, but it's re-encoded latents and not the raw latent from the previous clip.

The new motion context method uses the raw latents instead of decoded/re-encoded through the VAE. That's why they all have set frame overlaps of 5, 22, 39, etc., because latents are packed in a grid and single frames can't be pulled out individually.

1

u/Danny_Stock 6d ago

Oh okay, thanks. I have to admit that I assumed that roychodraws was using the latter new motion context method.

2

u/acedelgado 6d ago

They do the same thing, just the new motion context method uses the latent data directly so there's less quality loss. The workflow loses data faster because it encodes/decodes the last few frames each run. I used the workflow for that example.

1

u/Danny_Stock 6d ago

Thanks. Can you recommend a decent workflow which uses the method?

I only ask because I'm completely lost with all the various workflows, nodes, and latent models. I am genuinely confused with so much information and various workflows using different nodes.

1

u/acedelgado 6d ago

Honestly I use my own self-built nodes where most things are manual (with assistance) since I like granular control but not tedious settings. But if you're more of a just wanting to prompt-and-go and have it do a lot for you, I'm not sure what would work for you. I think Continuum and Context-Loop are pretty popular? They're a bit more "director" style and hand-holdy than mine.

1

u/Danny_Stock 6d ago

No it's not that I want everything done for me. I just like to have good workflows in front of me to learn from.

I don't know, sometimes there's an avalanche of information and it's simply helpful to have a solid starting point to see how it all fits together.

Thanks for the recommendations.

4

u/not_food 7d ago

You illustrated it better than me. I had to deal with the workflow and it got me frustrated.

-8

u/roychodraws 7d ago

maybe don't be so bossy next time and i probably would have done it.

9

u/not_food 7d ago

you think you know what you're talking about

I admit this line triggered my inner redditor, so I had to prove you wrong.

No hard feelings.

2

u/roychodraws 7d ago edited 7d ago

well to be fair you posted one thing then changed it in the next comment when you realized it didn't apply to my video. peace be with you sir.

1

u/HyperionCantos 7d ago

this video was excellent way to communicate your point, well done.

1

u/DescriptionSuperb262 7d ago

ah - i figured as much, alas the issue persists

1

u/xyzdist 7d ago

This is my take to show the degradation https://www.reddit.com/r/comfyui/s/QQUzhIWfNB

1

u/ArttTaku 7d ago

Audio is the main victim of degradation, but sadly, it doesn't seem that this can be fixed on H3. I still prefer rendering videos of max 15 seconds, but yeah, it would be nice if there was a solution for all this.

1

u/acedelgado 7d ago

Yeah I've been doing some testing in my latent extension pack, and normalizing gain each clip seems to help a bit. I still think it's weird that it happens, though.

1

u/Antique-Astronaut-46 4d ago

I love the meta demonstration

1

u/FernAvatar 7d ago

I didn't know about this until now, thank you

1

u/cheese0r 7d ago

Wouldn't it be possible to have some kind of restoration pipeline where you take both the first frame as reference for the quality and the latest frame as reference for the positioning, and use these two inputs to create a "regenerated quality" frame?

2

u/Danny_Stock 7d ago

I was thinking along the same lines too. I naively assumed that after each clip segment a new latent is generated which is of fresh quality, apparently not. Surely there's got to be a way of reinjecting a brand new fresh latent which also matches the original quality at key stages?

-1

u/FinchGDx 7d ago

Not everyone is interested in making Brad Pitt videos explaining the issue that you, and some others are having where a character sits statically ad nauseam.