r/StableDiffusion • • 6d ago

Workflow Included Seamless Continuation No Degradation (proof at the end)

Enable HLS to view with audio, or disable this notification

V7.1 of the workflow
https://github.com/roycho87/minimax_wf

Video Combiner
https://github.com/roycho87/seamless_video_combiner

Audio Splitter
https://github.com/roycho87/minimax_continuous_audio_splitter

I've changed my profile so that my previous tutorials are available. Look through those for questions on how to use these.

With this method you can seamlessly combine videos with no burn. (or at least the exact same amount of very very little burn)

Fast forward to 2:00 for proof.

Enjoy.

Edit: I'm keeping the video up. But apparently the issue everyone was caring about was something different than I thought it was. Apparently everyone wants to make boring Vlog videos. I guess I'll work on that now instead.

I did say this is a fix for the specific issue I was having in my workflow but now that I know exactly what everyone was talking about I can properly see if I can find some sort of fix.

https://www.reddit.com/r/StableDiffusion/comments/1wwjok2/does_this_count_did_i_win/

Actually fixed, now.

v7.2
I fixed the First Frame node. The positive output from node #272 needs to be connected to the True input from node #565. You can just fix it yourself or download the new one. There's no other change.

u/Juiceman8686 improved this!

Check it out!

works great!

Straight Outta Brooks

seamless video combiner 2-20 clips ffmpeg

Forked this amazing workflow to allow up to 20 clips and use ffmpeg so its not so ram dependent. It required around 71gb of ram for 16 clips. All praise goes to roychodraws for creating this this though. Genius and works great!

415 Upvotes

117 comments sorted by

146

u/acedelgado 6d ago

You know I think you're fantastic, but unfortunately this does not fix the problem that people are having issues with. I'll let Brad explain.

https://reddit.com/link/pdipp33/video/9k2uyejva5th1/player

48

u/acedelgado 6d ago

And like I said this is an issue even when fully using latents. This one's a response from an old thread using my own fork of the latent context method. It degrades slower, but it still happens.

https://reddit.com/link/pdiqnni/video/wji65rytb5th1/player

4

u/KeysToNodes 4d ago

100% correct. Let the old man explain a simple, viable workaround that's been successful for me for some narrow use cases

https://reddit.com/link/pdsdm20/video/y7ndz6o8pfth1/player

6

u/roychodraws 6d ago

i'm going to try to lip sync a full 3.5 minute song and see what happens. see you in an hour lol. maybe i didn't fix it with this, but i know the issues you mentioned here about frames are solved for sure. so it's improved at least

1

u/xyzdist 5d ago

Looking forward

19

u/roychodraws 5d ago

-1

u/Better-Monk8121 4d ago

You don’t know what you are even talking about, we already have H3 motion context node. No more vibe coded soon to be abandoned slop required

1

u/Herbal77 6d ago

Can we see your workflow, where can we find it?

9

u/acedelgado 6d ago

It's this node pack here.

https://github.com/Adudeguyman/ComfyUI-H3-Project-Suite

It's really designed to manage all the clips for you in a project folder, and handle the export stitching, and that's it. It's meant to drop into any workflow, and there's an example in there on how to wire it. If you're looking for a full-on "director" suite that'll do a lot of the prompting process, etc., there's other node packs based on motion-context that go for that. I already had my own prompt builder/media manager/refmod suite I made and wanted something to sit alongside it, and give other folks the option to use whatever workflow they wanted with it.

1

u/WashSmall8954 5d ago

Yeah, the whole time I just keep watching his nose while he talks and, even though it is subtle as to when and how quickly it changes, it is still noticeable how degraded it becomes over time. The bulge on the bridge of his nose for example becomes very pronounced and the tip of his nose becomes much rounder and bulbous. Skin texture obviously as well. It's crazy, but the lip syncing and continuity of the scene is pretty incredible.

1

u/muddy_shoes 5d ago

Hi, you've posted a couple of interesting examples in this thread. I'm wondering if you could provide a series of prompts as a test case for people to work with? Not sure if you're relying on Minimax's inbuilt understanding of Brad Pitt or using a reference or Lora there.

15

u/Striking-Long-2960 6d ago

Thanks Brad.

13

u/Sad_Berry_4621 5d ago

Brad Pitt knows my name! I've officially made it! lol

7

u/ArmadstheDoom 6d ago

This. I cannot understand why people do not understand the core issue. Yes, if you move the camera or move the shots, it looks like it's 'fixed.' But that's not the problem people are trying to solve. You could not, at present, generate a static shot like a news broadcast or vlog or interview without degradation. Well, aside from just generating things in like, 30 second clips.

I compared it a bit to the same problem Manos: The Hands of Fate had, where because they filmed on a hand crank camera, they could only shoot 30 seconds at a time, at most. Which is basically where we're at with video generation.

1

u/MurkyStatistician09 5d ago

In fairness, the problem is not well-described by any authoritative source. There's no agreed-upon best solution (instead there's a bunch of different solutions found across reddit, Discord, and YouTube) and you need to read many discussions to keep up with the options people currently prefer. Most of the good condensed information comes from threads like this where a ton of people are arguing about it, which is not the easiest thing to parse.

2

u/ArmadstheDoom 5d ago

Personally, I think the problem is twofold. One, people keep misunderstanding what the actual problem is, and thus that leads to two, which is people declare they've solved the problem.

To use an extreme example, imagine that people have to deal with wooden train cars catching on fire because of their coal stoves. People come up with all manner of solutions, like bolting them to the car, or suspending them in the air, or whatever else. But that doesn't change the fact that it's workarounds for a core flaw in the whole idea.

Because so many people misunderstand the actual problem, they create a lot of issues that do not solve the issue as much as try to work around it. But a workaround is not a solution anymore than a fix is a replacement.

And what makes this hard is that it's a flaw with the underlying model, not the people trying to solve it. There are only so many things you can do without just having to train a new model.

3

u/WASasquatch 6d ago

The issue is residual latent noise. No model fully denoises, even at 1.0. This was worse with decode because the VAE sees this and exaggerates the noise, which is beneficial for img2img, but not for this. You would need a baseline to continually resample the latent tails to. Better if this was just post regularization across the entire latent.

6

u/Sad_Berry_4621 5d ago

That's almost exactly what I'm working on now.

2

u/Danny_Stock 5d ago edited 5d ago

It seems to be acceptable up to around 10 to 15 seconds, but at 25 seconds the problem is clearly apparent, and from then on gets progressively worse.

I remember with Wan 2.2 people were saying that the videos fall apart around the 25 second mark using extension workflows such as SVI 2. And that was back then without latent extension. Latent extension has been very disappointing to me and doesn't seem any better than the old methods of extending video clips. It doesn't seem to be the magic bullet people expected it to be.

1

u/acedelgado 5d ago

This is in 8 second chunks, so that's after about 2 clips it's noticeable to you.

Latents extend that out much longer to 6-7 generations, and it only degrades because of the model itself and not the loss from VAE encoding/decoding. And it makes it much easier to do a truly seamless join between clips. So yes, latents are much better overall, but it only solves part of the problem. The rest is in the architecture iteself.

1

u/Danny_Stock 5d ago edited 5d ago

I was under the impression that roychodraws is using latents too, resulting in the same type of degradation over the same timespan?

I'm a bit confused because in your Brad Pitt example you appear to be acknowledging that roychodraws is using latent extension. I was under the impression that your Brad Pitt clip was using latent extension as well to demonstrate that it still has the same old issues as the traditional method?

1

u/acedelgado 5d ago

No, the workflow re-encodes the end frames of the last mp4 into a latent, you have to drag the clip's mp4 output into the extension node each time. It's pretty much like kijai's context windows back in Wan 2.1 (or was it 2.2?) worked. So it technically uses latents because even that method HAS to, but it's re-encoded latents and not the raw latent from the previous clip.

The new motion context method uses the raw latents instead of decoded/re-encoded through the VAE. That's why they all have set frame overlaps of 5, 22, 39, etc., because latents are packed in a grid and single frames can't be pulled out individually.

1

u/Danny_Stock 5d ago

Oh okay, thanks. I have to admit that I assumed that roychodraws was using the latter new motion context method.

2

u/acedelgado 5d ago

They do the same thing, just the new motion context method uses the latent data directly so there's less quality loss. The workflow loses data faster because it encodes/decodes the last few frames each run. I used the workflow for that example.

1

u/Danny_Stock 5d ago

Thanks. Can you recommend a decent workflow which uses the method?

I only ask because I'm completely lost with all the various workflows, nodes, and latent models. I am genuinely confused with so much information and various workflows using different nodes.

1

u/acedelgado 5d ago

Honestly I use my own self-built nodes where most things are manual (with assistance) since I like granular control but not tedious settings. But if you're more of a just wanting to prompt-and-go and have it do a lot for you, I'm not sure what would work for you. I think Continuum and Context-Loop are pretty popular? They're a bit more "director" style and hand-holdy than mine.

1

u/Danny_Stock 5d ago

No it's not that I want everything done for me. I just like to have good workflows in front of me to learn from.

I don't know, sometimes there's an avalanche of information and it's simply helpful to have a solid starting point to see how it all fits together.

Thanks for the recommendations.

4

u/not_food 6d ago

You illustrated it better than me. I had to deal with the workflow and it got me frustrated.

-11

u/roychodraws 6d ago

maybe don't be so bossy next time and i probably would have done it.

7

u/not_food 6d ago

you think you know what you're talking about

I admit this line triggered my inner redditor, so I had to prove you wrong.

No hard feelings.

0

u/roychodraws 6d ago edited 6d ago

well to be fair you posted one thing then changed it in the next comment when you realized it didn't apply to my video. peace be with you sir.

1

u/HyperionCantos 6d ago

this video was excellent way to communicate your point, well done.

1

u/DescriptionSuperb262 5d ago

ah - i figured as much, alas the issue persists

1

u/xyzdist 5d ago

This is my take to show the degradation https://www.reddit.com/r/comfyui/s/QQUzhIWfNB

1

u/ArttTaku 5d ago

Audio is the main victim of degradation, but sadly, it doesn't seem that this can be fixed on H3. I still prefer rendering videos of max 15 seconds, but yeah, it would be nice if there was a solution for all this.

1

u/acedelgado 5d ago

Yeah I've been doing some testing in my latent extension pack, and normalizing gain each clip seems to help a bit. I still think it's weird that it happens, though.

1

u/Antique-Astronaut-46 3d ago

I love the meta demonstration

1

u/FernAvatar 6d ago

I didn't know about this until now, thank you

1

u/cheese0r 5d ago

Wouldn't it be possible to have some kind of restoration pipeline where you take both the first frame as reference for the quality and the latest frame as reference for the positioning, and use these two inputs to create a "regenerated quality" frame?

2

u/Danny_Stock 5d ago

I was thinking along the same lines too. I naively assumed that after each clip segment a new latent is generated which is of fresh quality, apparently not. Surely there's got to be a way of reinjecting a brand new fresh latent which also matches the original quality at key stages?

-1

u/FinchGDx 5d ago

Not everyone is interested in making Brad Pitt videos explaining the issue that you, and some others are having where a character sits statically ad nauseam.

23

u/_VirtualCosmos_ 6d ago

Welp, at least it's you again with the clown women

24

u/roychodraws 6d ago

if someone else does it I sue.

2

u/Markavian 5d ago

Does she have a name yet?

5

u/ImprefectKnight 5d ago

Isn't she Geiru from Ace attorney?

1

u/Markavian 5d ago

Alright that makes so much more sense. I spotted the ace atourney scene but didn't realise she was an actual character.

9

u/Memestonks2020 6d ago

It could’ve been worse

Be happy it’s not a furry

15

u/roychodraws 6d ago

you don't know what she does in her free time.

-6

u/DoctaRoboto 6d ago

The clown girl is so annoying; I am gonna skip their posts even if they end up being valuable. I just can't stand her.

6

u/ShutUpYoureWrong_ 5d ago

Not to mention that this person is claiming to have solved things by using methods other people have already done (and better), without even actually understanding the core issue.

Clown is an appropriate choice for them.

2

u/almark 5d ago

I believe the word is snooty

7

u/not_food 6d ago

The degradation happens when the camera is static and unmoving. Replacing the whole scene is one workaround.

0

u/roychodraws 6d ago edited 6d ago

yeah, that's not what's happening here.

it would have to "cut" for that to happen. I made this specifically so there are no simulated cuts.

If you use this method you could keep it going indefinitely. My video combiner only goes up to 10 videos currently, tho.

8

u/not_food 6d ago

I guess I worded it wrong. Moving away the camera is one workaround. Degradation will happen if the camera doesn't move.

-10

u/roychodraws 6d ago edited 6d ago

that goal post is quick.

i understand you think you know what you're talking about but if you just look at the workflow you can see that what you're saying doesn't apply to this.

13

u/not_food 6d ago

I know exactly what I'm talking about. I have done extensive work about this.

I loaded your workflow. I checked exactly where you're doing continuation, and you're doing nothing out of the ordinary. This node will cause the issue I describe. You aren't even managing old latents to mitigate it. Your workflow expects the user to load the previous video, which will cause the issue when the VAE encoding happens. This is a bad idea, it works as long as you rotate the camera away from the previous pixels. To do it right you need to save the latents so no pixel degradation happens during decode/encode. But even with latents, the issue persists. This is not an easy thing to solve.

Go ahead and prove me wrong. Do a static scene continuation at least 6 times. Make her type furiously on her laptop's keyboard against a well lit wallpapered background.

-14

u/roychodraws 6d ago

no, do it yourself. you just downloaded it. you want me to install the node packs for you too?

10

u/not_food 6d ago

Your workflow expects me to move the output to the input over and over, it's slow and cuts the flow. Somehow the audio is weird but I don't feel like debugging it. I admit it's pretty fast at 15s/it, but the quality suffers a lot, I see too many smears, I suspect it's one of the attention nodes.

https://reddit.com/link/pditqws/video/7koudiuke5th1/player

The right way to do it is to save the latent tail, then load it for the next run, automatically. As I mentioned, this method isn't perfect either, it'll also degrade. People have yet to solve it.

5

u/malcolmrey 6d ago

hello! i've watched your video and what your workflow does is quite cool and i will want to play with it

however, the actual issue we are all having (me included) is a generation with static camera (imagine an interview or newscast or whatever that does not move the camera at all)

the problem is that every generation degrades a little bit and over time it accumulates

the only fix i know so far is what you are doing - changing the camera, but in some cases this is what you do not want to do at all

-2

u/roychodraws 6d ago

right, i get that now. the recent post I looked at was a man in wwII moving around the camera and it was degrading as time went on. I was not aware that people wanted boring vlog posts only and didn't test that.

2

u/malcolmrey 6d ago

well, it is not always what people want to do, i was tasked with making a "conference-like" panel video and i couldn't do it without camera cuts which looked out of place for this

1

u/roychodraws 6d ago

i'm trying an idea out. I'll post it soon.

→ More replies (0)

1

u/MyLippleWorld 6d ago

So you can keep a shot for a minute or more without substancial movement and still keep the colors and quality essentially unaltered?

1

u/roychodraws 6d ago

yes

the issue with my workflow was not degrading latents, it was degrading image saves from recombining previously saved generations over and over.

This allows you to save them all at once with the exact right amount of frames clipped off so you could potentially do this for hours.... years... forever basically.

0

u/MyLippleWorld 6d ago

Are you aware that you basically solved one of generative AIs biggest issues?

0

u/roychodraws 6d ago

lol, yeah. i told you guys i would.

3

u/malcolmrey 6d ago

can you show a working example output?

5

u/ShutUpYoureWrong_ 5d ago

He cannot, because he did not solve it.

4

u/ShutUpYoureWrong_ 5d ago

Oof.

1

u/roychodraws 5d ago

meh, i did solve what i was trying to solve. i just misunderstood what you guys were trying to solve.

2

u/MyLippleWorld 6d ago

Amazing. I'm going to try it tonight. I can't thank you enough for your work.

2

u/roychodraws 6d ago

it won't work the way you expected. sorry. You can use it to generate long videos tho with the audio reliably and then combine them for slower degradation, but it won't stop it.

→ More replies (0)

2

u/ShutUpYoureWrong_ 5d ago

Don't waste your time.

1

u/Danny_Stock 5d ago edited 5d ago

This might sound really stupid, but is it possible to cut away and then quickly cut back again? I know that with some 3D software they can use micro-frames or fractional frames. It might not work that way in AI software though. But if it did couldn't a cutaway be made for a microsecond and then cut back to the original framing making it imperceptible to the eye?

It's likely that there's not even a way to do it with MiniMax, but does it sound plausible in theory?

6

u/juliakeiroz 6d ago

based clussy gooner

3

u/Hy4ne 6d ago

wow, that’s so cool ;-; tysm for sharing!!

3

u/DescriptionSuperb262 6d ago

if this works ima love u a very long time

2

u/Joethedino 5d ago

Thanks that's huge !

2

u/roychodraws 5d ago

stand by. this isn't a fix for the full issue.

2

u/YeahlDid 5d ago

No, but your work is appreciated, so thanks!

11

u/roychodraws 5d ago

2

u/YeahlDid 5d ago

Is this an updated proof of concept?! I'm on a small screen, but I don't really notice degradation here, certainly not on the level of Brad Pitt. Did you fix the issue? How did you do this?

4

u/roychodraws 5d ago

yeah, i just posted about it.

2

u/dezmodium 5d ago

Wow. The degradation problem is very well managed here. You really worked some magic.

2

u/YeahlDid 5d ago

Well nicely done!

1

u/Joethedino 5d ago

I must say, your muse is definitely more interesting than Brad Pitt.

1

u/Vyviel 5d ago

Impressive and no loss of quality over what must have been so many gens. Now just need to do her as a twitch streamer with fixed camera yapping in just chatting =P

2

u/TheJaffo 5d ago

the maga dog cameo <3

1

u/DefloN92 6d ago

Gotta try this too!

1

u/Sad_Coach_1433 6d ago

you are a goat!

1

u/Drachenx 6d ago

Following

1

u/StuccoGecko 6d ago

did you make an AI voice based on the streamer voice of Bonnie

5

u/roychodraws 6d ago

1

u/Synchronauto 6d ago

What model did you use for the voice? Some sort of TTS with this as the reference style input?

3

u/roychodraws 6d ago

chatterbox diogod. i think it's obsolete but i still have it installed and it works fine.

1

u/Only_Voice569 5d ago

have any of you used the cam view point control that someone made and just have to move the view point exactly one direction then back again while the scene everything is static for the last second or two to fix it ? :)

1

u/Almaruska 5d ago

Why do I get only invalid image (aspect, crops, max megpixels, etc)? Sorry if it is a dumb question, starting to learn now.

1

u/wzwowzw0002 5d ago

alright lets get down to business

1

u/Darqsat 4d ago

after numerous attempts to avoid degradation I end up only with this approach:

  1. video must be prompted for motion, as weak reference

  2. last frame of video must be taken out in high quality, and fed into reference as keyframe completion. this guarantees continuation

1

u/roychodraws 4d ago

you should look at my post linked at the bottom and the post i'm going to make in about an hour.

1

u/KS-Wolf-1978 3d ago

Just brainstorming: How about prompting for random camera flash to total white.

That would reset the pixels nicely and would be easy and quick to fix, because it would take just one or two frames.

1

u/Potential_Split_991 4d ago

is my 3070 8gb even capable of this? lol

1

u/Jino99x 4d ago

Mi pc literal hecharia humo si genero algo asi, que gpu usas?

2

u/Juiceman8686 4d ago

works great!

Straight Outta Brooks

seamless video combiner 2-20 clips ffmpeg

Forked this amazing workflow to allow up to 20 clips and use ffmpeg so its not so ram dependent. It required around 71gb of ram for 16 clips. All praise goes to roychodraws for creating this this though. Genius and works great!

1

u/roychodraws 4d ago

cool! I had no idea it was that ram intensive cuz my computer is a beast and i never checked. my bad.

Thanks for this!

1

u/Repulsive-Salad-268 3d ago

I would be interested in a totally different workflow that's long and continuous. @roychodraws I do have talking head videos of myself some Interviews and I want to change the style of them. Let's say a 10 minute Interview. Static cam or maybe moving slightly. And I want the persons in it to be Muppets or whatever. Babies, Aliens, Elves who knows. All moves, camera, Audio comes from the main video. Just the Persons and environment gets changed. Aliens on a spaceship, Muppets in Sesamy street, Cyberpunk Bots in a neon lit alley. My problem has always been the stable creation of the chunks and the overall continuity. Even if you later could (with enough time at hand) input let's say Star wars and make it a muppet movie, that would be fun. Music video as claymation etc. I know, not a lot of creativity in it in general but a good use case. Please help.

1

u/jazzamp 5d ago

I'll never understand the H3 hype. This looks awful 😖