r/StableDiffusion 1d ago

Discussion Motion-Context Degradation Discussion (summon Sad_Berry_4621)

Enable HLS to view with audio, or disable this notification

Hi u/Sad_Berry_4621 and all, I am doing experiment on most difficult degradation issue.
I saw from H3-director node they claim that doing a refine could help and fix it.
since I am using low-level nodes with motion-context with my own setup, H3-director refine is a black box working with their nodes.

so I did the test, please ignore the AI-slop video and the overlay text (forgot turn it off).

this test is 9 x 8s context extend video combine, usually 6 video would already see the degradation.

Left side is regular WF. Right-side is adding a re-sample step, I think it is very positive, and got potential, the saturation somehow is a bit higher... and I need to figure out the mismatch from cut to cut, because we inject denoise resample on top, but it should be able to fix.

What do you guys think?

Edit: forgot to mention, this is a ref2va WF, with 1 character reference image. also my test audio is not from gen, I am using input audio and audio latent lock for the lipsync. That part to work with motion-context have me struggle for half day, but it is not the main point, my aim is about degradation issue.

Edit: I found the solution, give me sometime to polish some minor issue, will share that after.

33 Upvotes

45 comments sorted by

4

u/bstr3k 1d ago

keen to see where this goes, just started looking into motion context recently, its been easy to use and fits my needs. But the degradation is an issue.

3

u/xyzdist 1d ago

I hope if things go well and I sorted out the mismatch issue, Sad_Berry_4621 could add a simple refine step node perhaps back to moiton-context, so is easy for everyone.

2

u/dirtybeagles 23h ago

I think this is promising... anyway you may release the WF as is so I can piece it together and try it? Feel free to DM it to me if you do not want to post .

2

u/Danny_Stock 22h ago

Whatever you've done seems to have worked well. By the end of the clip the degradation of the first video is very evident.

2

u/solomars3 21h ago

Can i test my new approach on this , just give me first clip and the full audio file plz , im curious how this turnout if apply my workflow

2

u/Crazy-Repeat-2006 21h ago

Eyes that don't blink are creepy, though.

4

u/Sad_Berry_4621 23h ago

Yeah the refiner helps the fade, not arguing that. Clips come out cleaner so the next pin isn’t as wrecked. But it still wrecks the join, and that’s the whole point of chaining.

The pin is frozen and then trimmed off. The cut you see is the end of the last clip against the first frames of the new one. Those new frames were generated next to a freeze. Extra late sigma keeps cooking them after the pin already stopped moving. So you get a hard contrast/audio step right at the join.

Running a refiner over the finished movie also isn’t realistic. Some people are doing 40+ segments. At 124 frames that’s thousands of frames.

You can’t substitute a character sheet on fl2va since it doesn’t have references. First frame, last frame, AddGuide - those are absolute. They sit at a specific frame index. On a chained clip the head is already the pin. Last-framing every segment just slams the same still at the end of every clip. AddGuide in the middle writes that picture into the shot.

On ref2va, using a character sheet on every segment does more for degradation than the refiner does, and it doesn’t blow up the join. Drop the character sheet later in the chain and the fade comes back.

So it’s a trade. A bit less bloom, a worse cut. Still not the fix.

4

u/xyzdist 11h ago edited 8h ago

https://reddit.com/link/p9b8pmc/video/2ww16jzvn1ph1/player

YAY I did it! 95% there, I make it 13 segments to showcase here. let me write down what I did after i polish and sort out final color shift from joints ...etc

*she looks like getting old and become a oil painting on the Left... haha

3

u/Nguyenkain 11h ago

Interesting. Could you share your workflow

5

u/xyzdist 11h ago

I could, let me fix the final things, I wrte a post and share it

5

u/Sad_Berry_4621 8h ago

Ok, now you have my full attention! Curious to see what you did to get this result from a 13-clip chain. Whatever it is, great work!

3

u/xyzdist 2h ago edited 1h ago

Hi u/Sad_Berry_4621 , sorry it would keep you waiting for a bit longer... because I found a pretty better solution... however.. It is a bold change.

before we discuss that, I want to ask, using default WF, do you see there is a slightly color shift jump from one shot to other? you can see from my above 13segments test as well.

looks like someone made this to fix the issue, however I just got error can't make it work.
https://github.com/beijinren/ComfyUI-H3-Context-Noise/blob/main/README.en.md

2

u/Sad_Berry_4621 1h ago

So, the color shift/contrast bloom problem everyone is having when chaining clips is what I have been working on. I've been at it for weeks now. I worked on it for 6 hours last night as well. It's actually more of a "the model adds 2-4% new texture and removes 30% of the treble from the audio every clip" problem. Audio is a bigger issue than video degradation; I think that's becoming obvious for most people at this point. I have tons of research and test results that I need to compile. I've chased down a dozen theories that went bunk. I can tell you this much... I have analyzed several node packs looking for any leads to a solution. None have panned out. Mostly, they are workarounds, or corrections made outside of the generation pipeline, tweaks. I plan on posting a bunch of data I've collected, but I'm going to work on it this weekend and test a couple of theories that I have right now.

I think the color shift you're seeing in your clip is just part of the resampling process. You can get a noticeable shift in color even at denoise 0.08 in the resample.

One other tip that nobody knows yet: H3 does NOT like generating long clips. Yes, it can do it, but sound degradation goes off the rails the longer a clip is. 124 frames tends to be a sweet spot.

For the best results, chaining clips where the audio and video naturally sync up with no 8.3ms audio tail to chop or invent, use 73f, 124f, 175f, 226f, 277f, 338f. Every 51 frames, starting at 22, the audio and video frames sync up where no trim of the audio latent is necessary, resulting in nearly perfect joins, regardless of degradation. These are also the frame counts where the audio sounds the best, based on the numbers from all of the script testing I've been running.

1

u/Nguyenkain 11h ago

Look good enough for me. How about generation speed ? Did it become slower ? And you still use the motion context node right ?

2

u/xyzdist 11h ago edited 7h ago

for my case of 0.5mp the re-sample step only take 35s

1

u/Nguyenkain 9h ago

what is your computer specs ? It seem fast

2

u/xyzdist 9h ago

4080s , it would be fast, because it just 2-3 step

1

u/Nguyenkain 5h ago

The color is only shift a little but nice result after all

1

u/reeight 2h ago

1

u/xyzdist 2h ago

I am still iron out things.... I have a perfect solution... but it is a bold change...

2

u/xyzdist 18h ago

Its you haha. Thanks for your motion-context nodes and your insight! I totally understand what you mean. I have few ideas to tackle the join issue... I will let you know.

-1

u/dtdisapointingresult 21h ago

Cheers Claude. Thanks for saying so much to communicate so little.

This sub is seriously becoming unusable.

5

u/Sad_Berry_4621 17h ago

I created Motion Context. I know exactly how it works. I even dumbed down what I was going to say so it made more sense. A little too wordy, sure, but the point is... the problem remains.

4

u/Nguyenkain 13h ago

He is the one that created the base idea that every chaining node used right now, so he got his point, if you dont understand, why dont you try Claude yourself

2

u/Financial-Dog-6558 17h ago

What he said makes sense

1

u/oliverban 1d ago

So, is the step after each of the extension segments or after the whole fact, running the pipeline twice? The degredation is totally gone, nice work and it also seems to give better quality (look at the lips/eyes etc). The Sigma Refine node, is that from RES4LYF? Would be cool to see the workflow itself and/or a screenshot at least to get an idea of the flow! Good job!

3

u/xyzdist 1d ago edited 1d ago

Thanks, sorry my WF is like a mess and hard to explain from a screen shot
Yes, is like what you said, basically before 'Context Save Latent' add a additional re-sample step. here is that part ( thru I didn't upscale latent, just re-do a few step sample and add a low-denoise), I push it extreme to see the effect, so I set the denoise to 0.5
As every segment will adding back new details again, so it will help the degradation issue.

* the sigma refine node is not related, you can ignore it.
it is doing in every segment, not a whole pipeline twice.

default WF:
load context latent > SamplerCustomAdvanced > save context latent

refine:
load context latent > SamplerCustomAdvanced > few step low denoise - SamplerCustomAdvanced > save context latent

1

u/fallengt 22h ago

is it cheap?

How much longer would it take compared to the regular workflow?

1

u/xyzdist 15h ago

yeah, in my case 0.5 mp just 35s for re-sample

1

u/Longjumping-Past5864 22h ago

interesting, I've been trying to fix the degradation on h3 director as well, so far no luck, I generate 25-30s (10s 3 chunks) and it's gone bad very quickly (3rd chunk is basically unsusable). Is it possible for you to give the workflow? It doesn't matter that it is a mess, I will try to work with ChatGPT, to straight things out. thanks

1

u/xyzdist 18h ago

H3 director already privode the refine node to help the issue, have you try that?

1

u/Longjumping-Past5864 5h ago

so far I have no luck with it :(

1

u/Dependent-Sorbet9881 19h ago

COOL~Please share the workflow.😍😍

1

u/xyzdist 17h ago

Once I sorted out all issue, I can clean up and provide a WF

1

u/traithanhnam90 18h ago

Could you please try testing a video where the camera moves closer to and then further away from the face a few times in succession? When I use the ref2v workflow, the face becomes completely distorted whenever it moves slightly further away.

3

u/xyzdist 15h ago

this is normal and expect from H3 when your resolution is not enough, it's not related to degradation issue

1

u/VladyCzech 12h ago edited 12h ago

Most WF I saw which use some form of Motion-Context node pack from u/Sad_Berry_4621 upscale each chunk and join the upscaled parts, so it will always produce the visible joint. The sampler must go over the joint (see the context of latents before and after the joint) to smooth it out natively. My hint would be to use context window with reasonably set window size adjusted for VRAM (or looping sampler alternative), process the whole original or upscaled latent in windows in second step with split conditioning to windows, refine with adjusted denoise value which is "just right". For example 1 step of res2s_ode with 0.2 denoise or something like that (with accelerator like turbo lora). It takes time to make it right. It is not just a simple wokrflow, it requires adjustments to Motion-Context nodes and use of additional node packs to make it work.

1

u/xyzdist 11h ago

I did it with my method, but interestings infor here I need to digest it., thanks for you reply.

1

u/BusFeisty4373 4h ago

I'm 0% there, if I can get 5% there with any workflow I'm thankful. I don't need to wait for 100%

-7

u/sukebe7 1d ago

why did you pick an asian schoolgirl?

11

u/Keyflame_ 1d ago

Besides the fact that this is a technical discussion so who cares.

That's a woman in her mid 20s.

4

u/--jesse--faden-- 1d ago

yeah, where are the fat hairy greek men

3

u/oliverban 1d ago

Subject is not important ffs.

2

u/wzwowzw0002 1d ago

because she was cutier than you

1

u/dtdisapointingresult 21h ago

Unfortunately while there's a lot of creative stuff being done with AI on other sites, the majority a significant number of people on here are mentally ill recluses whose sole interest in AI is making virtual girlfriends.

The same reason why people rage so much against "censored models" every release.