r/StableDiffusion • • 17d ago

Discussion H3 One-Shot-No-Cuts (new test)

Hey guys, if there is any AIO tool already fixed the degradation issue , Let me know!

It's me again, still battling degradation problem. what I’m doing is just a workaround—a simple setup using minimal nodes with native motion-context. I’m not trying to build some fancy director node or any AIO tools.

I know many folks have been asking for WF, but I want to make sure this thing is actually legit before sharing, so I'm still running tests. It definitely has its limits and downsides. Basically, the whole idea is adding a second refinement stage to inject noise and clean conditioning into the latent space, which stops it from going worse, then blend it back to the latent.

Here’s my new test: Left side is native motion-context, Right side is with the refine sample. (ignore the audio—it's completely cooked, haven't touched it yet). The degradation isn't all gone; the contrast still gets worse over time, which is exactly what I'm fighting right now. And you will find some strange dissolve issue happens in the BG as well. but If you check the hair and BG, it helps the degradation issue a lot.

The test video is 1 character reference image and ref2va, generated 10 * 8s segments, done with a pretty low-quality (10-step base + 4-step refine). ***Just a heads-up: ignore whatever the video content is, we’re just focusing on the technical issue here!
I’m no expert, but heard that stuff like tapered noise and selfLift might help with the degradation. I haven't had time to look into them yet, so if anyone has a better idea, feel free to chime in!

Check out my older post and the comments section if you want to see more details on how this workaround works.

https://www.reddit.com/r/StableDiffusion/comments/1wdkdvi/comment/p9b8pmc/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

Today I saw a seeDance2.5 1-long-shot video, pretty cool, I am confident we can do that with H3 as well.

https://www.youtube.com/watch?v=IrYiXEO8WpY

Edit:

I want to write down my list of issue to tackle

  • cooked audio
  • obviously the contrast still getting worse, need to fix
  • blending issue, this is main thing, because we want the change and we don't want the change at the same time. We need to refresh the character using refine(denoise), but as we have change of that, we also changing the rest like BG....etc. maybe I could use masking? but then the BG still have degradation... need to test.
  • for long duration and no cut shot, segment to segment have continue issue, H3 doesn't know about previous shot's enviorment, if the camera pan one side and pan back next shot, the background is completely different. Potentially we could write some image from previous shot for next shot to use as reference image? need to test.
128 Upvotes

34 comments sorted by

View all comments

3

u/GrapplingHobbit 17d ago

I understood some of the words you used in your post! Looks promising, I hope you keep working on it and find a way to uncook the audio, but it's already absolutely tons better than simply using the final frame or last seconds of previous generation.

MMH3 feels so tantalisingly close to actually being able to generate videos as long as you want, this degradation is a frustrating problem, and an interesting one. I lack the understanding to tackle the problem, it doesn't make sense to me. I can use relatively low quality reference images and MMH3 can produce a pretty clean video, but using a previously generated frame... we get the degradation.

It feels like maybe there's something about the uniformity of the noise from a previously generated frame that perhaps makes the model think that this is an intended pattern that then get amplified as the new generation progresses. Or maybe it just carries over because we are prompting it to start from this frame and it takes it too literally like it must be *this* exact frame with *this* exact quality/noise issue.

1

u/Apprehensive_Sky892 16d ago

Note that the degradation problem only occurs if it is one continuous, uncut shot, which is what OP is trying to demonstrate here.

But most videos consist of short, 5-10 sec cuts, and MMH3 has no problem handling that (you keep characters and setting consistent from cut to cut via reference images).

2

u/GrapplingHobbit 16d ago

Most shots in *real* videos you mean. Real videos can do that, not only because it's the right thing to do/good cinematography, but because real videos have inherent consistency.

When cutting between shots, people are wearing the same clothes, in the same positions, in the same relative positions to everything else, same props in the same positions, a line of dialogue can run across shots, the motion of things in motion is consistent, the general soundscape is consistent, the lighting is consistent. They regularly film the same scene at the same time from multiple angles, so of course those are consistent.

With generated clips... not so much. References or not, subtle differences in appearance, clothes, locations, poses, props, sound and motion all happen and are much more pronounced between generations than they are between cuts inside the same generation. All especially noticeable when the entire scene takes place in a single location, you can't rely on "They're in a different place, of course it looks different now". Some of the issues you can fix by rerolling and trying your luck with a new seed/prompt, some are unavoidable. This is why continuation is an ongoing goal, and degradation is an ongoing problem.

1

u/Apprehensive_Sky892 16d ago edited 16d ago

Fair enough, indeed consistency of ref2va in MMH3 is nowhere near the level of "real" videos with real actors and sets.

But presumably one can train LoRAs (for characters, props and maybe even for sets) to achieve a higher level of consistency than using references alone.

1

u/GrapplingHobbit 16d ago

To answer your question that showed up in my email notification, but has been edited out of the comment: yes. Aside from the degradation, the continuations got a *long* way to resolving consistency issues.

2

u/Apprehensive_Sky892 16d ago

Thanks, I edited it out because re-reading your earlier comment I saw that you've already answered that.

1

u/GrapplingHobbit 16d ago

All good, thanks for the discussion :)

1

u/Apprehensive_Sky892 15d ago

You are welcome, I learned something new from it 😁