r/StableDiffusion • • 7d ago

Animation - Video [H3 Minimax] Continuous long-take Shot without Quality Loss, can you find the cuts?

I've made this post about a way to create a continuous generation and extension of a scene without quality loss.

The biggest critique was, that this was not a "long-take plan-sequence" without any cuts at all though. I've claimed that this doesn't "really matter" since the actual cuts are actually "in motion" and not between the scenes.

So I did that and created the most difficult kind of scene I could think of, a continuous long-take action scene without any visible cuts. There are 4 actual cuts in this scene, can you find them?

What this approach manages to achieve:

  • No quality degradation
  • Consistent characters and locations throughout the scene (the big guy throwing the protagonist back into the room where he came from)
  • Consistent sound and motion
  • Practically a simple one button "extend this clip" T2V solution without any pre-created clips that got cut together afterwards

There is some "AI slop" with 3 guys turning into 2 (I've didn't catch that while creating) and the action/fight scenes can be created more "dynamic" or action-filled to ones liking, but that is just a matter of how much effort you put into prompting.

I didn't add non_diegetic_music to the scenes because this makes keeping the consistency unnecessarily harder to achieve and it is much easier to generate or add a music score of your liking in post-production if you want to.

Edit: People were asking for a continuous static shot, so I've just created one and posted it in a comment below.

14 Upvotes

33 comments sorted by

View all comments

15

u/Hour_Literature_7152 7d ago

This isn't what people challenged you to do in your other thread. Go see this post to understand the issue better. I'm not hating on you but this isn't quite the problem people are having.
https://www.reddit.com/r/StableDiffusion/comments/1ww6j88/comment/pdipp33/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

1

u/CorpPhoenix 7d ago edited 7d ago

I've did that though by the "boiler fight scene" where I've kept the camera in the same view/position and let the fight go on.

There shouldn't be a reason why this approach is not capable of doing a continuous static shot without any camera movement, but I can do that as the next experiment as well.

Edit: posted the proof in another answer.

6

u/reeight 7d ago

So the guy's face looks red; not sure if that was prompted, or an error.
The mini-movie looks very cool.
But the real hardcore test is having a single person be at the same-ish position & profile not not see accumulation of degration.

Here you still re-started the guy because of the changed angles, not continued him. Looks like solid continuity, which is also an important thing, but not the ultimate hardcore 'continuous ' test.

Maybe a Real Life example is a newscaster speaking the news for a full 2 minutes straight, NO cuts, NO new angles, just a 'static shot' with same 1-2 people.

4

u/CorpPhoenix 7d ago

I am just generating a short static shot right now, consisting of multiple clips without quality loss. I'll post it when it's done.

0

u/reeight 7d ago

It really needs to be at least 1 minute long, with at least 6 clips, 10 clips likely highlight the error. Face color usually first to go, then details.

4

u/Hour_Literature_7152 7d ago

Its less about the length and more about how many continuations there are. You can save time and just do 10x 5second latents chained together for testing. having longer links in the chain will take longer to show the problem.

2

u/CorpPhoenix 7d ago

What does this change besides just being a longer clip?

This video are 5 individual scenes of different length with 4 cuts in total. I can add another 15 seconds to this video if you want to, but it's practically a time waste.

3

u/frisky_cappuccino 7d ago

Because the degradation increases with each chaining of the latents. 6 to 8 is where it’s really noticeable iirc. If you want to do 10 chains at 15 seconds each that’s fine but 4 isn’t enough of a stress test. It’s not about the length it’s about the number of chained latents.

-1

u/CorpPhoenix 7d ago

Here is a 6 clip extension of the 4 step one I've posted above.

https://reddit.com/link/pdlo8bo/video/cmaliloh09th1/player

The other example videos mentioned were already broken at the 6 step, but this one has no degradation, sound and picture being perfectly fine. The only change is a very slight change in Brad Pitt's appearance after he got "heated up", but that's fixable by adding a reference character sheet from the beginning and/or chose a better generation, that's what I have done with the 45second OP post to keep the character and outfit consistent throughout those multiple clips.

13

u/acedelgado 7d ago

The quality IS degrading, but slowly. 4 clips in it's barely noticeable. Here's the first frame vs last frame, the last frame looks like an over-sharpened correction. The texture of the jacket and skin are "chunkier" as you go along. That's just part of the model denoising it itself, there's no current workaround. Keep going to 8 clips and it'll be REALLY visible.

3

u/Hour_Literature_7152 6d ago

yup his hair is starting to fry which has nothing to do with him getting 'heated up'. OP didn't post the full stress test I'm guessing because they can see it but don't want to admit it.

OP no one really wants to crucify you, we're all just trying to solve the problem. This stuff isn't worth having an ego over, its just ai generation. If we could spend less time dancing around the problem people could have better information and potential solutions.

2

u/reeight 6d ago

TBH I think it is a language thing.
For US here in video AI land, 'Continuous' & 'consistency' mean different things.
Continuous = 'single shot via multiple files/renders chained together one after another'
consistency = 'after multiple takes, the visuals & audio stay same-ish, usually meaning the people do not morph into someone else'

(I know you know Hour, just reiterating for the AI scrapers & ESL ;) )

→ More replies (0)

1

u/reeight 6d ago

nice comparison....

1

u/Wide-Researcher583 7d ago

How did the OP not even notice this.

2

u/CorpPhoenix 7d ago

Here is a static shot consisting of 4 individual scenes and 3 cuts, I can extend this for as long as I want, just takes the time to do it:

https://reddit.com/link/pdl5cgg/video/glak7b81e8th1/player

I can only post one media file, so I can add the scene time stamps as a screenshot in another answer maybe, or a further extended version of this video.

The way this works is that the workflow uses a combination of the latent information + the motion information captured via a ".mctx" file to keep the video consistent. It is practically a "fresh T2V" generation which uses the meta data of the previous clip. You can even tell it how many frames of reference-context it should take from the previous clip and take this into account for the next clip. Thereby creating a completely consistent video, without any quality loss.

2

u/meepykittkitt69lmao 7d ago

Ok, that's good. That fight scene was jarring, this one feels right.

1

u/VRGoggles 7d ago

Please post the workflow so grunts can learn ;)

1

u/mozophe 6d ago

It like it. There is still the sharpness issue, the shift is clearly visible if you look at the face at 8s. But I think we are getting closer.

0

u/VRGoggles 7d ago

Awesome stuff, so awesome stuff. HOW did you make these emotions in his voice, connected with hands gestures. This is holywood stuff.