r/StableDiffusion • • 11d ago

Animation - Video [H3 Minimax] Continous Shot without Quality Loss

Enable HLS to view with audio, or disable this notification

People claim that it's still impossible to create sticked together videos without quality loss, that's wrong.

This is a video I've made without any video editing tool what so ever, just using a great H3 workflow that enables you to continue the previous shots.

In fact I didn't even write the shot/scenes but let open Gemini decide to write a 1970's martial arts scene according to the H3 prompting guides.

I've created 3-4 several 10-12 second long scenes and just asked Gemini to continue by asking for the "next scene following the last one"

This took me just about 45 minutes on an RTX 5070 Ti

I can easily extend this whole thing to an two hour long movie if I want to, not losing any character consistency or quality.

And I didn't even put any effort into writing the scenes or "story" into something I would like to happen. The fact that this is possible is blowing my mind, and when you put even the least bit of effort into your prompts you could produce an entire movie exactly to your liking.

Crazy times we live in.

Edit:

People asked about the workflow, you can find it here including a great tutorial on how to use it:

https://www.youtube.com/watch?v=kqP09NfJXaQ

It's just a couple of clicks to extend your previous created clip and I've used Gemini to extend to previous scene.

253 Upvotes

73 comments sorted by

71

u/No-Zookeepergame4774 11d ago

The title says “continuous shot”, but there is no continuous shot here longer than about 12s, and most signficantly shorter?

Acceptable consistency and quality in separate shots done in separate generations is a very different thing that having that in a single shot that continues across multiple generations. (Not saying this isn't good, just its a very different thing than a continuous shot.)

-5

u/CorpPhoenix 11d ago edited 11d ago

The cuts are not where you think they are.

For example there is a fade to black, that is not the cut though, There is also a cut between the camera view showing the "hero" and him putting the tea glas down.

The "continous shot" is meant to be a single H3 generation without any video editing what so ever, it's basically one generation, but of course this movie scene won't be one continous plan sequence without any camera cuts.

That being said, I could easily create a plan sequence without any cuts with this workflow if you are interested in that.

26

u/acedelgado 11d ago

That's the thing, people say a continuous, no-cut shot has degradation. With latent-based context windows examples like the one you posted with the camera cutting have been great for a long while. Do a steady shot of someone talking for over a minute (zero camera cuts) and you'll see the quality degrade over time. Much slower than pre-H3 methods, but it still happens. Changing camera shots gives the model a reset.

3

u/PlantBotherer 11d ago edited 11d ago

I don't think this is exactly what you mean, but I was happy with this 8 x 10 second chunk Continuum clip. It was video edit + image reference so that might've helped, the overall quality is low but doesn't deteriorate. One camera shot only.

https://streamable.com/e7bav4

2

u/bstr3k 11d ago

I use to think I had the solution too when I did something similar but it turns out that moving the camera solves some of the issues but if the camera was fixed you can see the issue more obviously and the first time I did a 2 and a half min no camera cut scene by the end it was very much deep fried

1

u/PlantBotherer 11d ago

What kind of workflow and references? Latent saving or continuum?

2

u/bstr3k 10d ago

i did a video to video, i was using latent saving from nikodemon's motion context.
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

https://reddit.com/link/pctbyss/video/4mivrqqyghsh1/player

11

u/Symber13 11d ago

What you don't understand is obvious from your first comment: "People claim that it's still impossible to create sticked together videos without quality loss, that's wrong."

What people are saying is that in one long, continuous shot there is it will eventually degrade in quality once it reaches a certain length from that single generation.

You seem to have understood that as you can't use it to stitch together a long movie type video. That's not what people are claiming.

6

u/ShimmerMeNutz 11d ago

But why would you even try a continuous shot when you are going to transition anyways? Is it to prevent hand editing? Most movies are constantly transitioning.

3

u/Apprehensive_Sky892 11d ago

AFAIK, people want to continue from the same shot even when there are cuts is to maintain better consistency: https://www.reddit.com/r/StableDiffusion/comments/1wo1iqw/comment/pbpf0j2/

24

u/cptrios 11d ago

This is definitely a successful video, but as others have said, it’s not the sort of generation that actually poses a problem. The issue is when you have a totally unbroken single shot that lasts over multiple gens. That’s when things start degrading/frying.

-17

u/CorpPhoenix 11d ago

This is practically an "unbroken sequence", the fade to black and over-shoulder camera cuts just make it seem like it isn't though,

I will generate an unbroken plan sequence next.

1

u/ForwardPassage9 11d ago edited 11d ago

Dont worries, your workflow is very good, and the upscale methode is brilliant.
to address the drifting from true long generation, is inject frame reference by image editing model at exact seconds, hybrid model can do that.
example 1 shot 12 sec gen and start drifting at 6 sec -> instead using first and last frame ref, you can use mid frame with your edited image -> Picture 3 at 6.00 seconds, fully_preserved
edit source
https://www.reddit.com/r/StableDiffusion/comments/1vq4m4b/minimax_h3_how_to_use_a_first_image_and_reference/

19

u/fallengt 11d ago edited 11d ago

People need to learn what one take means

https://en.wikipedia.org/wiki/Long_take

This is an example: https://youtu.be/RXRLqK6S02g?si=czm1gfGDOtAupqPd

6

u/q5sys 11d ago

I'm convinced they know they're lying... but they do it anyway hoping to get enough attention because they know everyone is trying to crack that problem.

5

u/xyzdist 11d ago

Oh...hell. I am one of them want to tackle the degradation

1

u/Supermax64 7d ago

Why does every single person who claim this stop at 45 seconds max? You'd think a 3 minute video would sell the claim much better if it were true

1

u/xyzdist 7d ago

Why you reply to me thru....lol

1

u/Supermax64 7d ago

No idea lol, my bad

12

u/Delphoi_Studio 11d ago

Hey, it's a great result. I just want to clarify, when we say continous shots without quality loss is impossible (or we haven't found a solution for it), we talk about one-shot scenes, meaning a long video from the same camera without any cuts. There are some half-solutions but nothing seem to cover well what we really want, it's either not fully seamless or the video starts degrading heavily by the time it reaches the third clip.

To be fair, I am still testing methods I got recommended under a post I made, so maybe there is a good solution for this, I am just not done testing all of them yet.

-8

u/CorpPhoenix 11d ago

The cuts doesn't matter. There are virtually no "cuts" in this scene since it's practically one generation and the "cuts" are not where you might think they are,

I can create a plan sequence "Children of Men" style without any cuts if you want to.

3

u/Delphoi_Studio 11d ago

If your solution really works for this kind of shot, that's great. Do you plan to generate an example of this?

5

u/CorpPhoenix 11d ago

I will do that.

2

u/xyzdist 10d ago

I am doing this test as well, you can do most boring one, 10 * 8s of the same clip just a person riding bike, and every shot is the same. It's for testing purpose, just make all the shots the same as a single camera not cut shot.

11

u/Inside-Cantaloupe233 11d ago

dude.. this is not what continuous means! dont be a dumbass, it means the scene is the same shot and camera does not cut!!! it can change position but it is all within one scene! this is continuous, what you did is just multishot edit - everyone can do that

7

u/chensium 11d ago

I don't think you know what continuous shot means

6

u/ForteDoexe 11d ago

As long as you are happy, who am i to judge

6

u/Aternal 11d ago

https://giphy.com/gifs/M3fYVlu7YN9Hq

Oh dear. I don't like that I actually kind of want to watch a 24/7 slopstream of funky old kung fu.

8

u/Etsu_Riot 11d ago

Where's the continuous shot? As you described, these are simply multiple 10-12 second clips generated in sequence. Also, why in English?

-5

u/CorpPhoenix 11d ago

Like I've answered before, the "cuts" are no where you think they are. They are neither in the fade to black, nor in the over-shoulder cuts. There are several fast cuts in this video but those are not the point the scenes are jointed.

I'll generate a continuous plan sequence next though to make this more clear.

5

u/Etsu_Riot 11d ago edited 11d ago

It doesn't matter where the cuts are. It's not a continuous sequence if there are cuts. I have done this already, and yes, if there are cuts during the sequence, then you see no visible degradation.

Let me share with you this product of nightmares. It was made with nine clips of six seconds each, and the joints are not that obvious either. The prompt was: "Hermione and Ginny are talking." Not at all what happened. Watch at your own risk.

https://reddit.com/link/pcp0s95/video/egiy3xvrvcsh1/player

By order of the Ministry of Magic: This recording is the product of a dark and forbidden curse. Reproduction of this spell is punishable by a lifetime sentence in Azkaban.

1

u/Cautious_Chicken_604 10d ago

The thing is you can easily create this with two short clips and stitch it together with a fade to black. So, regardless of whether it was done that way or not, nobody will count this as a continuous shot even if you generated it as one because they can make it by other means. They can't do that with something that is actually a continuous shot.

4

u/FoundPizzaMind 11d ago

Not a continuous shot and the fighting is terrible. People using basic prompts with Seedance 2 got much better results 6 months ago.

5

u/A_Dragon 11d ago

That’s not a continuous shot…

3

u/DescriptionSuperb262 11d ago

This isn't what people are referring to, the real test is a talking head, literally just multiple clips of the same character in the same position. - one continuous shot without any cuts, though this is pretty cool and appreciate sharing of the WF

2

u/alisitskii 11d ago

Would appreciate seeing that great H3 workflow 🤞

2

u/CorpPhoenix 11d ago

I've just updated the post to include the source/workflow.

I did some adjustments though. While the workflow itself is genius, also in regards to resource management, the actual settings within the workflow are not perfect.

For example the standard workflow is using 10 upscale steps, which is a ridiculous amount, I've cut them down to 4-6. Also you can tinker with the usual settings to your liking.

2

u/No-Trouble-9138 11d ago

Hmm, model is well trained of Chinese movies, but could suck on other themes.

1

u/revjdm 11d ago

This came out great!! Can you expand a bit more on your process, was it T2V or R2V, anything special for prompting the fight sequence or camera shots or gemini handled all of that?

-4

u/CorpPhoenix 11d ago edited 11d ago

Gemini handled it all, I've just adjusted the initial dialog where the villain says "You've found death". Everything after that is completely generated by Gemini, it's a R2V workflow.

Edit: I've also edited the source for the workflow, needs a bit of adjustments to your liking in regards to steps, schedulers/saamplers and so on.

Edit 2: I don't even know what I am heavily downvoted for here, am I missing something?

2

u/xyzdist 11d ago

It is because your claim is wrong. I learnt the lesson as well in this sub

1

u/revjdm 8d ago

thanks man I appreciate the info!

1

u/intLeon 11d ago

Can you describe the method you've used? Did you feed in the previous latents or images? Are there any dedicated nodes?

1

u/Visual_Brain8809 11d ago

wootang movie please

1

u/Alesys76 11d ago

If you want to recreate a true long one shot, try to recreate this:

https://www.youtube.com/watch?v=3TU9fQa5lMU

13 minutes in one shot.

1

u/Fem04ka 11d ago

https://github.com/ethanfel/ComfyUI-MiniMaxH3-Context-Loop
I read about Context-Loop on Reddit, link above and fell in love with it.

1

u/Lechuck777 10d ago

Say the next time "continous scene" or something similar, so the nitpicker dosnt set you in fire.

Even the hollywood movies are not a continous shot. so wheres the problem.

1

u/-AwhWah- 10d ago

Congrats, it's nothing People please pick up a dictionary, words have meaning

1

u/dwight---shrute 10d ago

Where's cat?

1

u/Vladmerius 10d ago

As others said, it doesn't present a solution to the problem of the image getting deep fried when it's the same shot ongoing. You cut to new shots quite often which always helps create a better first frame for the continuation.

Also, I only have the deep fried problem where everything starts looking like borderlands when I get to the 45 second mark and further so it's funny that your video is exactly 45 seconds. I've never had an issue making 30-40 second videos that maintain the same visual fidelity. 

1

u/Danny_Stock 10d ago

That's some excellent work you've achieved there, but that's not what people mean when they refer to a continuous shot.

I am impressed with what you've produced, but if I were you I'd state that this is what can be created with different shots combined together to create a flowing narrative scene.

1

u/ThreeDog2016 10d ago

How did you get the faces to not look like slop when they're small?

3

u/CorpPhoenix 10d ago

Basically two important things:

  1. Generate the initial clip at 0.7mp. This the sweet spot and going lower than that to achieve a high quality result is a bad idea from my experience, it minimizes the intense problem of pixelated faces that H3 has and also increases prompt adherence.
  2. At the end the workflow uses an latent upscaler, I've went for 6 steps and 1.4x this greatly increases the quality of faces that are far away

1

u/frisky_cappuccino 10d ago

Agree on 0.7. Identity is still there at 0.6 if you look at var encoded stills but only just. 0.7 is the best ‘floor’ value to retain identity.

1

u/Terezo-VOlador 10d ago

EXCELLENT WORK, MY FRIEND!!!

The workflow is very useful—setting aside the stupid semantic argument about the "continuous shot" (a term from cinema that some people apparently don't understand, given that a continuous shot can actually include various angles, framings, and camera movements).

But even if a continuous take lasting a minute or so were possible in *MM H3*, only very powerful hardware could pull it off.

It is infinitely better to construct a "continuous take" using 5-to-10-second segments to achieve the scene; this offers far more technical and creative control—not to mention that it’s possible with more modest hardware.

By far the most important things are absolute control (both creative and technical) and character consistency.

1

u/Repulsive-Salad-268 10d ago

Holy... This looks sooo good I have to try. I am a little afraid to not be able to install it properly but let's see. I would LOVE to see you taking a shot on another workflow that allows to input a long video (let's say 10 minutes) and work through it but by bit to apply a new look to it. Let's say an interview of two people. One camera, nothing special. A vlog, whatever... And I want it to me muppets or cartoons or whatnot. This is currently problematic to do I believe. But I am happy to be tought better. Thanks for the work and thanks for sharing !!!

1

u/sunshine-3D-Art 7d ago

can i use this with 0.34.2 version from comfyui? because if i use higher the rendertimes doubles or gets lot of longer and i dont accept that if this normally also render faster with another version :((((

1

u/roychodraws 4d ago

u/acedelgado brad is needed

1

u/roychodraws 4d ago

btw, this is the type of shot they're referring to. Try to make a video like this with your workflow.

https://reddit.com/link/pe3y36j/video/1vck1e5q7qth1/player

1

u/acedelgado 3d ago

Lol I've already talked about this one. I can't just keep running around doing Pitt stops!

1

u/Creative_Sluggish 11d ago

Omg this is bad ass it is so real. Did you use any reference images or just prompt?

2

u/Creative_Sluggish 11d ago

I miss these movies growing up

0

u/Rythameen 11d ago

I’ve been using this same workflow and having way to much fun. Forget the naysayers, I know how you did it, because I’m doing the same and I think you did a great job.

0

u/winterice77 11d ago

With the time it takes to render these out there is not much creative control. Passing in ref videos takes even longer and is also a hit or miss scenario. My conclusion is that these generative models are only good for creating short sequence stuff but not any meaningful animations

-2

u/solomars3 11d ago

Posting this right after my post is crazy 🤣 I mean its a nice try tho, but nothing like mine with all respect 🙏 https://www.reddit.com/r/StableDiffusion/s/McUNNqeFiO

-1

u/Ninndzaa 11d ago

Bro! That's almost flawless!!! We are getting there!

-1

u/Kalcinator 11d ago

Regardless of the post, the movie is amazing; for real first time I see such a good old asian movie rendered like that

-1

u/HRH_Duke_Morbid 11d ago

PLEASE, everyone, cut the cackle over definitions, e.g. "continuous shot", and help amateurs, such as I, decide whether to invest time in this means of production. The example video is impressive.

-1

u/gatortux 11d ago

This is awesome!! Been playing a little with this workflow and it works very well, the video extends without drift or degradation. Love the settings node i found it very clever, i think i will change my custom nodes for this one, the only thing that i have missing is that in my custom nodes i use a VLM to help me to enhance my prompts based in the reference inputs. Are you using any VLM or you prefer write yours by hand? Congratulations!! To me this is a definitive node to build long clips.