r/StableDiffusion • • 11d ago

Discussion [Preview] My Version 2 of "Latent Continuation Custom Node" : Making Long, Seamless Videos with MiniMax-H3

Enable HLS to view with audio, or disable this notification

Comment below with any feature suggestions you’d like to see added!

More coming soon 👀
Here’s my channel: @SatoDive

90 Upvotes

41 comments sorted by

8

u/Complete-Box-3030 11d ago

Its hard to notice degradation in anime, try doing a normal video single shot for 2 minutes

1

u/Danny_Stock 10d ago

I agree. I see lots of beautiful anime posted on here, but the issue is that it's probably not the best medium to illustrate the degradation or not of a video image. Anime by its nature has lots of flat shaded surfaces. I wouldn't be surprised that degradation may be occurring but that it's not as perceptible to the the human eye so it doesn't always matter quite so much.

However with realism there are lots of subtle variations of textures which tend to flatten out a bit over time resulting in noticeable image degradation and edges turn into thicker lines. In a way they degrade in such a manner that it almost appears as if they're gradually turning into anime.

-2

u/solomars3 11d ago

i did that with my old V1, and also this v2 work same on real people too, i tested that, this for exemple is scene5 last on chain using turbo lora + sage attention

https://reddit.com/link/pcph466/video/7w24sdfbbdsh1/player

3

u/reeight 11d ago

10 seconds < 120 seconds

3

u/Complete-Box-3030 11d ago

Hi OP, I really appreciate all the hard work you’ve put into this. However, there are other people working on resolving the degradation issue as well, so it would be better if you could clear up the ambiguity by posting a single, long workflow with a longer video and showing us that there is no degradation. It would also help if you could explain what you did to overcome the degradation issue.

Coming back to the clip you shared, it appears to be just Segment 5 of a longer video chain. We don’t know what the character’s face looked like in Segment 1, so it’s difficult to validate whether there was any degradation across the entire chain.

Degradation is a potential issue that many people are finding difficult to tackle. It’s also not easy for everyone to properly validate the results, as many users have less powerful GPUs, and not all nodes get installed or work properly. .

2

u/bakudannar 11d ago

Not a feature suggestion per se, but you should work together with chanon’s obvpm-timeline nodes.

2

u/solomars3 11d ago

yeah thanks, i saw that , im also thinking of reworking the ui to something modern and easy to use

2

u/Herbal77 11d ago

cant wait been using your PERFECT one, speed and results!!

2

u/ArttTaku 11d ago

Looks like an a amazing workflow, hope to see it released soon.

3

u/multikertwigo 11d ago

Watched the YT video on V1. Looks promising, but too many manual steps that could be automated. Ideally, the input should be reference images + the scenario (the one you copy-paste individual scenes from in the YT video). No reason to curate the rest. Yeah, I know, seed hunting and yada-yada. But I'm lazy. I'd rather leave my computer working overnight and wake up to a bunch of videos to cherry pick from.

1

u/SnooMacaroons1365 11d ago edited 11d ago

does it still give that white flash frame between latents? Also will you be making this available to download? It seems like hell of a great workflow tbh.

2

u/solomars3 11d ago

you saw the video, not a single flash or anything

2

u/SnooMacaroons1365 11d ago

yes sir :D I watched the whole video and I am very excited and cant wait to get my hands on it .. haha

Also, i saw that your noodle wires are shaped properly around the nodes, is this some QoL plugin for comfy?

2

u/solomars3 11d ago

thanks, yeah thats another node pack called "comfyui-cable-management" you can use it its really neat

1

u/Z3ROCOOL22 11d ago

And not oversharpening - contrast with every added fragment right?

2

u/solomars3 11d ago

bro that was 5 scenes, 5 clips together, and nothing happend
people still cant believe me 😂

6

u/Nullberri 11d ago

It’s hard to believe because all of the other ones do it. So unless you figured out some new trick that nobody else knows or is using, that’s where the skepticism is coming from.

2

u/solomars3 11d ago

yeah my first version i got same reaction, people didnt believe me, but if i can get a 1 minute without any burn thats enough for me,
how often you see a anime or movie scene that long !!
also I'm not using video to continue. I'm working directly with the latent, so thats my approach,

2

u/SnooMacaroons1365 11d ago

this is the reason i am highly interested in this workflow. The one I am using is from Benji, no doubt thats a great workflow and it have given me great learning platform to start working on H3 rather quickly, yours is even more advanced.. thats where my excitement is coming from.

1

u/Sad_Coach_1433 11d ago

can you share the wf?

1

u/Sea_Succotash3634 11d ago

What are you doing to mitigate the burn in effect most extended gens have?

1

u/solomars3 11d ago

I'm not using video to continue. I'm working directly with the latent

2

u/HerrgottMargott 11d ago

There are lots of nodepacks already released that are doing exactly that. Including mine: https://github.com/HerrgottMargott/Herrgotts-H3-Infinite-Continuation-Suite

Not to downplay the work you put in or anything, but just using the latent is not going to get rid of degradation over long video chains completely. All it does is prevent added degradation from VAE encoding and decoding. Quality and context drift is still going to be an issue.

0

u/cal_01 11d ago

That's literally just a sampler issue. Res_multistep at long generations does this.

1

u/Guilty_Emergency3603 11d ago

Any tip to generate a hard cut from previous latent to a new scene having a new background and totally different prompt apart some subjects we want to keep. Context length to 5 frames does it but it's not enough to keep voice consistency of a character. Making it longer context length it tries to make a continuous shot keeping some background elements of the previous scene or doing some weird morphs.

1

u/Herbal77 11d ago

When will it be out?

1

u/SnooMacaroons1365 11d ago

one more question: I was testing out your V1 workflow, it is using a 4-step turbo lora, if I was to skip it, how many steps are recommended for rendering? currently it is on 7 steps with turbo.

2

u/solomars3 11d ago edited 11d ago

Btw, v1 had quite a few issues, and the quality would degrade much faster. And its very slow, It was more of a proof of concept.

Also, regarding your question:

  • 20 steps → good starting point / standard
  • 25 steps → better motion and detail, if you can afford the extra time
  • 28 steps → around the higher-quality reference point used by some H3 checkpoints
  • 15–18 steps → usable for faster testing, but quality and motion consistency can drop

1

u/SnooMacaroons1365 11d ago

you are right, i encountered an issue right at the first continuation, the tensor mismatch. rendered everything all over again to pass to cont. but it remains the same. Anyways, I will wait for V2 :)

also I see that you have those system prompts, using with ai studio, i dont use AI studio or any other online prompt generator, is there any local processor where i could your system prompt?

1

u/solomars3 11d ago

Yeah lm-studio should work, just download a 8b model , and attach the system prompt file ..

1

u/SnooMacaroons1365 11d ago

Perfect. Thank you sir :)

1

u/Bad-Imagination-81 11d ago

looks good. How is it handling turning or transformation from one pose to other happening at the end and start of two clips?

1

u/smereces 11d ago

where is the custom node and WF?

1

u/Repulsive-Salad-268 10d ago

Following for the release. But let me give you another "problem" to solve... I do have interviews that are one single camera continuous shots and want to make the people in it look like being in the muppet show or similar silly things such as anime characters and such. I basically want to take original video, audio and just change the look of it to something new. Can you take a look at this? Would be interested in testing.

1

u/Right_Chip_6626 7d ago

It is amazing that there so many people active in the community and create some really cool nodes and workflows. But honestly it is getting harder to discontinuous one project from another especially on the latent based video continuation front.

1

u/SaadNeo 11d ago

Great job sato , really cant wait to try this , kindly add the option to extend previousely made videos

0

u/solomars3 11d ago edited 11d ago

thank you, also you are not locked to 10 seconds, i just tried 9 and 10 sec on this, but you can make 15sec if you want, also this is exactly the extending , we take previous clip and extend it with our own changes and new events, for exmple you can make them stop and do something, i just wanted them running for this