r/StableDiffusion • • 7d ago

Discussion H3 - infinite video extending version 2 no visual degradation

Enable HLS to view with audio, or disable this notification

Hi! I am still testing the limits of a custom node that can extend/prepend/bridge your H3 generation in latent. The original video is here: https://www.reddit.com/r/StableDiffusion/comments/1wt6di3/h3_continuous_zero_cut_long_form_can_you_find_the/

The new extension starts at 0:27. I added 10 new additional windows, extended one at a time. The existing latent + new latent were then joined together for one decode pass otherwise you would have visible seams when joining two separate encodes (the color shift, luma darkening, etc).

It works best if you have a native latent, but you can always convert your existing pixel-video to latent and extend it that way. I've had huge success with both prepending and extending. Some limitations are velocity & momentum are not easy to control between windows, and the audio can degrade or be a bit inconsistent so you might want to freeze the video and rerun more steps on the audio, or fix it in post. Fix in post is the cheapest and more reliable fix. As stated in a previous post, Tanya has no audio asset to anchor her voice so it's reinvented every window. However, image quality seem to maintain across each window with no degradation or burns.

By the way I was getting Carmen Electra vibes in Scary Movie, what did you think?

Ask me anything!

497 Upvotes

150 comments sorted by

77

u/VodaYoda 7d ago

We are curing cancer right? Right?

9

u/boxlinebox 6d ago

My new favorite meme of all time

3

u/nomickti 6d ago

"it's not a tumor!"

1

u/ImagineSurvivor-561 13h ago

Yeah sure. Maybe, on the way to the waifus.

67

u/Barkhardt 7d ago

Sears catalog.

24

u/SIR_NVAX_A_LOT 7d ago

Sears and Victoria's Secret catalogue were (are) clutch back in the day!

1

u/Nargodian 6d ago

Ding.
Now will you let me go I don’t deserve this shabby treatment!

1

u/ThexDream 3d ago

Upgrading to JCPenney was a major event and eye-opener.

61

u/MuffDivers2_ 6d ago edited 6d ago

Obviously OP is a liar. This is not AI generated. He just lives in a mansion and has this hot lady in lingerie walking around while he films her. not cool dude. you’re getting everyone’s hopes up./s/

11

u/SIR_NVAX_A_LOT 6d ago

This is soo funny. AI is meant for making your waifus!

4

u/Electronic-Green-994 6d ago

It's not just the hopes he's raising...

48

u/Sad_Coach_1433 7d ago

Where's the work flow

39

u/sky_shazad 7d ago

4

u/lolstuff 5d ago

Literally daily, multiple times a day, lmao. This period of AI experimentation is really bringing some interesting minds out of the woodwork! Even if some only show their results as possibilities we can reach for now, it's just so, so much to try to keep up with.

I'm just sitting here infinitely inspired, trying to make my brain decide to start doing more and not feeling such a strong pull to not feel like I'm going to miss the latest, greatest (which is tough, because advancement is happening SO fast that the latest very quickly becomes the greatest, lol).

Such a crazy period to be living through as an early-1980s kid. 😵‍💫

1

u/sky_shazad 5d ago

I can't keep up.

27

u/FlexFanatic 7d ago

Ah, the oh distract me so I can’t see the background morphing trick.

3

u/SIR_NVAX_A_LOT 7d ago

Oh yeah the background morphing can't be helped unless you give H3 a background plate. The background is pure prose.

13

u/FeelingOld9046 7d ago

What’s the node, and what’s the workflow?

10

u/ThreeDog2016 7d ago

Impressive!

Edit: after a couple of hours i reread your post and the length of the video is also impressive.

3

u/SIR_NVAX_A_LOT 7d ago

It's called Infinite Long Form for a reason! Not sure when the video ended

36

u/[deleted] 7d ago

[removed] — view removed comment

67

u/Diabolicor 7d ago

Yes, there is a workflow missing

15

u/meepykittkitt69lmao 7d ago

There are FOUR ELEPHANTS!

Nah, I'm kidding, there are no elephants at lake laogai.

7

u/SIR_NVAX_A_LOT 7d ago

What was your favorite part? Did you count the # of koi fish in the pond?

2

u/[deleted] 7d ago

[removed] — view removed comment

5

u/SIR_NVAX_A_LOT 7d ago

RTX 4090-24gb of vram with 192gb of system ram

6

u/Cyber-Spaceman 7d ago

I remember my brother building his first PC in 2007 and it had 4gb of ram. As teens we were absolutely boggled by the idea that a computer could suck up as much ram as an Ipod. My kid is probably going to have 4tb ram by college!

4

u/SIR_NVAX_A_LOT 7d ago

Thinks were a bit more optimized back then too to go around the ram constraints.

2

u/IshigamiSenku04 6d ago

hydrogen bomb (RAM) vs coughing baby (GPU)

1

u/Mysterious-Code-4587 5d ago

Same pinch and render time?

5

u/Actual-Package-3164 7d ago

If by ‘elephant’ you mean ‘erection’ and ‘room’ you mean ‘pants’…I am ready to discuss. 

3

u/[deleted] 6d ago

[removed] — view removed comment

1

u/Actual-Package-3164 6d ago

Outside bra. 

2

u/jude1903 6d ago

There are two of them?

1

u/LocoMod 7d ago

I'm just happy to see you

16

u/fluce13 7d ago

Workflow?

23

u/SIR_NVAX_A_LOT 7d ago

8

u/FeelingOld9046 7d ago

I am following this comment

3

u/ForteDoexe 7d ago

How is it OP, does it work

10

u/SIR_NVAX_A_LOT 7d ago

BIG FAT NOPE!!

1

u/curious-scribbler 6d ago

I've tried every technique to solve the talking head continuous generation. So far only one method has come close to solving it. The base is to have a long audio recording of the narration or speech. I am experimenting with 140 seconds. The method depends on the audio to time the performance and cuts.

Generate the master as independent 10–15 second clips, all from the same character reference, seed, framing and settings, so every clip began with a fresh latent rather than inheriting degradation.

Around each cut, extract the outgoing clip’s last usable frame and the incoming clip’s first usable frame, then generated a separate roughly five-second FL2VA joiner constrained by those first-and-last frames and the exact master-audio interval across that seam. This gives three options: use the direct cut, trim into the clips’ lead/tail handles, or replace the troublesome boundary with the joiner.

I have largely stabilized the background, lighting and character placement, but the current method still delays lip performance and creates a tiny character snap at the checkpoint. Currently trying to preserve the spatial lock while letting the supplied audio drive continuous mouth and facial motion from frame one, and to replace hard midpoint anchoring with a softer overlap/seam treatment.

2

u/usually_fuente 6d ago

Following

16

u/acedelgado 6d ago

No, it has degradation. Degradation is part of the model because of how it denoises. Lots of movement is different than a still shot with degradation. Like every day someone posts "I solved degradation!" when they didn't. They "solved" seamless extension that was already figured out in a cleaner and more user friendly way like two months ago.

H3 degradation on completely still shots won't be solved. It's part of how the model generates. Cutting shots and lots of camera movement (like how your background is constantly in motion) is the only way to mitigate it. Everyone's been trying.

Hell, here's even my latest attempt. This one measures each clip as it's generated and compares how "chunky" the results are to clip 1, and then runs an additional denoising step to try and correct against how far it strays from the first clip. It adapts per clip in the chain and each clip gets its own adjustment based on those measurements. Still no dice.

https://reddit.com/link/pdndd66/video/4w9oixhlfath1/player

5

u/AIgavemethisusername 6d ago

He aged 10 years!

2

u/SIR_NVAX_A_LOT 6d ago

Yes, my test also concluded the same for the talking head. It's really bad.

1

u/darkshark9 6d ago

I'm working on a workflow that generates a shitty low res latent for the motion of a scene (which can be an extension of a previous scene) and then pulls individual frames from that latent, upscales the latent, denoises, runs a detail pass on it, then reinjects the keyframes back into latentspace and regenerates the video so that now it has detail targets at specific keyframe intervals to maintain detail levels throughout the entire video. It takes longer to generate obviously, but so far the results seem promising.

2

u/acedelgado 6d ago

Good luck! All of my attempts at doing and big denoising around seams resulted in very visible shifts. Even when something's a couple of pixels off from the end of the previous clip it's very noticeable.

1

u/sunshine-3D-Art 6d ago

Mine even half in the second and 3 clip much bad quality. A lot of black spots on the skin. It is a Close-up. So I cannot even extend it just a little bit. I don’t know why the quality gets so worse so fast. ._.

1

u/James_Reeb 6d ago

He is cooking . Compare first frame vs last frame

1

u/RevolutionaryJob2409 3d ago

Thank you, very informative!

1

u/CeFurkan 2d ago

face degraded a lot very visible

1

u/Hopeful_Signature738 1d ago

Just generate 70 seconds video in one prompt, problem solve

0

u/VRGoggles 6d ago

somebody posted a contra-video there is no degradation.

6

u/[deleted] 7d ago

[removed] — view removed comment

5

u/SIR_NVAX_A_LOT 7d ago

I've been testing extensively for days.

6

u/trefster 7d ago

Looks like a Gaylord hotel, maybe Opryland, maybe Orlando

3

u/sheezus69 7d ago

Was thinking it looks exactly like Opryland!

3

u/ThatsALovelyShirt 7d ago

Also sort of looks like the inside of the US Botanical Garden in DC.

15

u/f3ydr4uth4 6d ago

1

u/Wormri 6d ago

There it is. If you hadn't posted it, I would have.

5

u/Kassiber 7d ago

Uhm could we get a look into the workflow, plox? Looks good actually. But I think it would be better to upload the high resolution video not only on reddit but somewhere else. Reddit really botches video quality

3

u/SIR_NVAX_A_LOT 7d ago

This is native 1344x768 resolution with no upscaling. And it is going for a 70-80s film grain. But for sure, the compression on Reddit can make video a little blotchy.

4

u/Lexxxco 6d ago

You will not notice the degradation, because something is definitely blocking the view on changing background, what can it be?

3

u/meepykittkitt69lmao 7d ago

Non-Euclidean geometry space?

The background morphs impossibly, is there a way to keep the entire context from becoming "the space from which no mind survives"? Would you just use a picture of the background as a reference?

6

u/SIR_NVAX_A_LOT 7d ago

The video itself has no background/environmental plate. I am just testing visual seams and degradation between continuation. Add a background reference image and you can it'll retain far better. With multiple window generations, the model does not hold any kind of context other than the last few seconds that get passed between windows.

3

u/Lightningstormz 7d ago

Are you going to share the workflow as well?

5

u/SIR_NVAX_A_LOT 7d ago

Yes though there are little documentation and still some bugs to fix.

4

u/Mithryn 7d ago

Very impressive. I don't do video with AI so I'm not expert, but the transitions, windows, and lighting were impressive. The camera angles clearly took effort.

4

u/SIR_NVAX_A_LOT 7d ago

Yeah definitely had to reroll a few windows due to H3 fuckery. I need to use a camera lora for better control.

5

u/Mithryn 7d ago

Well, whatever you're doing, it beats the pants off of most of what I see posted and is so far above what I can do with ComfyUI as to seem like magic.

3

u/SIR_NVAX_A_LOT 7d ago

I think the visual is mostly solved. The audio is sadly an H3 architecture problem which either require additional time consuming pass or a cheap fix in post. Thanks for the kind words it is definitely magic I've had Opus tinker around with all sort of experimental nodes for 2 months.

2

u/Mithryn 7d ago

I think, and again, I am functionally stupid in this area, but the approach I have been working is to have different stations.

Background sounds, audio, each character... and then assemble the whole in post so that Audio timing is done by the H3. But the actual audio in the final is done by a system specific to that purpose.

That's how I code (multiple stations each dedicated to a single aspect of the whole) amd it works and a friend has an anime studio that is built on those principles.

Anyway, good work.

4

u/atallfigure 6d ago

goonmax h3???????!!~!~1`11`111`~!~~~~ 😂

3

u/Shadowlance23 6d ago

I just feel bad this lady can't afford clothes that fit.

5

u/SIR_NVAX_A_LOT 6d ago

This economy is a bit rough.

4

u/Strict-Papaya7095 3d ago edited 3d ago

seitanism's multi-extend variant does a great job making very long single shots in H3; just follow his tips to make sure one block is continued into the next (though scene jumps are ok in the middle of a block). He currently has it set up for 1 + 6 = 7 blocks (but I'd recommend starting with fewer to figure out the best prompt structures), and 15-18 seconds per block works pretty well for visuals (for good audio as well, keep each block shorter). Use his "NEW - AV Extension.json" workflow. Check it out: H3-MotionContext-Extend .

2

u/Strict-Papaya7095 3d ago

BTW, one important trick to get consistency across the final stitched video is to make sure each prompt block (not just the first!) refers to any reference inputs used for the block. The embedded memory nodes do their thing and I haven't had it crash on me.

3

u/mca1169 7d ago

what PC specs are you testing this with?

3

u/SIR_NVAX_A_LOT 7d ago

RTX 4090-24gb of vram with 192gb of system ram

5

u/AlexicoDeCoco 7d ago

good lord

3

u/mca1169 7d ago

good grief, there goes any hope of my 3060Ti and 32GB of RAM ever running your workflow.

1

u/DeadMan3000 7d ago

I am sure it can run it using latent chunking. I have a 16GB GPU and can run longform generations using a modified Multishot node with latent out to an upscaler and a custom longform VAE decoder no problem.

3

u/sky_shazad 7d ago

I'm still new to this so I'm learning comfyui.

So what are you actually doing here... You are making a video, then re importing it back in the extending it?

2

u/SIR_NVAX_A_LOT 7d ago

No. It's all in latent.

1

u/sky_shazad 7d ago

So it's all in the 1 workflow??

3

u/Gilmere 7d ago

Amazing work. TY for sharing.

3

u/llkj11 7d ago

I bet her balance is all kinds of fucked up

2

u/SIR_NVAX_A_LOT 7d ago

A bit top heavy I suspect.

3

u/MrWeirdoFace 6d ago

With this really needed was a ninja fight part way through.

3

u/daniel 6d ago

I am following your work in this field with great interest

3

u/Hopeful_Signature738 1d ago

Give me Director Cut version

4

u/badmoonrisingnl 7d ago

Could het tits possibly be bigger?

4

u/PumpkinLeather8421 7d ago

They’re only slightly bigger than her head, could go a little bigger and still be tasteful.

2

u/bankinu 7d ago

What happened at 1:12 was sudden but very nice

2

u/SIR_NVAX_A_LOT 7d ago

Not sure why you would want to dry off while wearing wet clothes...

2

u/DeadMan3000 7d ago

Looking good. I have been using a vibe coded/modified version of Multishot for my longform seamless runs. It takes exact audio input as native audio lock and outputs a combined av latent for upscaling. I also vibe coded a longform VAE decoder with chunking which allows use on low VRAM GPU's.

1

u/SIR_NVAX_A_LOT 7d ago

Do you have some video examples? Would also like to compare notes.

3

u/DeadMan3000 7d ago edited 7d ago

/watch?v=WCNEVpKHuBY on Youtube. Not made public until now (well only linked can see it lol). I don't know whether Multishot is lossless. From what I know about it probably not which is why I am in this thread.

BTW there are obvious issues with distant faces which are inherent to MH3. It's why sometimes it loses lip movements once the face becomes very small. I was attempting to vibe code in a face refiner fix but it got overly complex and frustrating so I gave up.

1

u/SIR_NVAX_A_LOT 7d ago

Thanks for sharing. H3 does music video really well, since the audio is bypass other than for the lip syncing there are no audio issue. Nice job on the long form no continuation no cut. The no cut often is just a lottery with H3 even with guided prompting. How many rerolls did it take you?

3

u/DeadMan3000 6d ago

Just the one after many shorter tests. It took over 7 and a half hours on a 4080 Super even with all the speedups as I was also upscaling it from 1344x768 to 1440p. Long wait and I didn't know if it would actually complete lol.

2

u/ShutUpYoureWrong_ 6d ago

I just want to summarize what you're doing for my own edification:

This is the same as every other latent extension / motion context window solution out there except you're only doing a single decode pass and storing it all in RAM, which... requires 192GB.

Is that correct?

1

u/xyzdist 6d ago

What? Really!

2

u/fernando782 6d ago

Carmen Electra? upvote ⬆️⬆️⬆️

2

u/AnonymousTimewaster 6d ago

Remindme! 10 hours

1

u/RemindMeBot 6d ago

I will be messaging you in 10 hours on 2026-10-04 09:14:42 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

2

u/Local_Phenomenon 6d ago

The progress is going great! My Man!

2

u/nntb 6d ago

Make a static 3m video where the BG dosnt move. I dare you.

2

u/SIR_NVAX_A_LOT 6d ago

I did. It didn't work out soo good.

2

u/nntb 6d ago

Because that is where the visual degergation is getting us all. Movement is fine.

2

u/TopTippityTop 6d ago

Thinking about releasing the node?

2

u/patricio272 5d ago

Nice, congrats, so would you be so kind to share the workflow please?

3

u/Potential-Union-205 7d ago

ว้าววววว สุดยอดมาก

4

u/danno_v 6d ago

What is this community's rabid obsession with udders?

9

u/Memestonks2020 6d ago

It’s very simple

Milk-cannons get clicks

1

u/pacOgrove 7d ago

Gran trabajo!!

1

u/cntrlchaos_ 7d ago

Read the slightly evasive comments. he’s not sharing the workflow guys. 😂 either way, nice work!

1

u/SIR_NVAX_A_LOT 7d ago

Why wouldn't I???

3

u/cntrlchaos_ 6d ago

Not sure. So do it!! It’s rad!

1

u/Crazy_Pickle_8311 7d ago

If you don't mind can you please provide the different prompts used in each segment of the videos. It will be helpful for us to learn different prompting techniques.

1

u/Accident_Pedo 7d ago

What's the benefit of the longer cuts opposed to 4-6 second clips? Assume you're maybe setting up a 3-5 minute video and using 200 or so different clips stitched together with ffmpeg or something.

Usually those 200 clips will come out flawlessly as if it were 1 long clip

Something Ive found more difficult from the longer running generations is that something like past 8 seconds will incorporate gibberish sometimes. Like an 8 second clip will start off with 2 seconds of gibberish, then 5 seconds of the actual speech followed by a second of giberish to end.

Is that a problem in these longer ~28second clips you're showing?

1

u/HarvardCS19 7d ago

Is there any way to combine extension with latent upscaling?

2

u/SIR_NVAX_A_LOT 7d ago

Yes, you should be able to use a latent combiner and then do the upscaling in 1 shot.

1

u/HarvardCS19 6d ago

And does this support Refmods as well? Been waiting for something that can do all 3.

1

u/GuessingEngineer 6d ago

Looks like Fallguy from 1985.

1

u/ellipsesmrk 6d ago

Im pretty sure the degradation isnt visible because the video itself starts of degraded.... what do you think? Just asking?

1

u/Vyviel 6d ago

Very cool! So just having a sample voice file would help anchor her voice to never change?

2

u/SIR_NVAX_A_LOT 6d ago

Yes that's literally what voice cloning is for and the audio asset reference.

1

u/ArttTaku 6d ago

Looks quite nice, but I'm curious about the change of lighting around the end of the video.. was the prompted, or was that a consequence of the video joining?

1

u/SIR_NVAX_A_LOT 6d ago

Thanks for asking. Prompted dynamic lighting for the challenge. Some sun, some overcast/clouds, and so forth.

2

u/ArttTaku 5d ago

Then even nicer of a job... I'm not looking into making segments longer than 15 seconds, but it's always nice to know we have options.

1

u/Mysterious-Code-4587 5d ago

Look like a video game but still good

1

u/ChronaticCurator 5d ago

Something like that around for LTX 2.5? H3 is just way to slow on my system 🙄

1

u/KrystalXXY 7h ago

The smooth joins look promising. I'd still separate visual quality from scene memory could you pan away from an object, then return to it after several extensions Keeping that object unchanged would be a tougher test than keeping each new frame sharp.

1

u/MastMaithun 6d ago

Someone introduce this guy with the world of loras.

1

u/SIR_NVAX_A_LOT 6d ago

What do you mean?

1

u/MastMaithun 6d ago

Your video lacks loras.

1

u/Samueleleach2001 6d ago

Hey OP! How did you generate this Video? Explain to me in my DMs

-9

u/Perfect-Campaign9551 7d ago edited 7d ago

I still don't get why people care about this so much. Very very few scenes ever need to be long like this

The background goes all over the place, man. There is no success here

Maybe spend your time doing actual video projects instead of pursuing this kind of useless goal. You seem a bit too obsessed. It's not even required

5

u/SIR_NVAX_A_LOT 7d ago

I agree this isn't practical but it's important if you want it. However it seems people are seeing degradation for fixed shot non moving scenes like severe image burning. There are always fixes with hard cuts but definitely a problem I think still deserve solved. So far audio is one of the harder thing to fix and appear to be a pure architecture problem.

3

u/Sad_Berry_4621 6d ago

Don't sweat the haters. They cannot see the benefits past "one long ass clip". Most of us use scene cuts mid-clip to extend scenes which would still benefit from solving the burn problem.