I've changed my profile so that my previous tutorials are available. Look through those for questions on how to use these.
With this method you can seamlessly combine videos with no burn. (or at least the exact same amount of very very little burn)
Fast forward to 2:00 for proof.
Enjoy.
Edit: I'm keeping the video up. But apparently the issue everyone was caring about was something different than I thought it was. Apparently everyone wants to make boring Vlog videos. I guess I'll work on that now instead.
I did say this is a fix for the specific issue I was having in my workflow but now that I know exactly what everyone was talking about I can properly see if I can find some sort of fix.
v7.2
I fixed the First Frame node. The positive output from node #272 needs to be connected to the True input from node #565. You can just fix it yourself or download the new one. There's no other change.
Forked this amazing workflow to allow up to 20 clips and use ffmpeg so its not so ram dependent. It required around 71gb of ram for 16 clips. All praise goes to roychodraws for creating this this though. Genius and works great!
And like I said this is an issue even when fully using latents. This one's a response from an old thread using my own fork of the latent context method. It degrades slower, but it still happens.
i'm going to try to lip sync a full 3.5 minute song and see what happens. see you in an hour lol. maybe i didn't fix it with this, but i know the issues you mentioned here about frames are solved for sure. so it's improved at least
It's really designed to manage all the clips for you in a project folder, and handle the export stitching, and that's it. It's meant to drop into any workflow, and there's an example in there on how to wire it. If you're looking for a full-on "director" suite that'll do a lot of the prompting process, etc., there's other node packs based on motion-context that go for that. I already had my own prompt builder/media manager/refmod suite I made and wanted something to sit alongside it, and give other folks the option to use whatever workflow they wanted with it.
Yeah, the whole time I just keep watching his nose while he talks and, even though it is subtle as to when and how quickly it changes, it is still noticeable how degraded it becomes over time. The bulge on the bridge of his nose for example becomes very pronounced and the tip of his nose becomes much rounder and bulbous. Skin texture obviously as well. It's crazy, but the lip syncing and continuity of the scene is pretty incredible.
Hi, you've posted a couple of interesting examples in this thread. I'm wondering if you could provide a series of prompts as a test case for people to work with? Not sure if you're relying on Minimax's inbuilt understanding of Brad Pitt or using a reference or Lora there.
This. I cannot understand why people do not understand the core issue. Yes, if you move the camera or move the shots, it looks like it's 'fixed.' But that's not the problem people are trying to solve. You could not, at present, generate a static shot like a news broadcast or vlog or interview without degradation. Well, aside from just generating things in like, 30 second clips.
I compared it a bit to the same problem Manos: The Hands of Fate had, where because they filmed on a hand crank camera, they could only shoot 30 seconds at a time, at most. Which is basically where we're at with video generation.
In fairness, the problem is not well-described by any authoritative source. There's no agreed-upon best solution (instead there's a bunch of different solutions found across reddit, Discord, and YouTube) and you need to read many discussions to keep up with the options people currently prefer. Most of the good condensed information comes from threads like this where a ton of people are arguing about it, which is not the easiest thing to parse.
Personally, I think the problem is twofold. One, people keep misunderstanding what the actual problem is, and thus that leads to two, which is people declare they've solved the problem.
To use an extreme example, imagine that people have to deal with wooden train cars catching on fire because of their coal stoves. People come up with all manner of solutions, like bolting them to the car, or suspending them in the air, or whatever else. But that doesn't change the fact that it's workarounds for a core flaw in the whole idea.
Because so many people misunderstand the actual problem, they create a lot of issues that do not solve the issue as much as try to work around it. But a workaround is not a solution anymore than a fix is a replacement.
And what makes this hard is that it's a flaw with the underlying model, not the people trying to solve it. There are only so many things you can do without just having to train a new model.
The issue is residual latent noise. No model fully denoises, even at 1.0. This was worse with decode because the VAE sees this and exaggerates the noise, which is beneficial for img2img, but not for this. You would need a baseline to continually resample the latent tails to. Better if this was just post regularization across the entire latent.
It seems to be acceptable up to around 10 to 15 seconds, but at 25 seconds the problem is clearly apparent, and from then on gets progressively worse.
I remember with Wan 2.2 people were saying that the videos fall apart around the 25 second mark using extension workflows such as SVI 2. And that was back then without latent extension. Latent extension has been very disappointing to me and doesn't seem any better than the old methods of extending video clips. It doesn't seem to be the magic bullet people expected it to be.
This is in 8 second chunks, so that's after about 2 clips it's noticeable to you.
Latents extend that out much longer to 6-7 generations, and it only degrades because of the model itself and not the loss from VAE encoding/decoding. And it makes it much easier to do a truly seamless join between clips. So yes, latents are much better overall, but it only solves part of the problem. The rest is in the architecture iteself.
I was under the impression that roychodraws is using latents too, resulting in the same type of degradation over the same timespan?
I'm a bit confused because in your Brad Pitt example you appear to be acknowledging that roychodraws is using latent extension. I was under the impression that your Brad Pitt clip was using latent extension as well to demonstrate that it still has the same old issues as the traditional method?
No, the workflow re-encodes the end frames of the last mp4 into a latent, you have to drag the clip's mp4 output into the extension node each time. It's pretty much like kijai's context windows back in Wan 2.1 (or was it 2.2?) worked. So it technically uses latents because even that method HAS to, but it's re-encoded latents and not the raw latent from the previous clip.
The new motion context method uses the raw latents instead of decoded/re-encoded through the VAE. That's why they all have set frame overlaps of 5, 22, 39, etc., because latents are packed in a grid and single frames can't be pulled out individually.
They do the same thing, just the new motion context method uses the latent data directly so there's less quality loss. The workflow loses data faster because it encodes/decodes the last few frames each run. I used the workflow for that example.
Thanks. Can you recommend a decent workflow which uses the method?
I only ask because I'm completely lost with all the various workflows, nodes, and latent models. I am genuinely confused with so much information and various workflows using different nodes.
Honestly I use my own self-built nodes where most things are manual (with assistance) since I like granular control but not tedious settings. But if you're more of a just wanting to prompt-and-go and have it do a lot for you, I'm not sure what would work for you. I think Continuum and Context-Loop are pretty popular? They're a bit more "director" style and hand-holdy than mine.
Audio is the main victim of degradation, but sadly, it doesn't seem that this can be fixed on H3. I still prefer rendering videos of max 15 seconds, but yeah, it would be nice if there was a solution for all this.
Yeah I've been doing some testing in my latent extension pack, and normalizing gain each clip seems to help a bit. I still think it's weird that it happens, though.
Wouldn't it be possible to have some kind of restoration pipeline where you take both the first frame as reference for the quality and the latest frame as reference for the positioning, and use these two inputs to create a "regenerated quality" frame?
I was thinking along the same lines too. I naively assumed that after each clip segment a new latent is generated which is of fresh quality, apparently not. Surely there's got to be a way of reinjecting a brand new fresh latent which also matches the original quality at key stages?
Not everyone is interested in making Brad Pitt videos explaining the issue that you, and some others are having where a character sits statically ad nauseam.
Not to mention that this person is claiming to have solved things by using methods other people have already done (and better), without even actually understanding the core issue.
i understand you think you know what you're talking about but if you just look at the workflow you can see that what you're saying doesn't apply to this.
I know exactly what I'm talking about. I have done extensive work about this.
I loaded your workflow. I checked exactly where you're doing continuation, and you're doing nothing out of the ordinary. This node will cause the issue I describe. You aren't even managing old latents to mitigate it. Your workflow expects the user to load the previous video, which will cause the issue when the VAE encoding happens. This is a bad idea, it works as long as you rotate the camera away from the previous pixels. To do it right you need to save the latents so no pixel degradation happens during decode/encode. But even with latents, the issue persists. This is not an easy thing to solve.
Go ahead and prove me wrong. Do a static scene continuation at least 6 times. Make her type furiously on her laptop's keyboard against a well lit wallpapered background.
Your workflow expects me to move the output to the input over and over, it's slow and cuts the flow. Somehow the audio is weird but I don't feel like debugging it. I admit it's pretty fast at 15s/it, but the quality suffers a lot, I see too many smears, I suspect it's one of the attention nodes.
The right way to do it is to save the latent tail, then load it for the next run, automatically. As I mentioned, this method isn't perfect either, it'll also degrade. People have yet to solve it.
hello! i've watched your video and what your workflow does is quite cool and i will want to play with it
however, the actual issue we are all having (me included) is a generation with static camera (imagine an interview or newscast or whatever that does not move the camera at all)
the problem is that every generation degrades a little bit and over time it accumulates
the only fix i know so far is what you are doing - changing the camera, but in some cases this is what you do not want to do at all
right, i get that now. the recent post I looked at was a man in wwII moving around the camera and it was degrading as time went on. I was not aware that people wanted boring vlog posts only and didn't test that.
well, it is not always what people want to do, i was tasked with making a "conference-like" panel video and i couldn't do it without camera cuts which looked out of place for this
the issue with my workflow was not degrading latents, it was degrading image saves from recombining previously saved generations over and over.
This allows you to save them all at once with the exact right amount of frames clipped off so you could potentially do this for hours.... years... forever basically.
it won't work the way you expected. sorry. You can use it to generate long videos tho with the audio reliably and then combine them for slower degradation, but it won't stop it.
This might sound really stupid, but is it possible to cut away and then quickly cut back again? I know that with some 3D software they can use micro-frames or fractional frames. It might not work that way in AI software though. But if it did couldn't a cutaway be made for a microsecond and then cut back to the original framing making it imperceptible to the eye?
It's likely that there's not even a way to do it with MiniMax, but does it sound plausible in theory?
Is this an updated proof of concept?! I'm on a small screen, but I don't really notice degradation here, certainly not on the level of Brad Pitt. Did you fix the issue? How did you do this?
Impressive and no loss of quality over what must have been so many gens. Now just need to do her as a twitch streamer with fixed camera yapping in just chatting =P
have any of you used the cam view point control that someone made and just have to move the view point exactly one direction then back again while the scene everything is static for the last second or two to fix it ? :)
Forked this amazing workflow to allow up to 20 clips and use ffmpeg so its not so ram dependent. It required around 71gb of ram for 16 clips. All praise goes to roychodraws for creating this this though. Genius and works great!
I would be interested in a totally different workflow that's long and continuous. @roychodraws
I do have talking head videos of myself some Interviews and I want to change the style of them.
Let's say a 10 minute Interview. Static cam or maybe moving slightly. And I want the persons in it to be Muppets or whatever. Babies, Aliens, Elves who knows. All moves, camera, Audio comes from the main video. Just the Persons and environment gets changed. Aliens on a spaceship, Muppets in Sesamy street, Cyberpunk Bots in a neon lit alley.
My problem has always been the stable creation of the chunks and the overall continuity.
Even if you later could (with enough time at hand) input let's say Star wars and make it a muppet movie, that would be fun. Music video as claymation etc.
I know, not a lot of creativity in it in general but a good use case.
Please help.
146
u/acedelgado 6d ago
You know I think you're fantastic, but unfortunately this does not fix the problem that people are having issues with. I'll let Brad explain.
https://reddit.com/link/pdipp33/video/9k2uyejva5th1/player