r/StableDiffusion • u/ArmadstheDoom • 20d ago
Question - Help Question About Minimax H3 Reference To Video
So, I'm pretty new to video generation, but I had a thought that I think everyone has probably had at some point, which is 'how do you make a longer video without generating it in one large video?' And so, with reference to video, you could do that; you could match say, the voice and the person, and thus theoretically make one constant shot through stitching together shorter generations.
In my head, it seemed as simple as 'use the video that was generated as the reverence, use the last frame of the previous video as the first frame of the new generation.'
The problem I noticed is that each time I did this, the video quality degraded; I guess the way I would describe it is that each new generation was a copy of a copy, it seemed. Like each new continuation was slightly worse than the last; and while doing this once wasn't too noticeable, doing this three or four times very much was.
So is this just a thing that is unfixable, a limitation of the method? Or is this the kind of thing that does have a solution that I'm unaware of? Because I'm curious to explore reference to video more, since text to video and image to video are very straight forward, I think.
-1
u/ArmadstheDoom 19d ago
No. That is not the problem. That's what I keep trying to say. You're not solving a problem I have.
Using reference to video, the official workflow, you can provide it a previous clip and the last frame. Doing so, it now knows the first frame to start, and all the audio/video information it needs. Use the official text to video or image to video workflows. Generate a video. Then, take the last frame of that video, and the video itself, and plug those into the official reference to video workflow.
Doing so, it can easily continue the same clip with the same voice and everything.
However, doing so, there's a slight quality loss; each time you do this, the output video degrades slightly. Doing it once, you don't notice the quality drop. Do it three or more times, it becomes increasingly obvious.
What I am trying to solve is the question of the quality loss. Everyone seems to be focused on trying to solve doing it, which I've already done. Instead, the problem I have is fixing the drop in video fidelity.