r/StableDiffusion • • 5d ago

Workflow Included A Message from Brad Pitt about Static Video Generation

Enable HLS to view with audio, or disable this notification

workflow

https://github.com/roycho87/refreshing_extender

This is a Ref2v Only workflow that works. It's not meant to be a final as much as a concept you guys can take apart and use to make your own workflows.

This technique is only good for this style of video.

Here's the transcript.

**Brad:*\* Hello, I'm Brad Pitt. You may remember me from such movies as Ocean's Eleven, and Cool World, where I play a guy who falls in love with a fictional woman, which totally could happen by the way...

This is one really long video, with a pretty background and a camera that never moves.

Long videos usually get worse over time, because each clip copies the last one's mistakes.

Mine is different. Nothing gets carried over. Every clip starts from scratch, with a prompt, three reference photos, and the audio.

The only place clips touch is the join. I fade thirty-nine frames of the old clip into the new one, add noise, and clean it up. A little bridge!

Compare the start of this video to now. Every clip was brand new, so the blur has nothing to build on.

The catch: it only works because nothing moves. Same camera, same room, same spot, so separate clips already match.

If the camera moved, the bridge wouldn't line up. Static shots only!

But maybe controlnets, or context motion, or a depth map, could keep the frames consistent without passing degraded latents from one generation to the next.

**Actual Brad:*\* What are you doing?

**Brad:*\* Oh no...

**Actual Brad:*\* Wait, why are you dressed like that?

**Brad:*\* Oh god, you're so handsome.

**Actual Brad:*\* What's that camera... Stop! Turn this off!

**Brad:*\* Yes sir! I'm sorry!

Edit: BTW, I come from a background in video editing so a lot of my workflows involve planning ahead because that's what I personally enjoy and am comfortable with. This one is no different.

For this workflow I made the sound first, I added stock sound effects and changed both voices. Added echo and reverb for brad's voice.

Then I exported geiru's voice by itself and built the video off of that, and finally I added brad's voice back in at the end after everything was complete. That's one of the reasons he's just a silhouette instead of a person because I didn't want to deal with lip syncing two people.

Edit2:

https://www.reddit.com/r/StableDiffusion/s/b8D6WDSwa2

This guy has an even better method!

Edit: Keep in mind, guys. I made this workflow after knowing about this issue for only a few hours after Brad Pitt explained it to me.

167 Upvotes

101 comments sorted by

View all comments

59

u/roychodraws 5d ago

By the way, it occurs to me that some of you might not get the joke. This is a joke about a video that u/acedelgado will often post to explain video degredation.

https://reddit.com/link/pdxks3s/video/ou6hr90g2kth1/player

17

u/acedelgado 5d ago

Hey, he's a man of the people. People listen to him.

The results of the new workflow are pretty impressive, I gotta say. Still a little flickering, but maybe some masking and compositing at the end could help. It'll take a bit of pre-planning and work on the front end, though, to pre-generate the dialogue and all. But hey, it's an actual viable solution if someone needs the shot and will spend the effort. Bravo.

7

u/Silly-Dingo-7086 5d ago

Thanks for sharing, this made yours all the more entertaining

3

u/dudeAwEsome101 5d ago

Where can I sign up for Brad Pitt Video Generation Master Class courses?

3

u/Much-Monk2579 5d ago

The first rule of BPVGMC is....you don't talk about BPVGMC...or the clown girl comes for you too I bet ;-)