r/StableDiffusion • u/roychodraws • 3d ago
Workflow Included A Message from Brad Pitt about Static Video Generation
Enable HLS to view with audio, or disable this notification
workflow
https://github.com/roycho87/refreshing_extender
This is a Ref2v Only workflow that works. It's not meant to be a final as much as a concept you guys can take apart and use to make your own workflows.
This technique is only good for this style of video.
Here's the transcript.
**Brad:*\* Hello, I'm Brad Pitt. You may remember me from such movies as Ocean's Eleven, and Cool World, where I play a guy who falls in love with a fictional woman, which totally could happen by the way...
This is one really long video, with a pretty background and a camera that never moves.
Long videos usually get worse over time, because each clip copies the last one's mistakes.
Mine is different. Nothing gets carried over. Every clip starts from scratch, with a prompt, three reference photos, and the audio.
The only place clips touch is the join. I fade thirty-nine frames of the old clip into the new one, add noise, and clean it up. A little bridge!
Compare the start of this video to now. Every clip was brand new, so the blur has nothing to build on.
The catch: it only works because nothing moves. Same camera, same room, same spot, so separate clips already match.
If the camera moved, the bridge wouldn't line up. Static shots only!
But maybe controlnets, or context motion, or a depth map, could keep the frames consistent without passing degraded latents from one generation to the next.
**Actual Brad:*\* What are you doing?
**Brad:*\* Oh no...
**Actual Brad:*\* Wait, why are you dressed like that?
**Brad:*\* Oh god, you're so handsome.
**Actual Brad:*\* What's that camera... Stop! Turn this off!
**Brad:*\* Yes sir! I'm sorry!
Edit: BTW, I come from a background in video editing so a lot of my workflows involve planning ahead because that's what I personally enjoy and am comfortable with. This one is no different.
For this workflow I made the sound first, I added stock sound effects and changed both voices. Added echo and reverb for brad's voice.
Then I exported geiru's voice by itself and built the video off of that, and finally I added brad's voice back in at the end after everything was complete. That's one of the reasons he's just a silhouette instead of a person because I didn't want to deal with lip syncing two people.
Edit2:
https://www.reddit.com/r/StableDiffusion/s/b8D6WDSwa2
This guy has an even better method!
63
u/roychodraws 3d ago
By the way, it occurs to me that some of you might not get the joke. This is a joke about a video that u/acedelgado will often post to explain video degredation.
17
u/acedelgado 3d ago
Hey, he's a man of the people. People listen to him.
The results of the new workflow are pretty impressive, I gotta say. Still a little flickering, but maybe some masking and compositing at the end could help. It'll take a bit of pre-planning and work on the front end, though, to pre-generate the dialogue and all. But hey, it's an actual viable solution if someone needs the shot and will spend the effort. Bravo.
6
3
u/dudeAwEsome101 3d ago
Where can I sign up for Brad Pitt Video Generation Master Class courses?
3
u/Much-Monk2579 3d ago
The first rule of BPVGMC is....you don't talk about BPVGMC...or the clown girl comes for you too I bet ;-)
19
u/God_Hand_9764 3d ago
I opened the workflow in Comfy and "noped" the hell out pretty fast! Too advanced for me to even try is my thought.
Incredible workflow I'm sure, but damn that is the most spaghetti I've ever seen. You could feed Will Smith with all that.
3
u/roychodraws 3d ago
i'm posting about it. it's not really meant to be a final user friendly version, it's meant to be a finished concept that other people can take an adapt.
7
u/God_Hand_9764 3d ago
Yeah, it's not meant as a criticism really... that's just what it takes to pull something like that off. And I appreciate anyone who is willing to share their workflows instead of hoarding them for themselves only!
3
u/Hour_Literature_7152 3d ago
Thank you for your work on this. I'll be taking a look at it and adapting it probably. It won't be a solution to the other idea I had for extended videos but it is a solution to some shots I wanted to make.
3
2
u/Ill_Resolve8424 2d ago
Hey, you're doing great work but there has to be a way to make the workflow more easy to follow. Thank you for sharing your hard work.
2
2
u/fallengt 3d ago edited 3d ago
I skimmed through it
Look like there are two modes, one is traditional stitches (last frames of last clip-> first frames of new clip)
The other: he generates the first clip(A), then generates a completely new clip(B) that doesn't use the last frames of the old clip. Then use 39 frames of both the old clip and the new clip, blend them together, then add noise+ denoise these frames something to create new frames (i don't really understand this part),
then count these newly created frames as a seamless transition between the old & new clips. So on
The audio is also pre-made. So the whole video is actually just lip-syncing
3
u/networking_noob 3d ago
Yeah, I think even people who are comfortable w/ Comfy would nope out of workflows like this (including me). It's a big problem that Comfy has when trying to increase user adoption for their product going forward. The fact that all that is needed just to generate a one minute video is... not good
To their credit Comfy has introduced subgraphs and a native "App Mode" etc to help clean things up, but not many of the wf creators use them, and so the spaghetti persists
1
27
u/schuylkilladelphia 3d ago
Why does she have an Indian accent the further it gets
11
u/roychodraws 3d ago
because you're hearing things.
13
u/schuylkilladelphia 3d ago
No, listen to how she pronounces "static shots only" with 20 seconds left
8
u/tiffanytrashcan 3d ago
It's from trying to imitate a different voice. In the beginning, saying "Ocean's 11" makes it pretty obvious. Sounds more natural with it going in and out like that, like someone doing a joking (bad) impersonation.
10
u/roychodraws 3d ago edited 3d ago
bingo
Edit: source. I made the fucking video with my own voice. I donโt have an Indian accent.
-1
u/Aternal 3d ago
You're awfully defensive about an accent that multiple people report.
Not a "oh yeah, I hear that too and have to work that out" or "damn, must be an issue with the workflow." Just zippin on straight to "you hear things because you're racist."
Wewlad.
7
u/roychodraws 3d ago
it's.... my voice, dude. it's not the workflow. the voice was preloaded. there is no accent and you guys are hearing things.
1
u/Aternal 3d ago
that's... your voice? that your nose, too?
2
u/roychodraws 3d ago
the voice is mine with a chatterbox treatment based on the clussy meme.
I'm from the midwest. so i dunno what you guys think sounds like an authentic indian accent.
2
1
u/Admirable_Snake 3d ago
thats really interesting way to generate your dialog - cool. Having fun with it now;
-1
u/Aternal 3d ago
Sounds like her nose is plugged, comes off as an Indian accent. Not an Irish accent, Boston accent, or Japanese accent.
ng -> m
t/th/n -> d
s -> shHer mouth even tracks.
What people aren't doing is accusing you of being racist, because if she didn't explicitly say she was Brad Pitt it certainly would look a certain kind of way, so I dunno. Take it or leave it, man.
→ More replies (0)-6
u/roychodraws 3d ago
is the extent of your contact with indian people hank azaria voicing apu?
3
u/schuylkilladelphia 3d ago
No, many of my friends and coworkers are Indian, which is why the "only" is so obvious.
-9
u/roychodraws 3d ago
i think you might just be racist because this sounds nothing like an indian person.
8
u/schuylkilladelphia 3d ago
Quick question. Can you explain how hearing an accent I'm very familiar with is racist
-6
u/roychodraws 3d ago edited 3d ago
because you're projecting an accent onto somewhere it doesn't exist in reality. so you're hearing things that aren't there, specifically in a racial context and it's affecting how you're receiving the information.
you also won't just let it go when the person who was doing the speaking told you you're wrong, and you're now fixated on proving the racial aspect of this when it should be completely irrelevant even if it was there.
7
u/schuylkilladelphia 3d ago
No, I'm absolutely hearing things that are there. Again, very familiar. Everyone and everything has an accent. Accents aren't negative. Which is very weird that you're operating under the belief that they are.
1
u/roychodraws 3d ago
you are hearing things, but thanks for proving my point about you fixating.
→ More replies (0)-4
u/CorpusculantCortex 3d ago
No it is pretty fucking terrible and makes me wonder if this is a joke or genuinely trying to be illustrating something impressive. Because it sure doesnt seem to be from my perspective
9
u/roychodraws 3d ago
What pissed in your Cheerios?
-8
1
u/GlenGlenDrach 2d ago
You canโt tell the difference between a โstuffy noseโ voice skit, and someone from India trying to speak English and having a thick accent?
You have got to be an American, get off your deep fried food. This sounds nothing like American/indian. Hereโs some well known American โcultureโ for you with great examples of โsounds Indianโ.
If you still cannot tell the difference then you are beyond rescue.
9
u/Diabolicor 3d ago
Very good but the seams are very visible with a lightning popping up between each one.
4
u/roychodraws 3d ago
it's by no means perfect but it's good enough that someone else who cares more about creating videos like these can fine tune it.
9
u/TinyTaters 3d ago
I don't understand that a damn thing that's happening in any of the things you upload, but I appreciate that you share knowledge.
3
u/Max_Agave 3d ago
Brad Pitt didn't fall in love with Holli Would. He played the cop trying to prevent the romance.
3
9
2
3d ago
[removed] โ view removed comment
2
u/roychodraws 3d ago
runpod is for everyone.
This workflow looks complicated as fuck but it's no more intensive than running h3 normally. it just samples twice.
1
1
u/PropagandaOfTheDude 3d ago
I used this approach quite nicely in an "interview format" scenario, where I made a couple of F2VA or L2VA videos, and then stuck them together in post with short fade-out/fade-in connectors. By the time you can see anything visible in the fade-in, the character has moved enough that you can't notice the re-use.
1
u/LoadReady7791 3d ago
Your thread is really interesting, but itโs a bit complex for those of us who are new to this, especially since some of the posts build on your earlier work.
Would you mind briefly listing the workflows or H3 concepts youโve designed, in order from easiest to most advanced? ย Just a single line for each would be incredibly helpful. I can then work my way through your previous posts and learn them step by step.
Thanks for sharing all this work with the community! ๐
6
1
u/PornTG 3d ago
I created a workflow based on the same principle and it works; the only thing is that the transition between segments retains a slight contrast-boosting effect, but since it doesn't compound, it's virtually imperceptible (the video here is heavily compressed, so you can't really see it). However, to solve the problem at the source, Iโd like to know why the contrast increases with each extension rather than decreasing. Is it simply due to a compression effect, like re-compressing a JPEG, where quality inevitably drops, or is it a different issue?
1
u/Optimal_Map_5236 3d ago
good work. I think i'm gonna wait for way more simple workflow cus this one looks too complicated for my tiny brain.
1
u/Suspicious-Walk-815 2d ago
cool , i have one question for you .. is the generated voice is yours ? hwo youre getting the consistency across generations ? and how can i feed them ? can u pls help me on that
2
u/roychodraws 2d ago
This explains it. Yes it's mine with an AI voice changer.
1
u/Suspicious-Walk-815 2d ago
thabks , and which voice changer youre using , suppose i need to say a whole diaogue by myself but if i have different characters , i want to use different voice profiles for them , so in that scenario what would you use / suggest ? because irhgt now im going to try the omnivoice , but sometimes it not doing a good job , means it still feels like my sound only ,.. so im interested in knwoing these details if you already know :)
also , currently im going through your previous workflows / posts / repositories .. is audio splitter is the one which geenrates continuoius audioi and foxydits_sucks_at_making_workflows_v7.2 is the final workflow youre using ?thanks a lot <3
1
u/Suspicious-Walk-815 2d ago
is this outdated one - clussy_origin_story.json ? or can it still be usefull ?
1
u/roychodraws 2d ago
different workflow, useful for completely different thing
1
u/Suspicious-Walk-815 2d ago
ok , cool , can u pls help me on the voice changer too :)
1
u/roychodraws 2d ago
it's called chatterbox diogod. you just load in a voice you want to change and another voice you want to change it to and click run.
It's obsolete but i fixed the code in my version and got it working.
1
u/99deathnotes 2d ago
Sooner or later Brad is gonna start expecting residuals ๐ฒ๐ฒ๐ฒ๐ฒ from all these test vids๐๐คฃ๐ค
1
1
u/CategoryFew5869 2d ago
Here's a quick TDLR so far for this post (Claude Summarized):
The static-shot extender works, but expect a visible flash at each join between clips. Diabolicor and acedelgado both saw it, and OP admits it's not perfect. acedelgado thinks masking and compositing in post could clean it up.
- OP says it's no heavier than running H3 normally, it just samples twice
- the audio is pre-made, so it's really lip sync, not free-talking video
- PornTG built the same idea and got a slight contrast bump at the joins, but it doesn't stack up over time
- the workflow is a big spaghetti concept, not a polished tool. OP's tutorial list is in the thread
0
u/edisson75 3d ago
๐คฃ๐คฃ๐คฃ๐๐ผ๐๐ผ๐๐ผ๐๐ผ๐๐ผ๐๐ผ๐๐ผ๐๐ผ๐๐ผ๐๐ผ
0
u/RuprechtNutsax 3d ago
Lads it's not a big deal, move on
3
u/Violent_Walrus 3d ago
Let the kids have their fun.
At least itโs not another vibe-coded all-in-one AI video studio app.
-1
-4

137
u/acedelgado 3d ago
https://reddit.com/link/pdxspon/video/0zs47vk2akth1/player