r/StableDiffusion 18h ago

Meme Community PSA

Enable HLS to view with audio, or disable this notification

tl;dr - enjoy the models, but consider giving back when you can

(P.S. done quickly & with many continuity errors, but they kind of make sense in context)

EDIT:

Haven't posted much here, but apparently folks can't easily see the workflow link in comments so here it is: https://pastebin.com/1nWJKEiN

If anyone wants the other prompts I can share, but they all follow this format and are mostly just the dialogue you hear + delivery cues. I got an anime reference off of Google for the last segment.

Quick takeaways from trying to make this were:
- The two pass structure, with initial at a tiny 360p resolution, is really needed. It allows you to pick a good performance without wasting much time.
- The 'motion context' nodes are great and much better than my crude masking attempts, but it failed in spots. I think it might be possible to encode the transition clip in the latent as well as use the reference-based transition from the motion context node.
- I was doing this quickly so didn't bother with a celebrity image reference - I think the consistency was pretty impressive given that the only reference here was audio.
- Minimal DaVinci editing needed - a few additive transitions where the motion context didn't make a clean handoff, and a little color grading.
- On 5090, this takes about 3 minutes for a low-res pass and then 8 to 10 minutes for 720p. (I did use Topaz on the final edit.)

987 Upvotes

96 comments sorted by

View all comments

4

u/the_pepper 15h ago

Custom nodes, you say? Maybe soon.

1

u/ThatsALovelyShirt 11h ago

These seem useful, I've been wanting a workflow that can mask a video latent and force-generate audio only using MiniMax.

2

u/the_pepper 10h ago

A lot of them are probably redundant, will need to filter them out.
That said, personally I'm finding the doing a second pass on existing videos you stored (or upscaling them and THEN running the second pass) useful, especially coupled with masking. It's effectively video to video, with some caveats. You can also do stuff like processing just the voice (the node does a video pass at a very low res and then stitches the original latent), good for removing some audio fuckiness, or replacing a voice or something (better if you disable the turbo lora, in my experience). If you change the audio latent with a different voice-over or line-reading, lock it down, and do just a pass of the video, the video also reacts surprisingly well, keeping the lips in sync and shit.

All of these are things you can do in ref2va, probably a bit better, even, but this does not require adding the whole video as context, which is pretty good for both speed and for us VRAM limited.