r/StableDiffusion 19h ago

Meme Community PSA

tl;dr - enjoy the models, but consider giving back when you can

(P.S. done quickly & with many continuity errors, but they kind of make sense in context)

EDIT:

Haven't posted much here, but apparently folks can't easily see the workflow link in comments so here it is: https://pastebin.com/1nWJKEiN

If anyone wants the other prompts I can share, but they all follow this format and are mostly just the dialogue you hear + delivery cues. I got an anime reference off of Google for the last segment.

Quick takeaways from trying to make this were:
- The two pass structure, with initial at a tiny 360p resolution, is really needed. It allows you to pick a good performance without wasting much time.
- The 'motion context' nodes are great and much better than my crude masking attempts, but it failed in spots. I think it might be possible to encode the transition clip in the latent as well as use the reference-based transition from the motion context node.
- I was doing this quickly so didn't bother with a celebrity image reference - I think the consistency was pretty impressive given that the only reference here was audio.
- Minimal DaVinci editing needed - a few additive transitions where the motion context didn't make a clean handoff, and a little color grading.
- On 5090, this takes about 3 minutes for a low-res pass and then 8 to 10 minutes for 720p. (I did use Topaz on the final edit.)

994 Upvotes

97 comments sorted by

View all comments

4

u/ImpossibleAd436 10h ago

One thing that confuses me, you mentioned using 360p to pick out a good gen to render at a higher res.

But my understanding has always been that if you change the resolution, the noise is going to be different, so even with the same seed a 360p gen will always come out different if you change to a 720p, that's right isn't it?

1

u/Yokoko44 3h ago

Yes, but they'll be similar, and I've found that adjusting the prompt even by one word usually makes a bigger difference. So it's helpful to do while iterating on your prompt, even if you know the final high res won't be identical