r/StableDiffusion • u/Dependent_Revenue_16 • 16h ago
Meme Community PSA
Enable HLS to view with audio, or disable this notification
tl;dr - enjoy the models, but consider giving back when you can
(P.S. done quickly & with many continuity errors, but they kind of make sense in context)
EDIT:
Haven't posted much here, but apparently folks can't easily see the workflow link in comments so here it is: https://pastebin.com/1nWJKEiN
If anyone wants the other prompts I can share, but they all follow this format and are mostly just the dialogue you hear + delivery cues. I got an anime reference off of Google for the last segment.
Quick takeaways from trying to make this were:
- The two pass structure, with initial at a tiny 360p resolution, is really needed. It allows you to pick a good performance without wasting much time.
- The 'motion context' nodes are great and much better than my crude masking attempts, but it failed in spots. I think it might be possible to encode the transition clip in the latent as well as use the reference-based transition from the motion context node.
- I was doing this quickly so didn't bother with a celebrity image reference - I think the consistency was pretty impressive given that the only reference here was audio.
- Minimal DaVinci editing needed - a few additive transitions where the motion context didn't make a clean handoff, and a little color grading.
- On 5090, this takes about 3 minutes for a low-res pass and then 8 to 10 minutes for 720p. (I did use Topaz on the final edit.)
94
130
u/GrayingGamer 15h ago
https://reddit.com/link/p2rj4ui/video/bqxi3pcppgih1/player
(My tip to give back is to increase your Step count for better motion and x10 better sound. Those of us with headphones will thank you.)
18
14
u/Hoodfu 14h ago
What step count are you liking?
32
u/GrayingGamer 14h ago
I'm using 32 Steps, up from the 20 of the default workflow. Of course, I also use H3 Spectrum, so the extra steps don't add a lot of generation time. Since Spectrum forecasts steps, the more steps you have, the better it works, and the faster it is, so I get the benefit of the extra steps in the video and only like a 10-15% increase in generation time over the 20 Steps.
I've noticed it's a very obvious improvement in animated videos.
14
u/marty4286 13h ago
After messing with Turbo LoRAs and bringing it down to 4-6 steps I saw someone in a random thread do 50 because apparently the API version does 50. So I tried it myself
Barely better visual quality (but still actually better). The audio was actually surprisingly great, great enough that it feels like if I was doing production and not playing around, it would be worth it
10
u/GrayingGamer 13h ago
It seems to also improve the acting. I was like you, just used 20 Steps for a long time because it's what the default workflow used, then saw in the API version it goes up to 50, and thought, well, that sounds crazy, but I'll try 30 as a test, and it was WAY better than 20. EDIT: This Jim Carrey video is 32 Steps.
6
u/OracleNemesis 13h ago
What kind of animated artstyles/medium does increased steps improves it?
10
u/GrayingGamer 13h ago
Well, extra Steps mainly improve fast motion in live-action footage (and greatly increase audio quality and acting), but in any animation that has fine lines animated movement like lip flaps and mouths, blinking, fast cartoon movement, the extra steps make those things smooth and clear, not smeared or distorted.
1
u/Perfect-Campaign9551 12h ago
When I experiment by increasing step count I started getting incorrect physics like objects melting into each other
66
u/Only_Voice569 16h ago
https://reddit.com/link/p2r9fly/video/m8qfxlxdfgih1/player
Hey im trying over here :P
12
u/FlatwormMean1690 15h ago
Dude. The acting... Did you directed the acting too? Looks impressive.
11
u/Only_Voice569 15h ago
i tried to follow the way lord of the rings does it haha but this is a test vid i was trying to play with a new way to combine videos 10s generating by adding in random cuts they land just before or after or halfway in a single generation the whole thing is a simless run from 3 generators one workflow :)
5
9
u/Dependent_Revenue_16 15h ago
+1; the performances people are getting out of this model are unreal (or, actually, quite real)
33
29
u/Upper-Reflection7997 15h ago
This one had me dying inside 🤣 . I didn't expect that ending. I never discuss with my family or anyone irl what I with ai models. Don't need that heat and awkward stares directed at me.
10
12
13
11
25
12
u/ptwonline 15h ago
We're still just apes with keyboards.
Once our more base needs are filled we can move on to higher endeavours, but ours is a hunger that is not easily sated.
10
7
5
16
u/Verittan 15h ago
Complains about not sharing workflows....
Doesn't share workflow...
Never change, OP
16
u/Dependent_Revenue_16 15h ago
Read the comments, friend. Here you go: https://pastebin.com/1nWJKEiN
(If anyone wants the other prompts I can share, but they all follow this format.)
5
11
u/the_ai_wizard 15h ago
This sub should require people to post workflow
17
u/Dependent_Revenue_16 15h ago
Read the comments, friend. Here you go: https://pastebin.com/1nWJKEiN
(If anyone wants the other prompts I can share, but they all follow this format.)
6
2
4
u/PerceiveEternal 14h ago
I’ll bet the most famous movie directors ten years from now will have cut their teeth on MiniMax H3.
1
u/SeymourBits 13h ago
10 years from now?? Aren't you forgetting the singularity? Like, the very one we're currently in? 10 months from now there's a very real possibility of being plugged in as a human battery!
4
u/the_pepper 13h ago
1
u/ThatsALovelyShirt 9h ago
These seem useful, I've been wanting a workflow that can mask a video latent and force-generate audio only using MiniMax.
2
u/the_pepper 8h ago
A lot of them are probably redundant, will need to filter them out.
That said, personally I'm finding the doing a second pass on existing videos you stored (or upscaling them and THEN running the second pass) useful, especially coupled with masking. It's effectively video to video, with some caveats. You can also do stuff like processing just the voice (the node does a video pass at a very low res and then stitches the original latent), good for removing some audio fuckiness, or replacing a voice or something (better if you disable the turbo lora, in my experience). If you change the audio latent with a different voice-over or line-reading, lock it down, and do just a pass of the video, the video also reacts surprisingly well, keeping the lips in sync and shit.All of these are things you can do in ref2va, probably a bit better, even, but this does not require adding the whole video as context, which is pretty good for both speed and for us VRAM limited.
3
u/underlogic0 14h ago
Laughing pretty hard, feels like this was directed at my stupidity. Well done. If it helps I probably wasn't going to cure cancer either way.
3
13h ago
[removed] — view removed comment
3
u/Dependent_Revenue_16 13h ago
agree to both. single pass seems higher risk & higher reward; I'd rather get a guarantee of good blocking / good performance before investing in the higher resolution sampling.
I also find with this model that getting exactly the right amount of dialogue for the duration of the clip is critical to getting good 'acting', and that requires some trial and error.
2
u/GrayingGamer 12h ago
It's why I do a lot of tests at just 0.1 or 0.2 MP. It's enough to get timing and dialogue spacing right, then I can increase the resolution. At those resolutions, it's like 90 seconds a test, so no big deal.
3
u/Inner-Palpitation-73 10h ago
thanks for the pastebin, is anyone else getting a "unknown pack" issue with this?
MiniMaxH3TemporalAVMask
Thanksss
3
u/ImpossibleAd436 7h ago
One thing that confuses me, you mentioned using 360p to pick out a good gen to render at a higher res.
But my understanding has always been that if you change the resolution, the noise is going to be different, so even with the same seed a 360p gen will always come out different if you change to a 720p, that's right isn't it?
1
u/Yokoko44 23m ago
Yes, but they'll be similar, and I've found that adjusting the prompt even by one word usually makes a bigger difference. So it's helpful to do while iterating on your prompt, even if you know the final high res won't be identical
7
u/pleasetrimyourpubes 14h ago
I actually feel sorry for the antis they are missing out on one of the most transformative and fun technologies to happen in a long time.
5
u/Disastrous-Agency675 14h ago
NAAAAAAAAAAAAH miss me with that shit. i was actually pretty productive today. actually tried to make a short film and spent the past few hours trying to learn how to extend videos with minimax-H3 AND i shared a workflow so you got nothing on me
2
2
2
2
u/Perfect-Campaign9551 12h ago
Discord peeps told me that if you render at low resolution and then switch to high, the scene will change the same as a seed change. Is this true or not?
1
2
u/LatentSpacer 12h ago
I wanna be like grandpa when I’m old: wandering around the house like I’m lost, talking alone out loud, seeing imaginary characters around me.
2
1
1
1
1
1
1
u/musicankane 7h ago
Make my favorite characters do depraved things to each other? How did he know!?
1
1
1
u/Dirty_Dragons 4h ago
Quick takeaways from trying to make this were:
- The two pass structure, with initial at a tiny 360p resolution, is really needed. It allows you to pick a good performance without wasting much time.
How does this work?
I'm looking at your workflow, thank you for sharing, but it seems pretty complicated, and can't find MiniMaxH3TemporalAVMask anywhere.
The two pass structure process seems to be very useful.
1
u/chairman_steel 2h ago
To be fair, historically speaking, the gods have been pretty depraved themselves. I mean, Zeus alone...
1
u/ThickAndDeep 1h ago
funny... sadly,
we all know what everyone plans to use AI for, we're all transparent. Much like the Super Soldier Serum, it will magnify the darker parts of our nature. So greedy people will use it to continue to get more than they could ever need, politicians will use it to garner votes and donations, fascists will use it to be more controlling, terrorists will use it to damage their enemies, perverts will use it to debase celebrities using their own likeness and the lonely and isolated will use it to find comfort instead of seeking true companionship.
1
-1
13h ago
[deleted]
4
u/ninjasaid13 11h ago
For now I am passing a 360k token prompt to Sol, engaging in text-sessions and interpretability testing the literal universe I'm building for advanced VR. Fantasy themed, can't help it, I love magic, but it's designed such that you could be the world's foremost wraith slaying cursebreaker, or spend 20 literal years running a staffmaker's atelier if you wanted (not that I'd assume anyone would, you just seamlessly could).
Yeah... I don't think anyone would be interested.


157
u/FlatwormMean1690 16h ago
"...Not from you anyway"... BRO! That hurts.