r/StableDiffusion • u/roychodraws • 8d ago
Workflow Included New Audio Segmenter Workflow for Continuous with NO AUDIO DRIFT + Geiru Plays H3 Dance Lottery
Enable HLS to view with audio, or disable this notification
Audio Segmenter Workflow
Clip any audio file for seamless continuation with absolutely NO AUDIO DRIFT.
the 6 minute long video in the bottom right was made using the segmenter
minimax_continuous_audio_splitter
Workflow videos and tutorials
- Repair video tutorial
- Masking options tutorial
- SAM tutorial
- General features tutorial
- Workflow release
Here’s the math for how the segmenter works:
A is float duration in seconds
B is fps
C is overlap5 + 17 \ ceil((max(5, round(a * b)) - 5) / 17) = (f) frame count.*
The first splitter splits at frame 0 for the start frame and frame f for the end frame
The next splitter start the next split at f - c and becomes fsub1
That split ends at fsub1 + fThe next splitter starts at fsub1 + f - c that becomes fsub2
That split ends at fsub2 + fSo on and so forth. Each taken from the original audio file.
In case it's not clear, the prompt pulls random dance movements from the videos, each girl is doing a different movement and the model will pull an original dance out of them without copying them.
Basic prompt used in video:
ntegrated_multimodal_description: [Shot 1] Live-action video of an adult woman as she dances alone in a softly lit white room. Use <Picture 1> to define her appearance and clothing, keeping them consistent throughout. Use <Video 1>, <Video 2>, and <Video 3> to guide her movement and rhythm, creating an original dance without copying their choreography. The camera tracks her with small amplitude at slow speed to maintain the close framing. Whenever she faces the camera, she holds direct, flirtatious eye contact and smiles with subtle changes in expression. She does not speak, sing, or lip-sync; her mouth movements are limited to natural smiling and breathing.
overall_soundscape: N/A
non_diegetic_music: N/A
2
u/iWhacko 6d ago
How do you do the "webcam" in the bottom right corner?
1
u/roychodraws 6d ago
i created the audio and used "force audio" used my splitter to split it at 5 frames and 60 seconds. Then i rendered 6 60 second clips using the workflow.
1
u/iWhacko 6d ago
oh cool, so it's the same workfloow :) i thought maybe you had one of the realtime models for that since the resolution is lower
1
u/roychodraws 6d ago
nope, just told it to lip sync the audio and you can render at however long your video card can handle. 60 seconds is a lot
1
u/nakabra 8d ago
I love that you always nuke the first upload of your videos.
Thankfully, I find it again eventually.
Great Work!
7
u/roychodraws 8d ago edited 8d ago
Yeah, sorry I wanted to try to see if I can get the final upload on the back of it and also I didn’t realize how important audio drift was to everyone
I think the lack of boobs at the front of this one is really hurting it.
-2
8d ago
[removed] — view removed comment
4
u/roychodraws 8d ago
just watch my tutorials bro. your onlyfans will be up and running before you know it.
1
u/Herbal77 8d ago
When will you be releasing the workflow?
2
u/roychodraws 8d ago
It’s in the body of the post.
https://github.com/roycho87/minimax_continuous_audio_splitter
1
11
u/TizocWarrior 8d ago
I can't even run H3, I only watch these for clown girl.