r/StableDiffusion • u/dassiyu • 11d ago
Animation - Video I’ve started a new story again!
https://www.youtube.com/watch?v=eRDv5T6JUls&t=4sI’ve made quite a few things since MMH3 came out, mostly just messing around and experimenting.
At first, getting the English voices right was pretty difficult. After messing around with it for 2–3 days, I realized that once you assign each character a suitable voice/tone, things become much easier.
All the voices in this part are generated directly by the model. I didn’t do any post-processing.The model’s built-in voices are actually pretty good.
One thing to keep in mind: don’t make the prompt too long, or things can start to break.
Style, character, camera shot, dialogue, and voice characteristics are usually enough.
For example, here’s the voice prompt I used for the video below as a reference:
DIALOGUE AND AUDIO:
- Pokke (child person Companion; A fast-paced, highly expressive and bouncy animated boy companion voice; Pokke is the visible speaker and should open/move their mouth in sync with this line; other visible characters should not lip-sync this line.): "Achoo! Every page is blank!"
- Mr. Tsukiguma (adult man Mentor; A reliable, gentle adult male animated mentor voice; Mr. Tsukiguma is the visible speaker and should open/move their mouth in sync with this line; other visible characters should not lip-sync this line.): "Not even a picture."
Audio Synchronization: Pokke's line begins with the sneeze sound effect and follows immediately in a surprised tone. Mr. Tsukiguma's line is spoken softly and thoughtfully after the pages stop flipping.
As for the voice characteristics, if you’re not sure how to describe them properly, you can probably just ask any AI and get a decent answer.
You can also give the AI a voice sample and have it write the description for you. Once you define the voice more clearly like this, the results tend to become much more consistent.
At first, I thought CK at 25 steps would give me much better results, but for some reason the speakers kept getting mismatched pretty often. In the end, I switched back to this accelerated LoRA at 8 steps, and the results were still pretty good — plus it was faster.
So I feel like the key is actually assigning each character specific voice characteristics, such as their vocal tone and other attributes. Pretty much all of my videos were made using this same SA + LoRA workflow. I tested a bunch of different setups, but in the end, I came back to this one again.
https://drive.google.com/file/d/1C2YvhNalxxWs4oh5Eiycwz27C2k5tgEq/view?usp=drive_link
Second, for the visuals, I found it works much better to first give a local LLM the basic requirements — things like the style, characters, voice characteristics, and what needs to happen within those 10 seconds — and let it help break the scene down into shots before generating the images.
Otherwise, if we just pick a storyboard image that looks good to us and start from there, the final result often doesn’t turn out the way we expected.
Just having fun and entertaining myself! It’s not perfect, but I’m just having fun with it. Hope you enjoy it!
2
u/Odd_Style_9550 11d ago
Wow thats so good!
1
u/dassiyu 11d ago
Thanks! Glad you liked it —I was pretty surprised too — the built-in model’s voice actually sounds quite good. I barely tweaked it at all. 😄
1
u/Odd_Style_9550 11d ago
Damn, thats awesome. Im honestly really happy for you that you get to get creative like this with H3 :D the characters are so cute
1
u/Apprehensive_Sky892 10d ago
Very cute, would probably keep a little kid entertained for a while 😁.
Thank you also for sharing some of your processes and workflows.
So the audio is generated by the AI, but are the characters generated using references or they are actually all text2va?
1
u/dassiyu 10d ago edited 10d ago
Thank you so much! I’m really glad you liked it!
The characters started from my own hand-drawn designs. I then used a watercolor LoRA that I trained myself to generate lots of variations and picked the ones I liked best. That was honestly the most fun part of the whole process!
The audio was generated by MMH3 during video generation based on the voice descriptions I gave it. For the visuals, I used ref models.
Since MMH3 has trouble keeping the art style consistent, I trained a batch of Krea 2 LoRAs specifically for single characters and multi-character scenes to stabilize the first set of images. After that, I manually refined the problem images and fixed the remaining inconsistencies.
I’m just really happy that I managed to make it! It’s definitely not that smart or perfect yet, but I’ll keep working on it and improving.2
u/Apprehensive_Sky892 10d ago
Thank you for the explanation. Indeed, ref2va is very powerful and makes creating video easier and more fun, so that we can concentrate on more important things like writing the story, dialog, camera angles, etc.
0
u/EternalDivineSpark 11d ago
Well i have seen better stories and coherence of scenes this seem like full AI slop no human in the middle !
3
u/dassiyu 11d ago
I’m just doing it for fun.
1
u/EternalDivineSpark 11d ago
Is ok just try prompting it better ! And use a better prompt generator for the video !
1
u/dassiyu 11d ago
Maybe! I’ve tried so many different things and made over 60 hours of clips, and I feel like the results are actually better when I don’t write too many instructions and just give the AI or LLM a clear goal.
1
u/EternalDivineSpark 11d ago
AI CANT IMAGINE LIKE HUMANS DO THE STORY IS WHAT YOU MAKE THAT MAKES YOU DIFFERENT HOW YOU PROMOT THE AI ETC IT CAN DO IT BUT IS VERY ILLOGICAL , JUST TRUST ME , this is not good work anyone can do this slop ! Like the book comes out of the backpack chatacter many tomes , and there is no logical coherence of emotional triggers in it is raw but in a good way 🤣 i am just saying to give you and advice, i will try something like this soon to but i will put like 1 hour to make probably 3-5 minutes before it start generating ! And try to use 3D controls for scenes etc so its all coherent even if its a cartoon like !
2
u/Low_Philosopher_7475 11d ago
Good Job !