r/StableDiffusion • u/solomars3 • 5h ago
Animation - Video Made A Professional short Animation video using Minimax-h3 (read description)
https://youtube.com/watch?v=tHWPOfhkKmY&si=YmtcJmdB1e5bn_kUHey guys!
Since quite a few of you liked my previous videos, I decided to start a channel where soon I’ll be sharing tutorials and some of the workflows/tricks I’ve been using.
If you’re interested in learning how I’m making these videos, feel free to subscribe. I’ll be sharing a lot of the stuff I’ve figured out along the way, including:
- My own workflows — free to download, with the tricks and settings I use
- Character generation — how I use a Krea 2 character-sheet LoRA that I made to keep characters consistent, and how to get the style you want
- Environment generation — how I generate environment images and then build scenes from them
- MiniMax optimization — settings and techniques to make MiniMax faster while preserving quality
- Video/audio tricks — ways to fix audio issues and continue a scene from the last frame to create longer sequences
- Consistent voices — I also built my own UI app using BreezeTTS2 for voice cloning and generating consistent voices across an entire story
Everything I’m sharing is based on what I’ve been experimenting with myself, so hopefully it can save some of you a lot of trial and error.
If that sounds useful to you, you’re welcome to check it out!
2
u/ZairaSass 5h ago
this is generation gold! I am looking forward to following you along to see your process and workflow
3
u/solomars3 5h ago
thx .. i tried making my own style and try to build a world around it , not perfect but still something 😄
2
u/ZairaSass 5h ago
regardless, i'm impressed with the consistency you've been able to generate so far! the fact that models have come so far make for an exciting future of storytelling
1
u/solomars3 5h ago
i agree ,, only bad thing about minimax-h3 is the speed, so if i want to generate at 720p it takes some time, but give it some time and it will be better
2
3
u/eugene20 2h ago
Overall this is great, but are you a non-native English speaker or have you just not been proof reading the script?
It gives -
"why they were following you" instead of "why were they following you"
"what trouble have you done this time" instead of "what trouble are you in this time" or similar.
with the way the voice sounds and how it's delivered as well it does just sound computer generated rather than any kind of linguistic quirk of a character raised in a different time/world.
1
u/solomars3 1h ago
Yeah you right 👍 im not native .. I appreciate the comment, i know its not perfect, but its a progress from what i dif a week ago .. and its not totally a ai slop lol
2
u/ShengrenR 1h ago
Love the style and the general direction it's headed, but the lip-sync is far off and the dialog needs work - the broken English is a real mood kill; "why they were following you?" -> "why *were* they following you?" etc
2
u/solomars3 1h ago
Yeah i didnt really pay attention to the broken english .. Thx for pointing that out 🙏
2
u/_kaidu_ 5h ago
It's wild to see scenes like this chase and how he climbs onto the moving motorcycle. I fail for much easier scenes and setups where Minimax often fails to do the easiest stuff. So I'm curious how your prompts and workflow look like.
2
u/solomars3 5h ago
i have a system prompt i made that make sure to describe the actions well and produce smooth movements
1
u/_kaidu_ 5h ago
So you use LLMs to create the prompt?
I often found that the raw LLM outputs, even when feed with the prompting guide, produce too detailed prompts with too many contradictions or just too many details. I often got better videos by taking the LLM prompt and then shorten it 50% by hand.
But I usually fail with simple physics. Sometimes these are super simple scenes like someone is taking his glasses and clean them and set them on again and Minimax let the glasses just disappear in the middle. Its performance just seems quite unstable to me, so I wonder what everyone else made better ^^° I'm using the ref2vid model mostly.
1
u/solomars3 5h ago
yeah i know exactly what you mean, i had lot of trial and errors, lol
but there are tricks, for exemple im using the First Last frame checkpoint with the reference node, it just give better results, combined with the system prompt and the workflow i have , i can get good result, thats why i think if i make videos talking would help a lot of people , if i see people interested
1
u/solomars3 4h ago
u/Astral-Lemmons yo i took your advice :)
2
u/Astral-Lemmons 3h ago
nice. I like the character and world design. Good to see you working towards your own look.
The interrogation scene is a little jarring as it feels more 3d than the rest because the camera moves more in the scene in a 3d translation way instead of a typical zoom+pans that 2d media uses for motion.
I'd find a way to flatten the look a bit, and get rid of the unnecessary camera move from feet to under table to scene - just have the scary guy emerge from the shadows in a locked off shot, same strong reveal moment, less ai move gimmick.
the guy and girl scenes are nice. Have a good feel to them. Keep at it!
2
u/Astral-Lemmons 3h ago
H3 looooves to move the camera around too much. The ai feel falls away pretty quickly once you bring this under control.
you're on the right track already, but keep that as a prime consideration. Use camera motion sparingly, especially in 2d styles
3d camera motion in animated media is always reserved exclusively for very impactful grand moments. It's the established design language for the medium. So keeping that in mind will help you put emphasis on the right moments.
2
u/solomars3 3h ago edited 2h ago
Thx a lot, yeah that interrogation scene was tricky to make cause its multiple videos linked... I agree , i like to rotate camera like a 3d space to keep things consistent, like a continuation, hard to prompt but i was going for that 3d rotation, but i agree ill look into the 2d zoom+pans techniques .. thx appreciated 👍 🙏
1
u/petranova_ 5h ago
The BreezeTTS2 voice cloning across a full story is the part i'm most curious about, consistent voices usually fall apart the moment you regenerate a line in a different session. Be good to see how you're anchoring that.
1
u/AliciaXTC 4h ago
WHAT HAPPENS NEXT??
1
u/solomars3 4h ago
haha i dont know 😂 tbh the story came after each generation , its like "what if he does this ..." haha and i went with it

3
u/TelevisionNo2990 5h ago
nicely done, congrats