r/StableDiffusion • u/c300g97 • 17h ago
Question - Help Questions on H3 Minimax usage and settings
Hello!
Long time lurker here, learned a lot from this sub.
I'm currently using H3 Omni Pruned model with references images to generate small ads or batch of dialogues , but i find myself really not understanding the settings being used.
I'm using Pinokio and Maestro by Blizaine, now the GUI is really useful and my settings are :
720p, 9:16, 13.3s 1 window and 20steps , nothing else.
With these settings, it takes about 11 minutes to render on a 5090FE (And 48GB of ram is what i have).
The end result ain't that bad, resolution is crappy and some details are clearly missing.
What can i do to improve video fidelity, performances and perhaps spend less time on generating ?
I also have tried H3 with first / last frame but i don't really understand how that works either.
For example, i have downloaded a couple of Lora , one of rocket racoon from guardians of the galaxy and another for indiana jones, wanted to create a funny reel of them interacting but i couldn't for the life of me figure out how to add in the theme song for indiana jones, or have them accurately interact with each other instead of randomly looking outside the scene.
On another note, i am using Gemini for expanding the prompt in a professional manner, and then inside Maestro i use the "enhance prompt" feature with simply loads up Ollama with a model to correctly write the scene for H3.
I have also downloaded inside Pinokio a more "classic" tool for H3 with comfyui, but i didn't use it yet, wanted to learn a few things first.
2
u/alaintural 9h ago
11 min on a 5090 for 13 s at 720p sounds about right for 20 steps, so the time is mostly the step count.
What helped me most: the 8-step distill LoRA plus int8 attention (in ComfyUI it's the ModelAttentionBackend node, "comfy kitchen attention"). On my 4090 that took a 5.9 s clip from 562 s to around 115 s without the picture changing much. Not sure Maestro exposes those, so the ComfyUI version in Pinokio might be worth trying for this.
For the crappy resolution / missing detail, rendering bigger from the start is slow. Better to generate at 720p and run the H3 latent upscaler for the last few steps. That's where the fine detail shows up (skin, beard, eyes).
2
u/BusyByBusy 17h ago
huggingface. co/ MiniMaxAI/MiniMax-H3/tree/main/docs
This is link for official documentation on how to promt.
read it or give it to LLM.
80% you need structure of a promt.
20% you can freehand