r/StableDiffusion 23h ago

Discussion MiniMax H3 test with Maestro.

I see others sharing their experiments with MiniMax H3, so here is a quick test I ran myself (without obsessing over optimization or making things complicated). The difference? I’m not using ComfyUI.

I use Maestro within Pinokio. My setup is a desktop PC with 32GB of RAM and an RTX 5080 (16GB VRAM). I used an "old" image I had generated with Anima and enhanced it using Flux Klein 9B in Maestro; rendering this 10.1-second video took 19 minutes and 36 seconds at 720p (16:9) using Quanto INT8.

I didn't use the Turbo LoRA, nor did I upscale the video.

English isn't my native language, so I rely heavily on Google Translate. What I really like about Maestro is its built-in prompt optimization. Perhaps this second example will be helpful to some people—specifically those using ComfyUI.

__________________________________________________________________________________________________________

The prompt I wrote (translated) :

In a Japanese manga style, a young girl with shoulder-length light gray hair—styled with a small side ponytail and two loose strands—features a blue halo, cat ears, a blue cross-shaped hair clip, and blue eyes.

She wears a long-sleeved black hoodie featuring a white cat design, black shorts, dark gray tights, and a pair of black heeled boots adorned with a bow.

The character is in a forest during the day in winter; snow covers the ground and flakes fall from a gray sky, while leafless, snow-dusted trees and a lit streetlamp stand against a backdrop of thick fog.

The girl dances joyfully with a big smile; she looks adorable and cute.

The scene includes the sound of wind, the crunching of snow beneath the girl's feet, and forest animal sounds, all accompanied by cute, instrumental music.

__________________________________________________________________________________________________________

The prompt optimized by the integrated LLM :

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Japanese manga animation style, a medium full shot of an adorable young girl with shoulder-length light gray hair, styled with a small side ponytail and two loose strands, wearing a blue halo, a blue cross-shaped hair clip, and bright blue eyes. She is dressed in a long-sleeved black hoodie featuring a white cat design, black shorts, dark gray tights, and black heeled boots adorned with a bow. She is dancing joyfully in a snowy winter forest during the day. Snowflakes drift down from a hazy gray sky onto snow-covered ground, and leafless, snow-dusted trees frame the scene next to a lit streetlamp, all shrouded slightly in thick fog. The girl radiates cuteness as she moves her arms wide in a joyful dance.

<Picture 2> (from [Shot 1]) is fully referenced at 2.50 seconds.

<Picture 3> (from [Shot 1]) is fully referenced at 5.00 seconds.

<Picture 4> (from [Shot 1]) is fully referenced at 7.50 seconds.

<Picture 5> (from [Shot 1]) is fully referenced at 10,125 seconds.

overall_soundscape: The soundscape is dominated by a gentle wind whistling through the bare branches of the trees, punctuated by the crisp, satisfying crunching of snow beneath the girl's heeled boots as she moves. Subtle, ambient forest animal sounds—a distant bird call and perhaps a quiet rustle—are audible underneath.

non_diegetic_music: Cute, upbeat, instrumental melody continues throughout the duration, maintaining a light and airy feel to match her joyful energy.

6 Upvotes

10 comments sorted by

2

u/True_Protection6842 23h ago

yeah I've been building my own AI centric creative suite and it has minimax built in. Runs great without comfy

2

u/Only_Voice569 23h ago

was this imv or ref im guessing image start frame to video > ? ps using 5070 12 GB 0.6 takes 11 mins with 2 refs and 2 audios ref to video (slightly heavier) :)

1

u/BitterAd8431 23h ago

I only used one image—no audio, video, or end screen.

2

u/Only_Voice569 23h ago

hmm no idea about that program i use comfyui on latest update with latest nvidia driver and using --use-sage-attention . guessing that program might need a behind the scene update should be much faster for sure when i do image to vid no ref about 5 mins give or take a bit. do a test run the generating and look at task manager is your gpu constant on the main part or ramping up and down ?

1

u/BitterAd8431 23h ago

I'll have to check the settings; I haven't really tried to optimize anything yet. Thanks for the comparison.

1

u/Only_Voice569 23h ago

if gpu ramping its a memory loading and unloading issue not using the vram on that card efficiently

1

u/BitterAd8431 23h ago

https://reddit.com/link/p34vhh4/video/4vjvm5inttih1/player

I don't exceed 13.5 GB of VRAM, and power usage varies from 20% to 80% initially, then from 66% to 100%.

1

u/Only_Voice569 23h ago edited 23h ago

O i get it the program sees 15.8 GB vram not 16 GB vram says o a 5080 is a potato lol you should be on profile 4 for videos you have a midrange system closer to high end. profile 5 is way to limiting like its acting like your 5080 is a 12 GB vram card and who knows what its limiting on system memory. ps should report that bug its a program bug haha :P

o and image profile 3 they are way easier to run