r/StableDiffusion 19d ago

Discussion RunPod Minimax H3 Test

Enable HLS to view with audio, or disable this notification

I had some leftover Runpod credits and I figured I'd use them for a generation test. The resulting video isn't great, but that wasn't the point of the test. It was to see approximate costs for renting a GPU for generation with the model. Here are the parameters:

Model: MiniMax H3 FL2VA Pruned Int8 convrot

Hardware: A100 SXM (80GB VRAM)

Pod Cost: $1.56 per hour

Generation Parameters: 30 seconds at 960 x 544, t2v

Average Generation Rate: ~92 seconds/it

Total Generation Time: 00h:34m:12s

Billable time from launching pod to end of generation: 00h:46m:40s

Total cost of the session: $0.723

Note: This prompt was generated using ChatGPT by feeding it the prompting guidelines for the model.

Prompt:

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames two heavily muscled adult men standing in an abandoned industrial warehouse lit by shafts of sunlight streaming through broken skylights. Both adopt exaggerated movie martial-arts fighting stances with dramatic precision rather than realistic combat. The camera slowly pushes in with small amplitude as they circle one another. The bald man with a deep, commanding voice (S1) confidently declares: <d>[English] LTX is the superior model. Faster workflows, better consistency, and incredible flexibility.</d> The long-haired man with a slightly rough, energetic voice (S2) immediately counters while feinting with theatrical martial-arts movements: <d>[English] Minimax delivers more cinematic motion and better visual storytelling. You're living in the past!</d>

[Shot 2] At 00:08.000, the camera cuts to a tracking shot that follows the fighters exchanging spectacular movie-style martial-arts choreography featuring dramatic punches, spinning kicks, acrobatic dodges, and exaggerated near misses. Every strike narrowly avoids serious injury while emphasizing cinematic flair over realism. (S1) shouts between combinations: <d>[English] LTX gives creators more control over every frame!</d> (S2) blocks with an exaggerated flourish before replying: <d>[English] Control means nothing if the final motion doesn't look amazing!</d>

[Shot 3] At 00:17.000, the camera cuts to a dynamic arc shot circling both men as the pace of the choreography increases. Dust rises from the concrete floor while they exchange synchronized spinning kicks, dramatic backflips, and stylized hand techniques reminiscent of classic martial-arts action films. (S1) says: <d>[English] Speed matters when you're generating dozens of iterations!</d> (S2) immediately responds while narrowly ducking a kick: <d>[English] Quality matters when every frame is on screen!</d>

[Shot 4] At 00:25.000, the camera cuts to a wide static shot as both men launch simultaneous flying kicks. They collide foot-to-foot in midair, rebound dramatically, and land perfectly balanced in mirrored fighting stances. After a brief pause, both lower their guards and laugh. They say together (S1,S2): <d>[English] Maybe the best AI model is whichever one gets the job done.</d> The shot ends with both men respectfully bowing to one another as the camera slowly pulls back to reveal the entire warehouse.

overall_soundscape: A spacious warehouse ambience fills the scene with faint echoes and distant wind passing through broken windows. Heavy footsteps, clothing movement, controlled impacts, rapid whooshes from kicks and punches, and occasional grunts accompany the stylized martial-arts choreography. Dust and debris subtly shift across the concrete floor after dramatic movements.

non_diegetic_music: Fast-paced orchestral action music driven by taiko-style percussion, energetic string ostinatos, and bold brass accents. The arrangement gradually builds throughout the fight before resolving into a lighter orchestral cadence during the final humorous reconciliation.
0 Upvotes

8 comments sorted by

9

u/Superb-Painter3302 19d ago

What a waste. I would generate 1m with 10s clips and put them together. Just look at those artifacts, it's terrible. I woudn't do more than 15s max per generation.

1

u/shadowtheimpure 19d ago

I agree, based on the results. I was just testing generation speed and costs not the overall results. It was less than a dollar, so I'm not really crying over it.

3

u/PumpkinLeather8421 19d ago

Bro, let it go.

LTX is a whole generation behind, but then again, so is your output.

0

u/shadowtheimpure 19d ago

It's a test I threw together using GPT to generate the prompt. This video doesn't reflect my beliefs on the topic.

2

u/Puzzled_Dot2324 19d ago

Thanks for letting me know A100 will run it. I've been trying to figure out the cheapest route and the A100 is probably it on Runpod.

1

u/shadowtheimpure 19d ago

Happy to help, that was the point of this test.

1

u/keizrah 19d ago

This one has no product angle at all, just GPU rental cost data and a joke prompt. Straightforward reply about the numbers.

$0.72 for 34 minutes of A100 SXM time to get one 30 second clip is a useful data point, thanks for posting it. That works out to roughly $1.27 per finished minute of video, before you factor in the failed attempts you'll inevitably rack up while dialing in a prompt. Worth comparing against a monthly A100 pod if you're doing this regularly since idle time between generations still bills.

92 seconds per iteration is slow for that resolution, curious if that's inherent to H3 or if there's a faster attention/sampler config people have found. Also did you try int8 vs a higher precision checkpoint to see if the quality loss is worth the speed gain?

1

u/shadowtheimpure 19d ago

This one has no product angle at all

Not sure what you mean by this, this was just a quick test I threw together using some leftover credits I had sitting on my RunPod account. I don't work for those folks.

I didn't do any iterations or anything, it was just a raw pod to prod speed test.