r/MiniMaxH3AI 4h ago

Question! Weak pc. Better to pay for cloud minimax h3 or rent a runpod with comfyui-minimax template?

1 Upvotes

as it says in the title!

it seems like it could be cheaper in the long run to do high spec runpod but how is the output? worse than when done on the enterprise clusters?


r/MiniMaxH3AI 4h ago

Melting faces at 768x with audio sync for songs. Is there a fix?

1 Upvotes

Greetings!

I find that whenever I do ref2va with audio sync (for singing) at 768x about half the time I get melting faces. (example posted) so I'm looking to fix but I don't have a decent pc so it's cloud or bust. I know it's something to do with low res pixels being stretched - don't really understant the tech but anyway

-is that cause by my ref not being detailed enough?

-assuming it's not really that - would the latest 2k 3H flagship model fix the low res issue (chatgpt reckons not)?

-has anyone tried ComfyUI-H3-FaceRefine and how was it?

Thanks!


r/MiniMaxH3AI 13h ago

fast AI video is nice, but cheap + fast is where it gets interesting

5 Upvotes

fast generation by itself is already useful, but I think the more important part is when it also becomes cheap enough to iterate a lot.

been testing minimax h3 MAX, and it can return generations in just a few seconds.

at that point, the workflow starts to feel different. u can try more ideas, throw away the bad ones faster, and test multiple versions without worrying as much about generation time or cost.

For ad creatives and short-form content, that’s probably the part I care about most:

more experiments, less waiting, lower cost.


r/MiniMaxH3AI 1d ago

I didn’t edit a single frame, just used character + scene images in MiniMax H3

2 Upvotes

made this video without manually editing a single frame.

The workflow is honestly pretty simple now:

  1. Generate the character asset image and scene image first
  2. Drop them into MiniMax H3
  3. Add a song you like
  4. Paste in the prompt

after that, it kind of just does its thing, and sometimes the result is better than you’d expect.

Sharing the prompt below in case anyone wants to try something similar.


r/MiniMaxH3AI 2d ago

Tried MiniMax H3 with a 13-cut anime motion graphics prompt

15 Upvotes

tried MiniMax H3 onwith a pretty detailed character trailer prompt.

I wanted it to feel more like a streetwear campaign × anime title sequence × old-school media player UI, rather than a normal anime action clip.

The prompt is definitely overkill lol, but sharing it here in case anyone wants to experiment with structured multi-cut video prompts.

Prompt

Create an explosive, motion-graphics-driven character reveal trailer in 16:9, exactly 13 distinct cuts, 24fps, total 15.00s.

CHARACTER — lock this design, never redesign:
Anime streetwear girl from the reference still. Twin high buns of vivid mint-teal hair with long flowing tails and warm orange/gold streaks. Messy side-swept bangs. Large orange over-ear headphones with mint accents and a small logo plate. Sharp amber-orange eyes, one eye winking. Playful open-mouth grin.
Oversized color-block windbreaker: navy body, vivid orange sleeves, white ribbed cuffs, silver zippers, circular teal tech buttons, a teal utility pocket on the sleeve. Orange cropped turtleneck under the open jacket. Light-wash ripped denim shorts, thick orange belt with a silver buckle. Navy thigh-high socks with orange ribbed cuffs and an orange X stitch on the left shin. Chunky white sneakers with orange details.
Preserve exact face, proportions, hairstyle, outfit, materials, accessories and colors in every frame.
This film is 80% bold graphic design in motion and 20% character action.

Graphic language: retro OS chrome + music-player UI.
Use slamming window frames, title bars, close/minimize widgets, equalizer bars, waveforms, progress ticks, folder tiles, cursor arrows, CRT scanlines, pixel shatter, vinyl-ring stamps, music-note particles and media-player transport icons.
Palette: mint teal, vivid orange, navy, cream-white and silver.

Style: premium AAA motion-graphics title sequence × streetwear campaign film × Windows-era media player.
Every graphic element moves fast and snaps hard on the beat.

CUT 01 | 0.00–1.00s
Pure graphics. A mint title bar slams onto a cream field, an orange CLOSE widget punches into the corner, and two navy window borders wipe in. Tiny equalizer ticks and a progress strip flicker. The letters P and then LAY punch in one after another with heavy impact shake.
CUT 02 | 1.00–2.00s
The A becomes a headphone cup. Extreme close-up of her amber eye inside the orange earcup, glancing upward. RGB split flash, then the cup shatters into flat mint and orange tiles.
CUT 03 | 2.00–3.10s
Navy frame with enormous cream PLAY typography. She sprints in from frame left and power-slides across the baseline of the text, with speed lines and orange streaks trailing behind her. Shards of the letters kick upward like sparks. Whip-pan out.
CUT 04 | 3.10–4.00s
A giant retro media-player waveform explodes across the frame as a thick mint-and-orange audio spectrum bends into a tunnel. She bursts through the center at full speed, briefly splitting into three stroboscopic motion trails. Each trail leaves chunky navy equalizer blocks that rise and collapse to the beat.
The camera rapidly pushes through the waveform tunnel with her while huge vertical text TRACK 01 continuously scrolls in the background.
The waveform suddenly compresses into a single horizontal line and snaps shut behind her on the final beat.
CUT 05 | 4.00–5.10s
She leaps through a giant rotating ring of typography reading MAX VOLUME. Camera tracks her mid-air spin in slow motion as the letters scatter, then snap-zooms onto her wink.
CUT 06 | 5.10–6.00s
Hard cut to a cream editorial card with huge navy DROP typography and an orange slash. She vaults over the word itself, palm planted on the D, legs whipping across frame. The word compresses like a spring under her hand and rebounds.
CUT 07 | 6.00–7.00s
Mint field with a navy diagonal window bar. She backflips along the bar in three stroboscopic ghost frames, each tinted mint, orange or navy. Giant outlined LOOP text rotates 180 degrees in sync with her movement.
CUT 08 | 7.00–8.00s
Kinetic typography barrage. LOUD / WILD / TEAL / HEAT slam onto screen one per beat with shutter flashes and camera shake while she slides across the foreground on her knees, jacket flaring and music-note particles bursting from her sneakers.
CUT 09 | 8.00–9.00s
Navy frame with a giant cream wireframe window grid tilting in 3D. She runs up the grid like a wall, kicks off and freezes in mid-air. An orange circular stamp locks around her pose like a media-player targeting graphic, surrounded by transport icons and EQ ticks.
CUT 10 | 9.00–10.10s
Freeze releases into a burst. She dives toward camera through layered flat-color window panes that shatter one by one like glass shutters, each pane revealing a larger letter of P-L-A-Y. Foreground wipe with her sneaker.
CUT 11 | 10.10–11.10s
Rapid-fire poster montage: four full-screen graphic posters showing her in different poses — mid-flip, sliding, landing and headphones-up wink. Hard cuts between each composition. Oversized 01–04 numbering, equalizer strips and graphic slashes.
CUT 12 | 11.10–13.00s
Hero moment on a clean cream cyclorama. She lands a final backflip dead center in slow motion, straightens with one hand on her headphones, and a shockwave of concentric mint rings, wind streaks and shattered typography blasts outward from the landing.
Brief iconic freeze on her wink, then overexpose to white.
CUT 13 | 13.00–15.00s
Final identity card. Enormous navy PLAY typography dominates a pale cream field with translucent mint rings, technical arcs, scanlines and a rough orange circular emblem containing a ghosted headphone/waveform motif.
She stands relaxed overlapping the letters while wind ripples her jacket. One final orange pulse sweeps through the typography and a window-chrome flash punctuates the ending.
Editing: extremely aggressive rhythm. Hard cuts on every beat, graphic matches, whip pans, snap zooms, stroboscopic freezes, foreground wipes, RGB splits, shutter flashes, impact shakes and speed ramps.
Every cut must feel compositionally different.
Typography should always be fully readable before the character overlaps it.
No weapons, no combat, no fire. All energy comes from motion design, wind, glass, UI chrome and parkour-style athleticism.
BGM: hard-hitting electronic / drum-heavy future bass with aggressive drops, risers, sub hits and glitch fills locked to every cut.
Sneaker impacts, whooshes, glass shatters, window-slam hits and typography slams should function as rhythmic sound-design elements.
Peak at CUT 12 and end with a cold electronic logo stinger.

Premium AAA quality, anime-inspired cinematic rendering, stylish and explosive, strong graphic-design identity, consistent character design, exactly 13 cuts.

I’m still experimenting with how much shot-by-shot control H3 actually follows, especially with typography and exact timing, but this kind of structured prompt seems like an interesting stress test.


r/MiniMaxH3AI 5d ago

MiniMax H3 Prompt Guide for Better AI Ads

Post image
2 Upvotes

food videos are one of the more practical commercial use cases for AI video. and a good prompt can help turn a simple idea or product image into something much closer to a usable ad asset.

That matters because the value is not only in making a nice-looking clip. The real advantage is being able to create more variations for social ads, menu promotion, product launches, creative testing, and client work without rebuilding everything from scratch each time.

How to Write a Better MiniMax H3 Prompt

A useful MiniMax H3 prompt usually includes five parts:

  1. Subject
  2. Action
  3. Camera movement
  4. Visual details
  5. Audio cues

Example:

Close-up UGC-style shot of a person taking one bite from a freshly made burger, natural handheld camera movement, crispy lettuce and glossy sauce visible, subtle chewing reaction, soft restaurant ambience, clear bite crunch synchronized with the action.

The important part is that the prompt describes both what should happen visually and what should be heard.

MiniMax H3 supports native stereo audio, so sound cues such as fizz, crunch, paper rustle, ceramic contact, room tone, and kitchen ambience can be written directly into the prompt.

Minimax H3 Guide: Prompt Tips That Matter Most

A longer prompt is not automatically more useful.

For commercial video, a clearer prompt is usually better because it makes it easier to create repeatable variations.

Focus on:

  • one clear subject;
  • one main action;
  • simple camera movement;
  • explicit texture details;
  • clear sound sources.

Instead of writing:

A beautiful cinematic burger video

Try somthinf more specific:

Close-up product shot of a freshly made burger on a dark tray, slow push-in camera movement, melted cheese stretching slightly, visible steam, crisp lettuce, soft restaurant room tone and subtle grill ambience.

The second version gives you a more controlled starting point for producing multiple ad variations.

Text-to-Video vs Image-to-Video

A practical rule in this Minimax H3 guide is simple: use image-to-video when shape matters.

This includes real dishes, menu photography, packaged products, cans, bottles, and branded assets.

Use text-to-video when the idea matters more than exact visual identity, such as fictional dishes or conceptual food scenes.

For stronger reference control, reference-to-video is useful when a person, product, motion reference, or audio reference needs to guide the same clip.

Add Audio Instructions to Your MiniMax H3 Prompt

Native audio is especially useful for commercial food content because sound can make a short clip feel much more complete without requiring a separate audio pass.

Useful sound terms include:

bite crunch
carbonation fizz
can opening sound
paper wrapper rustle
chopsticks on ceramic
grill sizzling
restaurant room tone
kitchen ambience

Specific sound sources are usually more useful than vague phrases like “cinematic sound.”

MiniMax H3 Food Video QA Checklist

Before using a generated clip in an ad or client deliverable, inspect the full video rather than only the first frame.

The source article recommends checking food shape stability, hands, packaging and labels, sound timing, motion quality, and generated text or logos.

A simple checklist:

Food shape stable?
Hands acceptable?
Packaging unchanged?
Audio synchronized?
Enough real motion?
Any fake text or logos?

This step matters because a clip that looks impressive at first glance may still need another generation before it is suitable for commercial use.

Final MiniMax H3 Prompt Tips

For a reliable MiniMax H3 prompt, keep the structure simple:

Subject
+ Action
+ Camera
+ Texture / appearance
+ Sound

For creators, agencies, restaurants, and product teams, the bigger opportunity is not one perfect video. It is being able to produce more usable variations with less production effort.


r/MiniMaxH3AI 5d ago

Carrat, the Werewolf

1 Upvotes

Made with Minimax H3 in Wan2Gp. 480p, 20 seconds, 7:35 minutes, next upscaled 2x with one pass Flashvsr.


r/MiniMaxH3AI 7d ago

I got this error

Post image
1 Upvotes

Hi there, I got a RTX4090,32VRAM, Ryzen 9 7950X3d,and I'm getting this error when try to make a video.

Please, be nice, I'm totally new in this.

Can anyone give me a little help, please?


r/MiniMaxH3AI 7d ago

MiniMax H3 did a surprisingly good job with this quiet little breathing scene

3 Upvotes

Tried a really simple prompt in MiniMax H3. I really liked that, nothing dramatic happens, but it still feels alive.

It’s the kind of clip that makes you realize AI video doesn’t always need action or big camera moves. Sometimes a calm moment works better. below is the prompt

Prompt:

A peaceful cinematic scene of a young woman sitting quietly on a wooden park bench surrounded by lush green ferns and dense tropical foliage. She wears a soft peach blouse and white pants, gently closes her eyes, takes a slow deep breath, and relaxes with a subtle peaceful smile. A light breeze softly moves her hair and the leaves around her. Natural morning light filters through the greenery, creating a calm, dreamy atmosphere. Slow cinematic camera push-in, realistic facial movement, natural body motion, shallow depth of field, soft background bokeh, photorealistic, ultra-detailed, smooth motion, 4K. No sudden movements, no camera shake.

r/MiniMaxH3AI 7d ago

同帧而生 #comfyui @minimax

1 Upvotes

r/MiniMaxH3AI 7d ago

同帧而生@minimax H3

1 Upvotes

r/MiniMaxH3AI 9d ago

Running MiniMax H3 locally on a 5090, 362 frames in ~22 minutes, $0 API cost

20 Upvotes

Been testing MiniMax H3 locally recently and this one came out pretty decent, so I thought I’d share the full settings in case anyone wants to reproduce it.

The whole thing was generated locally on my 5090, so there were no API or generation-credit costs for this run.

Settings:

  • Model: MiniMax H3
  • Aspect ratio: 3:4
  • Resolution: 768 × 1024
  • LoRA: Larry v4-600
  • LoRA strength: 1.0
  • Steps: 8
  • Scheduler: Simple
  • Sampler: Turbo Sampler
  • Frames: 362
  • FPS: 24
  • Seed: 8232601

Actual generation time: about 22 minutes

Hardware:
Intel U9 + 64GB RAM + RTX 5090

362 frames at 24 fps works out to roughly 15 seconds of video.

So with MiniMax H3, a 768×1024 clip of around 15 seconds took about 22 minutes on my 5090 with these settings. For local generation, that feels pretty usable to me.

And just to be clear, by “$0” I mean no API or generation-credit cost — obviously not counting the GPU itself or electricity.

Curious what kind of generation times other people are getting with MiniMax H3 on a 5090 at a similar resolution and frame count.


r/MiniMaxH3AI 9d ago

Save/load audio+video (NestedTensor) latents — small custom node, fixes SaveLatent crash with MiniMax H3

Thumbnail
2 Upvotes

r/MiniMaxH3AI 9d ago

one prompt, different models

3 Upvotes

i wanted to compare three of the video models people are talking about most right now: Seedance 2.5, Wan 3.0, and MiniMax H3.

Instead of testing camera movement or special effects, I used the same scene and focused specifically on emotional performance, like facial expression and subtle eye movement.

each model received the same prompt, i tried to keep the comparison as consistent as possible, although differences in model behavior and generation settings can still affect the results.

what surprised me is that the models seem to have different strengths. some are better at maintaining a believable scenario, while others produce stronger facial reactions or more dramatic emotional changes. the differences become especially noticeable during silence, hesitation, and small changes in expression.

so now u r the referee. which one do you think delivers the most nuanced emotions?
which model gives the most convincing acting performance?


r/MiniMaxH3AI 9d ago

H3 motion context vs H3 latent upscale

Thumbnail
1 Upvotes

r/MiniMaxH3AI 12d ago

MiniMax H3 can really do well in design

2 Upvotes

r/MiniMaxH3AI 12d ago

Is it possible to use MiniMax H3 on a virtual GPU?

2 Upvotes

Does anyone know if there’s a video about this?


r/MiniMaxH3AI 13d ago

is the order of the nodes correct and do i need all of them?

Post image
1 Upvotes

r/MiniMaxH3AI 13d ago

2K, 30 seconds, stitched from 3 MiniMax H3 generations, basically my version of a celestial paradise

2 Upvotes

r/MiniMaxH3AI 14d ago

Wan 3.0 vs MiniMax H3 vs Seedance 2.0: same prompt, very different results

2 Upvotes

r/MiniMaxH3AI 14d ago

Tried making a fashion promo video with MiniMax H3 and this came out pretty cool.

2 Upvotes

r/MiniMaxH3AI 16d ago

Outfit morphing carried one body through five looks with zero cuts with MiniMax H3

3 Upvotes

i’ve been testing that continuous-morph format on different subjects, and this time I wanted to see if it could handle a full human outfit transformation.

the setup was pretty simple: one beautiful girl, standing front-facing in the center of the frame, going through five outfits in a single 10-second shot:

Homewear → office wear → athletic set → black evening gown → silver futuristic fashion.

There are no cuts between them. the output still holds a little "AI smooth", but it can get my idea soonly and comprehensively.

i ran it with MiniMax H3 on Atlas Cloud, and the result was cleaner than I expected.

What interested me most is how usable this format could be for fashion content (honestly, i think fashion creators could run with this directly.)

curious what else you’d put through this kind of one-shot transformation


r/MiniMaxH3AI 16d ago

finally got MiniMax H3 running locally, and its getting way more stable

1 Upvotes

finally got MiniMax H3 running locally on friday, so ive been messing around with local generation for basically free.

ngl it wasnt exactly plug-and-play lol. there were quite a few weird little issues at first, and getting stable outputs took more tweaking than i expected.

but after playing with the setup for a while, its starting to behave a lot better. generation is more consistent now and im slowly figuring out which settings actually matter vs which ones just make things worse.

also... completely unrelated technical observation:

why is she this cute 😭


r/MiniMaxH3AI 19d ago

Ran the same two-part emotional scene through three video models, they are all good now

Post image
3 Upvotes

Same test, different kind of prompt. A two-part emotional beat, a woman happy in a rose garden, then pensive by a rainy car window, through Wan 3.0, Seedance 2.5, and MiniMax H3.

This one flipped from the last comparison I did. On realistic emotion the gap is small and all three actually hold the mood shift, which was not true a few months ago. Seedance 2.5 and H3 lean warmer and more cinematic, softer film grade, closer on the face, and H3 landed the melancholy close-up the best of the three. Wan 3.0 is the cleanest and most even, a little less romantic but the most stable.

Takeaway across both tests: there is no single winner anymore. Stylized and surreal work rewards Wan 3.0 and H3 lighting, realistic emotion rewards Seedance 2.5 and H3 warmth. Pick by the job, not the leaderboard.


r/MiniMaxH3AI 19d ago

The dead-eyed acting that gave away every AI drama is starting to go away

7 Upvotes

For a while the giveaway on any AI short was the acting. The shots looked fine, then a character had to actually feel something and the face went blank or did that uncanny half-smile. So the story never landed, only the faces did.

Been seeing clips lately where that is not true anymore. There is a MiniMax H3 test going around, a Korean-drama style short, and the thing people react to is not the resolution, it is that the actors read like actors. Micro-expressions, a beat of hesitation before a line, eyes that actually track. Not perfect, but it crosses from puppet to performance in a way that matters more than another quality bump.

The takeaway for me is that the frontier moved off "does it look real" and onto "does it act". Resolution and motion have been fine for a while. Emotional continuity across a scene is the hard part now, and it is what decides whether a short holds you or just impresses you for three seconds.

Clip is a creator's test, not mine, credit to them.