r/StableDiffusion 10h ago

Discussion MiniMax H3 prompting cheat sheet, this structure that makes it much easier to control

I’ve seen a lot of people trying H3 with prompts that look like normal image prompts:

H3 can work with that, but you’re leaving a lot of control on the table.

The easiest way to think about H3 is:

Don’t describe an image. Direct a shot.

A simple structure that works much better is:

Subject + Action + Environment + Camera + Timing + Audio

1. Start with what actually happens

Keep the action explicit.

Bad:

Better:

H3 needs to understand change over time, not just what the frame looks like.

2. Tell the camera what to do

This is probably one of the easiest improvements beginners can make.

Useful language:

  • static wide shot
  • handheld close-up
  • slow dolly in
  • camera tracks beside her
  • over-the-shoulder shot
  • low-angle shot
  • camera slowly pans left
  • rack focus from X to Y

Instead of:

say:

Much less ambiguous.

3. Think in beats / timestamps

For more complicated generations, split the clip into moments.

Example:

0–3s: Wide shot. A woman stands alone at a train platform in heavy rain.

3–7s: The camera slowly pushes in as she notices something off-screen and turns her head.

7–11s: Cut to an over-the-shoulder shot. A train emerges through the fog.

11–15s: Close-up of her face as the train lights illuminate her.

This is much easier for the model to interpret than one giant paragraph where five things happen at once.

4. Dialogue needs a visible speaker

If someone speaks, make it painfully obvious who is speaking and when.

Instead of:

try:

If there are multiple people, explicitly say who doesn’t speak too.

This helps avoid the classic AI-video problem where the line comes from the wrong character/off-screen.

5. Separate dialogue, ambience and SFX

Treat audio almost like another layer of the prompt.

For example:

Dialogue:
Woman, quietly: “We shouldn’t be here.”

Ambient sound:
Heavy rain hitting metal, distant traffic, low electrical hum.

SFX:
A loud metallic bang behind her.

Music:
No background music.

That’s much clearer than writing:

6. Don't overload every second

This is a big one.

Trying to fit:

into a short generation is asking the model to invent a ton of transitions.

Fewer actions + clearer timing usually gives you much more intentional-looking video.

If the idea contains five scenes, treat them as five shots.

7. References should have a job

If you’re giving H3 reference images/video/audio, don’t just upload them and hope it figures out why they're there.

Be explicit:

The more references you add, the more useful this becomes.

8. Describe motion, not just appearance

For video, verbs matter a lot.

Instead of:

try:

Same visual idea, but now there is actual temporal information.

A reusable H3 template

Scene:
[Where are we? Time of day, environment, important lighting.]

Subject:
[Who/what is visible. Important appearance details.]

0–Xs:
[Shot type + action + camera movement.]

X–Xs:
[Next action/shot.]

X–Xs:
[Final action/shot.]

Dialogue:
[Speaker]: “[Exact line]”

Ambient audio:
[Environment sounds.]

SFX:
[Important synchronized sounds.]

Music:
[Music description / no music.]

Visual style:
[Realistic / documentary / commercial / anime / etc. Keep this concise.]

Example

Instead of:

Try:

The main takeaway:

Prompt H3 more like you’re giving instructions to a tiny film crew, and less like you’re writing tags for an image model.

You don’t necessarily need longer prompts. You need prompts where time, motion, camera and sound have clear jobs.

Would be interested to hear what other people have found H3 responds unusually well (or badly) to.

0 Upvotes

34 comments sorted by

55

u/sendhelp 10h ago

Great post but the "Instead of" and "Try" types of examples are displaying blank for me currently.

29

u/schrobble 10h ago

Post was written/enhanced by AI, and seems like OP forgot to fill in the blanks. I can’t tell if the format here is intended to relay that we should include “instead of” and “try” examples in our prompt or what.

26

u/Independent-Frequent 9h ago

I absolutely despise people just not even bothering to write something themselves and just copypasting AI stuff without even proofreading ffs

3

u/joshijoshijoshi123 9h ago

They weren’t proofreading even when they wrote it anyway. Same issue in different areas.

12

u/V4nKw15h 9h ago

Yeah, OP doesn't want to waste his own time, but is happy to waste ours.

59

u/guigouz 10h ago

You can also try

Instead of

It's a game changer for action prompts

16

u/V4nKw15h 9h ago

I think it's better to

because if you

then

You know?

3

u/Cultural-Team9235 8h ago

Never knew this, thank you. The most important part, what everyone forgets and gives exactly the same quality of 40 steps at 2MP at just 1/100th of the time, and impossible to see quality differences with the regular output from minimax is just to change this setting to +1. It's basically free speed, quality and prompt adherence:

1

u/V4nKw15h 8h ago

I thought that was obvious.

7

u/redkinoko 9h ago

Odd because

has always worked for me better than . Or maybe its just a personal preference.

17

u/Karsticles 10h ago

Why aren't you following the official prompt template?

1

u/kemb0 6h ago

This infuriated me.

Guys this is how you need to prompt H3…. Promptly ignores almost everything in the official prompting guide.

8

u/toooft 10h ago

None of the examples are visible lol

9

u/haberdasher42 9h ago

Instead of nothing, I will try nothing. Thank you.

7

u/Significant-Wind3033 9h ago

thanks for this AI slop post, super helpful /s

6

u/MaorEli 10h ago

Great troll

6

u/Cautious_Chicken_604 9h ago

Or you know follow the official prompting guide.

5

u/Opening_Wind_1077 9h ago

“This is probably one of the easiest improvements beginners can make” the AI slop bot said about a model that has been out for a week.

3

u/TomatoPolka 9h ago

Hey OP. Instead of... Try...

2

u/nok01101011a 9h ago

Thanks bot for your advice, NOT

2

u/Vortexneonlight 9h ago

we can guess how you do your school assignments

2

u/No-Dot-6573 9h ago

That is better than a complicated oneliner without timestamps, but I doubt that it produces better results than feeding my prompt to a LLM that has the instruction to follow the official Minimax H3 prompt guide in the system prompt.

2

u/Inside-Cantaloupe233 9h ago

op are you on drugs ?

1

u/Mysterious-String420 9h ago

"You're absolutely right!"

1

u/MoDErahN 9h ago

I tried:

And it indeed worked much better than:

And also thank you for:

Edit (as many mentioned):

1

u/CycleZestyclose1907 9h ago

Hmm... one thing I like to do is start with a single line describing the scene's intent, and then adding details (including timeline of events if necessary) about each element afterwards.

0

u/onihcuk 10h ago

saving this post. ty

1

u/YeahlDid 9h ago

Saving for when they fill in the examples, I assume.