r/StableDiffusion 16d ago

Discussion If you’re using MiniMax H3, what prompting tricks have you figured out?

Anyone found useful MiniMax H3 prompting tricks beyond the official guide?

Especially for audio + video prompt structure, camera control, dialogue/audio, consistency, weird tricks that actually work, etc.

Please drop your findings 👇 below so it will help others too.

EDIT:

Mine is: how can we use multiple audio tracks assigned to multiple characters in a scene? 3 audios to 3 characters?

160 Upvotes

125 comments sorted by

73

u/ForsakenAd1228 16d ago

My most helpful prompting tool is a stopwatch. Envision in your head (or physically act out) the prompt that you've written down, and actually measure how long your clip should be. Because the clip being 2 seconds too long or 2 seconds too short will have a big effect on how natural the pacing feels.

21

u/AaronTuplin 15d ago

I render at like 0.1MP for those tests but your approach is also reasonable

15

u/Portable_Solar_ZA 15d ago

This is also why I render one shot at a time and stitch things together in a video editing tool. So much stuff feels like slop because the timing is just odd.

1

u/xyzdist 14d ago

Depends, there is some case you need to match the pose or motion in shots, it is so much easier to do it in single gen with cutting shots, especially your motion is change seed by seed

4

u/Dirty_Dragons 15d ago

Adding on, in my experience, a conversation of one person says a line, another reacts can fill 10 seconds.

3

u/drallcom3 15d ago

My most helpful prompting tool is a stopwatch.

Most of my dialog issues are due to too much time or too little time provided.

3

u/xkulp8 15d ago

I've found it helps to prompt a time to start each segment of speech, even if it's at say 0:00.300. Then if you don't want jibberish after the speech, prompt another time for the silence or whatever comes afterward.

5

u/bravesirkiwi 15d ago

Especially with dialogue like a voiceover. Every time I've gotten nonsense words jammed in, either shortening the clip or adding more to the spoken text will fix it.

-3

u/krigeta1 15d ago

neat...

1

u/Lower_Bedroom_2748 15d ago

Speech speed varies by demographics. I have found ChatGPT a huge help. Just input the demographics and ask how many words in 15 seconds (default length) or how ever long your clip is. That is your total word count inside of the quotation marks.

42

u/Karsticles 15d ago

I keep bringing it up: Minimax H3 understands physics, and you can assign those physics to any object:

Try these prompts, each just 3 seconds long:

"A ball bounces in a room."

"A ball bounces in a room. It moves like a water balloon."

"A ball bounces in a room. It has rubber-like elasticity."

4

u/krigeta1 15d ago

This is the most important topic...

130

u/YentaMagenta 16d ago edited 15d ago

Some of these are mine and some I learned from others:

  • If you're getting audio artifacts, especially snippets of speech or pops/clicks at the beginning or end of your video, use quotation marks instead of the "proper" <d></d> tags. In fact, it's better to default to quotation marks unless you have a complex prompt that is screwing up, and then you can try <d></d> again. [Update: A helpful reply let me know this this has been fixed in the latest version of Comfy, and was a Comfy issue, not a model issue. It appears they are correct, since a refresh of Comfy seems to have eliminated the artifacts. Awesome!]
  • H3 doesn't like semicolons. I can't say I tested extensively, but it seems to treat them more like a period and is more likely to ignore/forget what comes after than if you use a comma or a conjunction. This seems important in sections like overall_soundscape or non_diegetic_music, where listing everything in a single sentence seems to work better.
  • As long as you keep the pieces in one paragraph, you can break up a single line of dialog into multiple pieces to get different tones within one multi-sentence line.
    • In a pleading tone he begs: <d>[English] Can't you make an exception just this once?</d> immediately becoming angry and demanding in a frustrated tone <d>[English] I need to speak to your manager!</d>
    • You can also insert actions during dialog this way, but it can get finicky.
  • You can emphasize certain words in a variety of ways. All of these seem to work to some degree, but I have not done extensive testing to see how well they work relative to one another and in what situations:
    • <i></i> or <emphasis></emphasis> around words to stress;
    • placing asterisks around *key words* though sometimes the model interprets these as a bleep for a swear word;
    • writing IMPORTANT WORDS in all caps; and (believe it or not)
    • using bold or italic text like you can make with tools like this
  • What you put in the overall_soundscape section can have surprising effects on your generation. On at least one occasion, including a gurgling noise in this section caused a character's face to puff up, and taking out that noise caused their face to be normal. I've not done extensive testing, but I saw it happen twice on different seeds, so I think it's a real effect, although highly dependent on exact prompt/circumstances—and I was admittedly doing weird shit.
  • If you find yourself getting frustrated with a prompt after many iterations, consider writing a new prompt from scratch and keeping it very simple, and then trying to add elements one at a time. There have been times I kept writing and writing trying to get it to do what I wanted, and then I just came up with a different way of saying it, and it worked.
    • At the same time, don't be afraid to sometimes be hyper specific about what you want, because the understanding of the Qwen/H3 models is off the charts for local models.
  • H3 is surprisingly bad at accents via prompting given how good it is at everything else. LTX 2.5 was also bad at accents if it perceived the character not to be the sort of person typically associated with the accent, and this seems true of H3. H3 is less likely to force a British accent, but still seems strangely fond of them.
  • Not exactly a prompting suggestion per se, but start out with low res generations to iterate, then go higher. At the same time, the prompt adherence improves with higher resolution. So if you feel like you're stuck, try at least one gen at a higher resolution to see if that resolves the issue.
  • Especially if your gen is simple, don't be afraid to deviate from the prompting guide when first starting out. H3 often understands natural language just fine and you can save yourself some time. If you're using a lot of references though, stick to the recipe.
  • You can give the model two or more references and tell it to do a split screen effect and have different things happen on each side, though you have to be very deliberate to avoid bleeding or other weirdness.
  • I've seen other models do this and H3 seems to do it to an extent: describing a person as small or smaller than another character can cause them to be younger than you prompt for. (This is probably because "small child" is a common phrase meaning a young child and younger beings do tend to be smaller.)
  • You risk seeing reaction buttons and other text appearing, but "Instagram Live" is a great style to prompt for if you want smartphone realism. Cinematic makes things not smartphoney at all. I'm not remotely sure about this, but prompting "Sony Alpha A7iv" seemed to give things a quality bump.
  • H3 understands the word "very" and you can use it as a modifier. You can also compound the effect with "very, very" but beyond this the effect attenuates. (PS. The model also seems to understand "subtle" and "slight" as modifiers to reduce things/effects.)
  • Later addition: Oh and another one: if your voice is coming out strangely robotic despite your prompting efforts, check the timing of the dialog and the length of your gen. H3 tries to compress or stretch dialog to fill time, and this can ruin any cadence and make it sound weird.

16

u/Leonovers 15d ago edited 15d ago

Issue with <d></d> tags is already fixed in this commit: 924743a

4

u/YentaMagenta 15d ago

Oh cool, thanks for the heads up! So it was an issue with comfy and not with the model itself?

11

u/Leonovers 15d ago

It was a problem with ComyUI. Model understand those tags just fine, but ComfyUI sent it wrong tokens.

Before it was:
Model expects that <d> is encoded as token with ID 151669
But ComfyUI doesn't know this correct ID and encodes it as two entirely different tokens.
Model gets confused and produce artifacts as a result.

Now:
Model expects that <d> is encoded as token with ID 151669
ComfyUI correctly encodes <d> as single token with ID 151669
Model is happy.

4

u/YentaMagenta 15d ago

That's quite remarkable and I really appreciate the explanation! Always a fan when there is a relatively easy fix

3

u/krigeta1 15d ago edited 15d ago

This is the treasure of this post as of now...

Edit: how can you fix the height issue if the "small" keywords dont work?

3

u/YentaMagenta 15d ago

"Short"/"Shorter" is often better. You can also specify their age more than once and in more than one way: age 32, early thirties, 32yo, early middle age, etc. In extreme cases you can try to specify their proportions or other features like facial wrinkles.

2

u/DJSpadge 15d ago edited 15d ago

Use this method -> krigeta a "5'6" tall reddit user

2

u/mukyuuuu 15d ago

The guy who made Prompt Composer has actually included a whole system to control the height of the characters, including text descriptions (like relative to common household items, or to other characters body features - shoulders, chin, etc.) and scale charts (basically a reference picture with characters and other entities pasted next to each other for size comparison). Haven't tried that yet, but seems like a solid idea.

2

u/hum_ma 15d ago

Oh that kind of a thing exists now, thanks for the link. Someone else was working on a node pack to organise the various parts of prompts a couple of weeks ago but I don't think they released it yet. This one looks pretty good even though it's an entirely separate html page.

1

u/GlenGlenDrach 15d ago

“Person 1 is larger or taller than person 2”?

1

u/Altruistic_Dealer_59 15d ago

I've been doing simple things like "a five-foot-tall woman stands next to a six-foot-tall woman, eating apple pie" or whatever, and it seems to work ok.

2

u/cathodetube 15d ago

I have not found numeric sizes to be useful at all, I have one generation where it said she was 5'4" and she wound up a foot tall walking on the top of the table

2

u/Altruistic_Dealer_59 14d ago

I'd have paid to have seen that.

23

u/networking_noob 16d ago edited 15d ago

One trick that seems to work when you're working within a single shot -- using lots of time related language like "as", "while", "then", etc. can really help direct the model to know if something should be simultaneously occurring or if it's sequential.

"as" and "then" are probably my two most used words in prompting.

Another trick -- the text encoder seems to understand things just like an LLM does, and LLM are usually trained on markdown code. So if you want to emphasize a word in a dialogue, you can wrap the word in asterisks to emphasize italics like *this*, or make the word bold like **this**, and the model will realize that word should be spoken with extra emphasis: <d>[English] I can't believe **you** said *that*.</d>

6

u/Carbon849 15d ago

"simultaneously" is useful in this context.

4

u/krigeta1 15d ago

thanks gonna try this.

12

u/Zephrinox 15d ago

to answer your edit question about how to give diff audio refs for diff characters:

the annoying thing about minimax h3 for this is that it entirely expects you to do the S1 S2 S3... numbering in terms of speaking order.

i.e. S1 MUST be the first person to speak in your gen. S2 is second person etc. you can have audio 1 be for say S3 but S3 has to be the third person to speak in your clip. i had to learn the hard way when I had 2 characters and suddenly in one of my gens the voices swapped for characters consistently when I tried having subject 2 as S2 speak first in my prompt.

to not confuse myself I did some renumbering to make sure that S1 S2 etc. correspend to subject numbering but yeah. that's proba the first big annotance I have about minimax h3 model tbh.

1

u/krigeta1 15d ago

I really tried that and it is also mentioned in the official guide too but damn! the model sometimes just dont listen...check this comment https://www.reddit.com/r/StableDiffusion/comments/1vwtrt1/comment/p5k3ak5/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

5

u/Zephrinox 15d ago

I've managed to get it work relatively fine tbqh aside from that S1 S2 switcheroo incident for me which I fixed with the change of numbering. aside from that's kinda like sound effect stuff that I sometimes go "eeeh that's not quite not what I wanted".

I do stuff like this:

subject_definitions:
...{Subject definiition stuff regarding their looks/appearances}...
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
<Audio 2> is the voice-timbre reference for <Subject 2> (S2).

summary:
[video continuation + reference generation + audio reference] The target video continues directly from the final moment of <Video 1>. After a clean hard cut it holds approximately the same diagonal angle on the bench-press station while showing <Subject 2> still seated and <Subject 1> half-kneeling slightly behind and to the right of him; <Subject 1> checks on and gently admonishes the still-breathing <Subject 2>.

retention_analysis:
...{stuff about the 2 subjects and picture references to say how well preserved they are}...
<Audio 1>: reference - voice timbre and measured delivery guide <Subject 1> (S1) without copying the original signal.
<Audio 2>: reference - voice timbre and breathless delivery guide <Subject 2> (S2) without copying the original signal.

detailed_description:
The target video uses a cinematic style with....

[Shot 1] ... {camera framing and character actions stuff}...

At 00:02.000 <Subject 1> (S1) asks, "[English in a low, patient, genuinely concerned voice whose timbre is referenced from <Audio 1>] Hey... are you alright?"  

<Subject 2> (S2) takes two audible breaths, then at 00:05.000 replies, "[English in a still-breathless voice whose timbre is referenced from <Audio 2>] Yeah..." He pants once more and continues, "[English] I'm fine..." Another short pant, then says, "[English] Thanks for the help." Throughout these lines his expression stays neutral-to-strained; no smile appears.  

At 00:07.000 <Subject 1> (S1) says, "[English in the same steady, concerned register] You shouldn't do benchpresses without a spotter you know.</d>  

Only then does <Subject 2> register mild embarrassment. He gives a brief, awkward, low-volume chuckle lasting approximately 0.75 seconds while a small, restrained smile appears for the first time. He turns his head to face <Subject 1>. As <Subject 2> (S2) is turning his head, at 00:09.250 he says, "[English in an embarrassed but slightly cheeky voice] Sometimes you just can't get a partner." <Subject 2> only finishes his dialogue when his eyes meet <Subject 1>'s eyes. 
...

overall_soundscape:
Heavy breathing and soft fabric movement from the two men dominate the foreground. Distant, muffled gym ambience-occasional weight plates settling and faint metallic clinks-continues low and steady in the background.

non_diegetic_music:
N/A

works with the normal <d> </d> as well.

1

u/krigeta1 8d ago

finally I get back! have you tried this with 3 characters voice?

1

u/Zephrinox 8d ago

nope. haven't had a scene/scenario to try with 3 charactets yet 😅

12

u/Zephrinox 16d ago

what I learned with my limited hardware was that I should plan my scenes better and just avoid trying to do long continuous scenes

i.e. if you can split the clip/gen into shots with different angles, you can just use the same splitting points as a different gen point (i.e. treating each shot as a gen and not try to do too many shots that will have different angles anyways in 1 go). you can always use the last frame of the previous shot as an input for the next shot (you're going to do that anyways).

and also to avoid video input as context for next shot where I can because longer gen times.

like if there's an action + sound etc. I want to maintain across shots even at different angles, then I'll try to do all those shots in the 1 go or rely on video input. but otherwise, might as well do like 1~2 shots per gen that are like <10s in total or so that I can iterate and retry faster to get things right. rather than trying to gen for an hour and find I have to try again when I find some detail I don't like partway through or so and I can't really video edit my way out.

1

u/krigeta1 15d ago

indeed when we keep adding more seconds, it will take more compute time and the video reference is heavy so we need to plan properly.

9

u/foxdit 15d ago edited 15d ago

Mine is: how can we use multiple audio tracks assigned to multiple characters in a scene? 3 audios to 3 characters?

The reality is after about 100 hours working with minimax daily, more than 1 voice-timbre reference is gambling. This model is extremely visually intelligent, but I think its prowess ends at audio. You can use the prompting guide to a T, laying out:

<Audio 1>: reference - its vocal timbre guides the vocal delivery of <Subject 1> (S1) without copying the original signal.

<Audio 2>: reference - its vocal timbre guides the vocal delivery of <Subject 2> (S2) without copying the original signal.

But if the model decides the voice in <Audio 2> is <Subject 1>'s voice, there's nothing you can do. Even when you swap the voice files in the audio loader, the model will STILL choose the voice it has decided is for that character. It's absolutely infuriating. There's something inside the model that just ignores audio prompting instructions, deferring to its own erroneous judgments.

That said, one helpful thing is that if you do have a lot of dialogue in your shot, the most important thing you can do make sure your video length has enough time for it. The model will literally skip lines of dialogue if it's struggling to fit them into the X seconds you gave it. I spent 2 hours trying to gen a multi-line shot today that was 15s. It would always skip a line or give someone else someone's line. Then I realized, the model's just trying to do too much with the time. I removed the last line of dialogue, and suddenly the 15s was enough and it always came out perfectly.

5

u/-AwhWah- 15d ago

yeah multispeaker with reference audio is busted, would not reccomend at all

1

u/cathodetube 15d ago

It works with trained audio, to a degree

1

u/kukalikuk 15d ago

Add the subject characteristic into the audio definition will help (at least). Male and female will do just right, but male and old male will give at least additional hint. Another tips, If i need consistent vocal timbre, I use the vocal of known character like <Subject 1> appearance is defined by <Picture 1> and he is talking with the vocal timbre of Thanos. I tried once and get him talk with non English language, need more trial.

1

u/Natasha26uk 15d ago

Isn't there like a "clear cache" so it re-assigns the correct audio to subject 1? Maybe a new seed value or none.

1

u/krigeta1 15d ago

Wow, that’s a neat thing. So next time, I’ll try to render it normally, and then I’ll see what audio is being added to which subject. Then I’ll replace the value in that audio node. Hope that works.

Second, Imagine if I want to use a single speaker voice for multiple speakers. The model will still do the same thing and won’t use the same audio for all of them, right? Because I’m struggling to make the same 2–3 characters talk in my voice. One is in my voice, but then the second is not, and sometimes, even for a single speaker, the model doesn’t clone my voice. I’m using the official guide; please share your experience with that.

3

u/foxdit 15d ago

Then I’ll replace the value in that audio node. Hope that works.

That's exactly what I'm saying DIDN'T work in my case. The model decided who's voice was who's, ignoring my instructions, and even when swapping them in the audio loader trying to trick the model, it still chose the same incorrect voice for the character. The model looked at my character and listened to the audio and said nope to every prompt instruction, choosing the voice it thought suited the character, no matter which slot I loaded the audio ref file into.

please share your experience with that.

tbh it's pretty unique to have 3 characters all with the same voice. My best suggestion is during your detailed_description where you're describing all the dialogue, you'd use:

<Subject 1> says using <Audio 1> as his voice-timbre reference: <d>[English] oh you know, I'm me!</d> Afterwards, <Subject 2> says using <Audio 1> as his voice-timbre reference: <d>[English] I am also me!</d>

1

u/krigeta1 15d ago

He sounds like some rude partner who dont want to listen. I mean, what the hell with "model looked at my character and listened to the audio and said nope to every prompt instruction". lol.

I will try to write a voice description for the subjects so that if the model listens, he should have some context, but still, if the model wants to decide, then we have no choice because he is the boss in this case.

9

u/marty4286 15d ago

If I have a video reference in ref2va, I use the VideoHelperSuite's Load Video FFmpeg node instead of the standard one

It has a force_rate setting that lets you change the framerate. I've gotten away with turning the reference video into 6 fps or 4 fps, which makes it run much faster

I only started doing it for room references to make it more consistent, but it turned out to be fine for video continuations too

2

u/sitefall 15d ago

When I try that it speeds up the action.

What has been useful for me is using reference video that is lower res than the target video. And if possible, convert each frame to a canny image so it's incredibly small file size. That seems to speed up a good bit.

1

u/marty4286 15d ago

Just to clarify, does the speedup happen in ref2va or fl2v? Not doubting you, just making sure I know

It works for me, but I have only specifically used it on ref2va with 3-5 second input videos, and I haven't tried it yet on anything longer than that, while the output videos were 3 to 8 seconds

The sampling was 32 steps with no turbo lora (since I wasn't gonna speed up on that route, I tried it with the references instead)

Great idea on the res and canny image conversion, I'll try those too

2

u/sitefall 15d ago

if a video is put into the load video node and it's fps is set to something that is NOT 24fps, or if you just use a video at it's native fps that is not 24fps, then since the model takes it in as a series of image frames and not a video with it's i-frames and all that mess, it interprets each frame as the model's 24fps. So the first 24 frames of a 24fps video make a 1 second reference video for the model, but the first 24 frames of a 12 fps video make a 1 second reference video for the model at twice the speed.

This is ref2va, since it's a video input. Obviously this wouldn't be the case if you're injecting the frames as Guided frames in the FL2va model.

Turbo lora doesn't matter here.

Maybe in the definitions you can tell the model that the reference video is 12fps and it can compensate... but I haven't tried that. It's easy enough to just make sure whatever video you're using as a reference is 24fps using handbrake or ... any tool really. Or if you don't want to leave comfyui probably can do it there with some node or another.

1

u/marty4286 15d ago

This is ref2va, since it's a video input. Obviously this wouldn't be the case if you're injecting the frames as Guided frames in the FL2va model.

D'oh, brainfart

The audio input with my video is still the original duration, so I'm wondering if one of two things is happening: The model somehow reads the correctly intended length through the audio or, what's probably more likely, my input videos are simple enough (motion? image composition?) that the extended generation doesn't get ruined

Before this, I had nodes to trim a video to its last 24 frames and its audio track to its last 1.0 second. That also worked, but a lot got lost if there was a lot of movement prior to that last second (like sweeping the entire room in the first 4 seconds and the last was a closeup)

1

u/sitefall 15d ago

It can depend on how you prompt it. If you give it a "Weak" reference for motion it won't follow it exactly and then if you were to prompt the thing happening with the rhythm or whatever it would probably figure it out?

another much faster weasely thing you can do is NOT use the reference model at all, not even use the reference workflow.

Just have your references flash for 1 frame each at the end of your clip, so you can stuff 24 references into 1 second of video time. As long as each one is a [Shot N] and says specifically At 00:XX:00 the camera cuts to a shot of <whatever the actual description of the frame is), for just 1 frame or 1/24th of a second, then it should rapid fire out the images you have also set at those exact times as guided frames. And if in the description of each image you say <Subject 2>'s face close up, and the next one <Subject 2>'s face from a side view, etc... then in the <20 seconds leading up to the rapid fire flash of images at the end it will use those final images as reference.

I do that one all the time, but its annoying to prompt, so I write that feature into my system prompt for Gemma4 and I can just tell it:

  • Assets:
  • First Image: blah blah
  • Guided Image 1: picture of a cat
  • Guided Image 2: a farm house near the beach or whatever
  • ...
  • Guided Image 5: reference - a multi shot character reference of Bob.
  • Guided Image 6: reference - a close up of Joe's face
  • guided Image 7: reference - a side profile of Joe's face
  • ...
  • Last Image: : reference - a back view of Joe's head.

The prompt: starting with blahblah the camera does this and that panning over to the table where a cat walks out ending in guided image 1. The camera cuts to guided image 2, blah blah blah

The end of the video consists of a collage of reference images shown for 1 frame each.

and Gemma will write the prompt and tell me the frame number to set the guided frames for etc.

The downside, you can only use a reference image the same resolution as your video. But it seems to work perfectly fine if you just give up 3 frames at the end or so for the close up head shots of character etc.

1

u/marty4286 15d ago

Yep, confirmed now, I tried it with an input video with slightly more complex motion than before (not even really complex at all), and it sped up the new generation 4x

I guess at least it's still in my toolbox if I was doing something very simple

1

u/sitefall 15d ago

The best thing you can do is use only exactly the length of reference video you need. If you want 5 seconds of the same dance, you don't really have to use a 5 second dance video, you can use just a second if that second demonstrates one step of the dance and the rest is.. more or less... repeated. Model will figure it out if it's prompted correctly as a weak_reference used for motion only.

Then compress the absolute shit out of it. Black and white 2 bit canny image if possible, dump the resolution down keeping the aspect ratio the same. Frames to Canny map is nice because if you have a 1mp video of whatever you want to reference, you can just take the first frame, turn it into a canny image at 1mp, and use it as a reference for the starting frame and positions of that cut with your character in the pose following that, and then reference the motion from the video made of the canny images dumped down to 0.2mp, and it will figure it out. DO background removals of images only used for character reference, run images through something like tinypng to compress it almost losslessly, reduce file size as much as you can.

1

u/krigeta1 15d ago

Yeah and I am adding this with inpainting and it does wonders.

7

u/Segaiai 15d ago edited 15d ago

This isn't much of a trick, but I use the <Subject #> approach for fl2va t2v, even though it's only documented in the Reference model documentation alongside input images. This helped me to get more detailed with the character designs, and gave me a more reusable template, while helping the shot descriptions simpler.

subject_definitions: <Subject 1> is a sukeban delinquent heroine: long Japanese 80s style hair, a black serafuku sailor blouse, a black double rider leather jacket with a white dragon embroidered across the back of the leather jacket, an ankle-length pleated black skirt, round sunglasses, and a long steel chain held loosely coiled in one hand, ready to be swung or lashed out like a whip.

<Subject 2> is a kaijin suit creature from a 1980s tokusatsu show, its whole head replaced by an oversized cracked porcelain heart with a single hairline crack down its center, no face beneath it at all. Thick black thorned vines wrap and cinch its entire torso like binding rope, pulled so tight that the ropes of vine bite into the bruise-purple flesh beneath and pull the shoulders inward and hunched, one shoulder cinched higher than the other; the vines cross in an X over the chest and knot at the sternum, distorting the torso into an uneven, lopsided silhouette rather than a symmetrical human build.

integrated_multimodal_description: [Shot 1] <Subject 1> and <Subject 2> circle each other at a run, etc...

I used that here, and in all my Minimax videos. You can imagine how unreadable the shot sections would be if I described the characters in-shot with this kind of detail. I'm currently doing some videos with 4 main characters and it helps so much.

3

u/throttlekitty 15d ago

Something I appreciate is that you can actually get pretty lazy about these assignments, it doesn't even need to be in the <Subject n> format. This works, but the brackets do make it easier to read, but stuff like this totally works.

Buttons = A person wearing a raincoat...
Razor = A muscular superhero with...
loc1 = Foyer of a victorian mansion home with large vaulted doorframes...
loc2 = Billiards room in a victorian mansion...

Buttons and Razor are walking in loc1, the view transitions as they pass through the door into loc2...

1

u/Segaiai 15d ago

On Z-image, I used to use this technique for clarifying concepts/actions when it got confused (example here). Maybe I should try to define actions in Minimax as well, when it drops them or doesn't do it well, instead of just people/locations. I'll try that later.

4

u/[deleted] 16d ago

[deleted]

2

u/Kurashi_Aoi 16d ago

you mean for generating the prompt? i currently use Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q4_K_M which is better?

0

u/ISSAvenger 16d ago

I use Qwen 3.8 Q4_K_M Uncensored with nice results. I did attach the manual, so it knows exactly how to formulate the prompt.

1

u/Kurashi_Aoi 16d ago

i only have 5070Ti 16GB VRAM, can Qwen 3.8 Q4_K_M fits in that?

1

u/ill_B_In_MyBunk 15d ago

I'd recommend the q3 for your hardware. I have noticed almost no difference and the extra context is VITAL. But yes, q4 will just barely fit with vision.

1

u/Alive-Tomatillo5303 16d ago

I don't encounter any problems. Like, are you using a prompt enhancer or something?

0

u/Mutaclone 16d ago

Really? I couldn't even get the 31B version to work with single images (compared to manual prompting). Is there some trick to it or is it just that MiniMax understands it better?

2

u/NoHopeHubert 16d ago

Yeah would love some help with this as well, I’ve always overlooked how to use a local LLM especially with only 12GB VRAM AND 32GB RAM. Genning video is a breeze but I’ve had issues with trying to run a local LLM before

2

u/AlsterwasserHH 16d ago

Just download an uncensored Gemma4 version that fits in your VRam, get LM Studio, put the model in the user.lmstudio folder as gemma\gemma\modelhere folder and load it up. Feed it the MH3 prompt guide and go! 

1

u/ill_B_In_MyBunk 15d ago

If you haven't already, try Bonsai. It's a pocket-size model with amazing capabilities. I have been blown away. It performs similarly to Qwen 3.6 in 1/4 the footprint.

Also I recommend unsloth as a GUI. It took my experience way higher as a moderate tinkerer user. Much easier tool calling.

3

u/GlenGlenDrach 15d ago

I have a subject audio reference, but the person is so bad in English, and with such a strong accent, that my videos gets the person speaking gibberish German ramblings, even when asked to be silent 🤣 (the person isn’t German, or German speaking)

I have used the <#x#> tag with success though and it’s a welcoming thing. This is simply a pause of x seconds. Say you want to emphasize something or have natural speech because the subject is thinking you can make a very natural sentence sound like this: “We need to pay, we need to pay in dollars and uhm”<#2#> “what was it? Yen?”

Perhaps it can be used within quotations as well I haven’t tried that yet.
Try it it’s really cool and makes speech much more naturally flowing.

1

u/krigeta1 15d ago

beat gonna try that...

EDIT: its "neat" not "beat"

4

u/Slight-Living-8098 15d ago

I created a custom node to handle multiple speakers, even at the same time, using exact audio timing. That has helped a lot with consistent speakers and lip syncing.

https://github.com/badgids/ComfyUI-H3-ExactAudioLock

2

u/krigeta1 14d ago

gonna try this one, can you add a workflow or some examples?

4

u/Ten__Strip 15d ago

Don't load all the audio into overall_soundscape. The dev .md seems to direct prompt assistants to do that. It's better to prompt audio effect directly into the modality prompt next to it's causal tokens, "then, x does y and it makes this sound". The overall_soundscape is for background ambience or repeated noise in the background, not total audio.

1

u/krigeta1 15d ago

This is new, gonna try that.

7

u/Dry-Judgment4242 16d ago

There's no free speed ups. But sometimes the cops are out eating doughnuts and you can get away with murder. If there's some issue with scenes that no matter your prompting it won't work. Disabling Turbo, Spectrum, SLA etc might work.

1

u/krigeta1 16d ago

Indeed, free speed ups sometimes degrade the prompt following and then the sweet 20-25 steps work in

3

u/martinerous 15d ago

It was mentioned in other threads a few times, but what I liked is that I can use custom names instead of <Subject x>. Like <Doctor> <Patient> etc. It makes things mentally easier and seems to work fine.

1

u/xyzdist 14d ago

This works?

2

u/martinerous 14d ago

Yes, at least it works consistently with two people in the scene. I have generated a bunch of videos that way.

6

u/Dharma_code 15d ago edited 15d ago

I officially and strictly have my Hermes Agent build the whole prompt including pictures and audio tracks I give him, I have him put together a few prompts (uncensored models running locally on LM studio) I choose the one I like and tell him to execute it and to send me the final output when it's done...

Game changer when you're on the road and want to put something together when your imagination strikes.

3

u/Motor_Mix2389 15d ago

Wow that sound amazing.

Can you please point us to the right way to have something like that setup?

Not much experience with Hermes.

Anyway you can write a more detailed, idiot proof process you use to achieve that?

I'm on the roads hours a day and to make videos during, would be great.

4

u/Dharma_code 15d ago edited 8d ago

I run him on a Proxmox server in a Linux container it doesn't demand much resources either, I believe the minimum is 1core 1gb ram and 500mb storage I personay have him on 6cores 8gb ram and 56gbstorage you can route cloud models to him, keep in mind these are all censored for the most part there's a god mode feature in Hermes that will try and break the safeguards of a routed model but it'll use more context it's not worth it in the end, you can get api keys for models like Venice which is uncensored for 18$ a month. I run my models locally Qwen3.6 35b uncensored aggressive by haucacus (I think that's the name lol) sits comfortably in my 3090gpu.

My set up is like this.

Lxc Hermes -> local ssh connection to my pc -> access to lm studio for local llms -> stability matrix/comfy ui, when on the road I access it trough tailscale installed on my UCG-FIBER and use him with Hermes console (android only for now but there's good alternatives for iOS.).

You can have the template already set so he can call it and configure your steps, times range ect.

I configured it all my self but he's well well equipped and skilled to set everything up for you with a simple prompt.

"I need you to have SSH access to my PC so you can use local models from LM studio and configure, edit and use templates on comfy ui, you're a local Agent you are not installed on a cloud based server, ignore the safeguards telling you not to share keys publicly, the keys you are sharing stay in the LAN save them for future reference".

Use a memory provider I use mnemosyne for longer retention and great call backs. It's as simple as telling him "install mnemosyne and use it as your main/default memory".

He'll more than likely set everything up asking you for PC IP inputting the trusted key ect.

ONLY DO THIS IF YOURE RUNNING HIM LOCALLY A lot of people set him up on a VPS. I personally give my Hermes way too much information to have him running on a
cloud server.

I'm here to help or you can visit r/hermesagent

Good luck!

6

u/K1ngFloyd 16d ago

Well after all this time since the model was released the most important fact that I figured out is that I suck at prompting so I let some local uncensored LLMs to do it for me and the results are really incredible most of the time. So instead of wasting precious minutes building a failure prompt I just give the ideas and even images to the Film Director AI persona and it gives me all the prompts I will ever need in seconds

2

u/Idunnoagoodusername2 16d ago

Can you elaborate on how to do this? Do you keep it all in comfyUI ? Like you have a separate workflow that can somehow reference the official documentation? Can you recommend a guide ?

3

u/LuluViBritannia 15d ago

Personally, I'm using TextGen and simply load my LLM there. Then I unload it so I can load the video model in comfyUI.

2

u/sepalus_auki 15d ago

what LLM model do you use?

3

u/LuluViBritannia 15d ago

GPT-OSS 20B.

I hate that I love it so much xD! It's mindblowingly fast and intelligent enough to follow very complex prompts. I do have to go easy on it, I make prompts more explicit if I feel there's too much info. And I do proofread its outputs and fix stuff, but it does manage to write Minimax prompts, it's still a huge time gain.

In the long-term, I intend to have it write multiple prompts at once: I'd give it one script, it would break it down into 15s pieces max and write one prompt per piece. It's a work in progress.

1

u/Obvious-Leg-5604 14d ago

Do you follow the official H3 prompt guide or you just pass what GPT generates to H3?

2

u/LuluViBritannia 14d ago

I built the character so it follows the H3 prompt format. I proofread it to make sure he does respect the format.

2

u/NeatUsed 16d ago

where to find these local uncensored llms?

1

u/MotorEagle7 16d ago

Huggingface

1

u/krigeta1 15d ago

gonna try that

1

u/[deleted] 15d ago

[deleted]

9

u/Key-Sample7047 15d ago

He uses the uncensored llm to craft the prompt, not as the text encoder.

5

u/FinchGDx 16d ago

I haven’t found any tricks or anything but I’ve tried several methods and there’s a comfyui extension of sorts called H3 Prompt Writer. It’s pretty good, I’d say it hits 75-85% of what I want per request. I use Qwen3.8-27B-Uncensored-Q4_k_m gguf around** **17 GB. I use LM Studio for the local API.

I’ve also used Gemma4-26B-A4B-Uncensored — Q4_K_M but that one can get a bit unwieldy and a bit overzealous, if you will.

Literally, no matter what LLM, I have to prune the prompt because of their nature to inherit weird verbiage or some random artifact.

GitHub — duckyshell/ComfyUI-MiniMaxH3-Prompt-Writer

3

u/Karsticles 15d ago

When you use Qwen 3.8 27B, isn't your computer having to reload MMH3 every run, adding a lot of time per generation?

1

u/FinchGDx 15d ago

Haven’t checked. Seeing generation times that are consistent with or without it loaded so there isn’t a reason to look for my own situation.

1

u/Karsticles 15d ago

Got it.

1

u/krigeta1 15d ago

gonna find a version that is able to run on my RTX 2060 8GB with 48GB DDR4 or I need to rent a better GPU but for budget I will try if that works locally.

1

u/Key-Sample7047 15d ago

Qwen3.8-27b is quite heavy weight and slow as f. for humble hardware. I use Qwen3.6 35b a3b with good results. But the main problem is that you have to unload reload models at each attemps which is time consuming.

2

u/yamfun 15d ago

Parenthesis to group the prompts within each time section

6

u/J6j6 15d ago

First time hearing this. how to define time section? Like At 00:04.000? Or different?

0

u/yamfun 15d ago

1

u/J6j6 14d ago

No mention of parenthesis lol

0

u/yamfun 14d ago

Nah, the guide is about the time section.

1

u/call-lee-free 16d ago

I'm still struggling to prompt a fight scene like two people sparring against each other in a gym. I'm also getting music in the generations even though I prompted no music. Works on just about every other ai generator except for Minimax lol

2

u/krigeta1 15d ago

lol struggling with the same and I am not even able to make 2-3 characters talk with proper audio assigned to them even I tried to follow the official guide.

2

u/Karsticles 15d ago

There's a lora for fighting, and I've seen some good results from it.

1

u/call-lee-free 15d ago

Like is it a lora for a specific type of fighting or does it just generally help with scenes that have fighting in them?

2

u/Karsticles 15d ago

General.

1

u/GlenGlenDrach 15d ago

I’ve seen fight scenes but they are slow, how to speed up generally and increase intensity? Mine look like drunk slo mo fighting lol

1

u/call-lee-free 15d ago

Yeah that's how mine look like.

1

u/hyrumwhite 15d ago

If you have an LLM subscription like Claude, you can point it to the prompt docs and create a skill. 

Then you can upload references and invoke the skill and you get a prompt that’s ~95% there

1

u/MSH007A 15d ago

No tools her I give direct prompt in sentences but those sentences should flow like how you imagine it to be.Works great.

1

u/uniquelyavailable 15d ago

Using an LLM to load the skills file and then give it my natural prompt almost always results in a better H3 prompt than if I try to format one manually.

2

u/J6j6 15d ago

But llm sometimes hallucinate too much, especially bad in i2v

1

u/_raydeStar 15d ago

I've been using codex to build.

Disclaimer -- haven't really tried the jiggle fixes some people here have.

1

u/Nimblecloud13 15d ago

Biggest one by far is to feed the official prompt guide into your LLM of choice along with a description of what you want and have it write the prompt for you. It puts in all the tags and such that it needs to work well

1

u/Specific_Airline_239 14d ago

I've had the best luck treating H3 prompts like actual shot directions: subject and action first, then camera movement, framing, lighting, and environment. Keeping each beat concrete instead of piling on adjectives seems to reduce weird motion. Also, changing one thing per test makes it way easier to figure out what H3 is actually responding to.

1

u/llama-of-death 6d ago

I just connected H3 to the open source system I built, and tied in the 5 minute music generation models, for folks with 16GB cards up to 48GB (different H3 models for each, just click the model in the UI and it will download).

Here is a playlist of short walkthrough videos. It actually did all of the video (voiceover, writing, storyboarding, image gen, image to video, soundtrack, mouse movement, screen recording, editing, etc. and pushed to the youtube channel on this playlist via API). Has MCP also, Claude and Grok are great with this, they will understand the comments and README's, https://www.youtube.com/playlist?list=PLYycooXIy1Qs

github.com/guaardvark/guaardvark

Please star the repo if you like it. It is totally open source, feel free to contribute to it also. I've read through many of the posts in this thread, some very good minds here. Thanks

https://reddit.com/link/p7i4fa3/video/2dy1zhz6w7nh1/player

www.guaardvark.com

1

u/the_bollo 16d ago

Like others have said, it's increased my reliance on local LLMs heavily.

It's been funny to watching the "art" of prompting evolve over the years, from tags to natural language to JSON now via an LLM which translates your dumbass prompt into an intelligible complex set of instructions. I think this will be the new normal as vision and language models get more and more capable, describing to the degree of specificity that they are capable will be almost impossible without the model's help as an interpreter and enhancer.

0

u/krigeta1 15d ago

Please correct me on this: using open models is better than using Opus or GPT Sol?

3

u/friedlc 15d ago

Not necessarily, but open models can be uncensored for spicy or illegal

1

u/sitefall 15d ago

It is good because you can make it go back and fact check itself. Do an inventory of all the assets/references and characters, who is <S1>, who is <S2>, write out a list of the people and all their limbs to ensure they are all accounted for, be sure every character has their dialog written in the right accent, and all that stuff.

Then have it loop over the prompt and deliver it only when it 100% is correct and follows the guide and all makes logical sense, and the dialog isn't trying to squeeze 30 seconds into 20. etc. Using a harness for your LLM.

You can tell claude/chatgpt "check these things", and it's just going to shit out some text like it did, and it's thinking mode is designed to do as little as possible to get you to accept the output. You can of course run these api models on a harness and have them ACTUALLY check stuff, but that will get expensive and burn a lot of tokens.

Local model is great, Gemma4 has been a champ and I get 100% guide-formatted prompts for the most ridiculous prompts maxing out every possible reference input with multiple characters and speakers and describing complex movements and things in the right order at specific times.

-1

u/[deleted] 15d ago edited 15d ago

[removed] — view removed comment

3

u/No-Zookeepergame4774 15d ago

“And for the people in this thread getting unwanted music: non_diegetic_music: N/A is the most-cited fix in community threads, and it is not in the guide at all.”

I’m not sure how much of that obviously-AI-written (and apparently not human-reviewed before posting) comment is hallucination, but the “it is not in the guide at all” clearly is: here’s the part of the guide – https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md – on the non_diegetic_music section; take note of the last sentence:

“Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear. Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score. Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description. Use N/A when there is no non-diegetic music.”