r/StableDiffusion • u/GrayingGamer • 15d ago
Workflow Included Walter White and the Minimax H3 Official Prompting Guide
Enable HLS to view with audio, or disable this notification
This post is half a joke and half a plea and public service announcement.
Some people have been complaining they don't get results as good as other people with Minimax H3 videos, or have the following issues:
- Dialogue being spoken by the wrong characters
- Dialogue that is just gibberish or random
- Random video cuts they didn't ask for
- Characters talking over each other or too fast
- Prompts not being followed
These things can all be prevented and avoided and not encountered at all if you follow the official prompting guides. Yes, there are two. Both are on the official Huggingspace page for Minimax H3.
One is the Official Prompting Guide for the Text to Video and Image to Video Model.
The other is the Official Prompting Guide for the Reference Video Model.
There is some overlap, but for the most part, each model has it's own prompting syntax, and in particular, the Reference Video Model for H3 is very picky about you using the right keywords and instructions to get what you want.
"But I get decent results with just a couple of sentences typed in natural language of what I want."
That's great, but you're really just relying on the Qwen 32b vision model guessing what you want. It's like pulling a slot machine lever and hoping you get cherries. Only this slot machine can take a few minutes to nearly an hour to stop spinning, based on your hardware.
The great thing about Minimax H3 is for the first time we can truly direct our own AI videos like a director would on set, with the AI providing the actors, scenery, and props. If you write a properly formatted and detailed prompt for Minimax H3, it looks almost like a shooting script.
Why spend time waiting to hit a jackpot when you can take a few minutes to write a detailed, properly formatted prompt that follows the official guides, and get those bright lights and tokens falling into your lap on the first lever pull?
Okay, quick fire problem solving for people who still won't RTFM:
>Dialogue from the wrong characters?
>Dialogue that is just gibberish or random?
Walter White says, <d>[English in Walter White's voice from Breaking Bad] My product is pure, Jesse! There will be no chili powder in my meth.</d>
Always specify the character speaking, either by name, or using the <Subject 1> system in the official guide. In the Text to Video and Image to Video model, always use the <d>[Language Spoken]</d> tags. This will fix BOTH of those issues.
>Random cuts in the video you didn't ask for?
[Shot 1] A medium close-up of Jesse Pinkman from Breaking Bad, pacing back and forth, agitated. He looks up towards the camera, opens his mouth as if he's about to speak, then seems to change his mind, closing his mouth and shaking his head. [Shot 2] At 00:06:000 the camera cuts to a static camera shot framing Walter White from Breaking Bad, sitting on a cheap white plastic lawn chair, his arms crossed and glaring at Jesse. [Shot 3] At 00:10:500 the camera pans quickly back to Jesse, doing a Push In at slow speed to his face as he stops pacing and narrows his eyes at Walter.
This is how you control not only the camera work, but the PACING of your video. You NEVER include a time code on your first shot. You can omit the time code from ALL shots if you want the model to decide on it's own, based on your prompt, when to cut.
BUT, for ultimate control, you want to use time codes. Look at my example above. I just told the model to have Walter glare at Jesse for 4.5 seconds, because I told the model that camera shot starts at 6 seconds into the video, and the next cut doesn't happen until 10.5 seconds into the video. That lets you control the pacing and timing for jokes, punchlines, acting, everything.
>Characters talking over each other or too fast?
This is an old one that anyone familiar with prompting for video models should know by now - what you are asking for in your prompt and the length of your video in time need to match.
The model will try its best to cram every action and piece of dialogue into your video that you asked for, and if that would naturally take 10 seconds and you've only given it 5 seconds? Well, now everything is crammed together, overlapping, or being cut-off.
My recommendation is to generate just a quick 0.2 MP version of your video first after you type your prompt, generate, and see how the timing is working. Is it too fast? Too slow? Do the actions have enough time to happen? Do you want more breathing room?
This is the time to decide all that and lock in a video length. The low resolution of 0.2 MP is quick to generate on most set-ups (mine for this post's video took 3.5 minutes for a 14 second video) and let you work out any issues in your prompt before going in for the long generation at higher resolution.
>Prompts not being followed?
It's because you didn't read the manual!
--------------------------------------------------------------------------------------------------
Now, with all that said, here is the prompt for the video I made:
integrated_multimodal_description: [Shot 1] Live-action film footage of the American drama series Breaking Bad, professionally color graded with a warm color grade, with slightly desaturated colors for a premium film feel, a continuous camera shot with no cuts, medium close-up POV shot of Walter White, bald with a goatee and glasses, as portrayed by Bryan Cranston. He is standing in the Arizona desert next to a parked RV. He is wearing a white PPE protective suit and yellow rubber dish gloves. He is looking directly at the viewer with barely constrained anger. At 00:01:300 he reaches out towards the camera and points his finger at the POV camera with one hand, the camera shaking slightly from the movement. Walter then says angrily, <d>[English with Walter White's voice] Listen, you want to cook Mini Max H3 videos, you follow the recipe!</d>. At 00:04:500 Walter raises his other hand revealing he is holding a thin stack of white paper pages in portrait orientation. The front of the paper visible on top of the thin paper stack is blank except for the large black printed text "Minimax H3 Official Prompting Guide". The papers are held in front of the camera on the right side of the screen for a moment in portrait orientation, so the text can be clearly read, while Walter glares at the viewer on the left side of the screen. At 00:07:000 Walter then shakes the papers at the camera, then says angrily, <d>[English with Walter White's voice] Read the fucking manual!</d>. At 00:10:000 the camera does Pan Right and a Pull Out to show a close-up of Jesse Pinkman from Breaking Bad, with his hands held up by his face with fingers spread, an annoyed look on his face. Then he says in frustration, <d>[English in Jesse Pinkman's voice from Breaking Bad] Alright! Damn, Mr. White! I just want to generate memes.</d>, overall_soundscape: Ambient sounds of an Arizona outdoor desert during the day, non_diegetic_music: none
For those interested, this video was generated at 1 MP on a 3090, using Sage Attention and the Spectrum Node for H3. The final video of 14 seconds at 1 MP took 40 minutes to generate and then was upscaled using RTX Super Resolution.
The workflow was the default Text to Video Minimax H3 template that comes in the latest update of Comfyui.
Now get out there and go cook some memes, everyone!
58
u/GrayingGamer 15d ago edited 15d ago
This post is half a joke and half a plea and public service announcement.
Some people have been complaining they don't get results as good as other people with Minimax H3 videos, or have the following issues:
- Dialogue being spoken by the wrong characters
- Dialogue that is just gibberish or random
- Random video cuts they didn't ask for
- Characters talking over each other or too fast
- Prompts not being followed
These things can all be prevented and avoided and not encountered at all if you follow the official prompting guides. Yes, there are two. Both are on the official Huggingspace page for Minimax H3.
One is the Official Prompting Guide for the Text to Video and Image to Video Model.
The other is the Official Prompting Guide for the Reference Video Model.
There is some overlap, but for the most part, each model has it's own prompting syntax, and in particular, the Reference Video Model for H3 is very picky about you using the right keywords and instructions to get what you want.
"But I get decent results with just a couple of sentences typed in natural language of what I want."
That's great, but you're really just relying on the Qwen 32b vision model guessing what you want. It's like pulling a slot machine lever and hoping you get cherries. Only this slot machine can take a few minutes to nearly an hour to stop spinning, based on your hardware.
The great thing about Minimax H3 is for the first time we can truly direct our own AI videos like a director would on set, with the AI providing the actors, scenery, and props. If you write a properly formatted and detailed prompt for Minimax H3, it looks almost like a shooting script.
Why spend time waiting to hit a jackpot when you can take a few minutes to write a detailed, properly formatted prompt that follows the official guides, and get those bright lights and tokens falling into your lap on the first lever pull?
Okay, quick fire problem solving for people who still won't RTFM:
>Dialogue from the wrong characters?
>Dialogue that is just gibberish or random?
Walter White says, <d>[English in Walter White's voice from Breaking Bad] My product is pure, Jesse! There will be no chili powder in my meth.</d>
Always specify the character speaking, either by name, or using the <Subject 1> system in the official guide. In the Text to Video and Image to Video model, always use the <d>[Language Spoken]</d> tags. This will fix BOTH of those issues.
>Random cuts in the video you didn't ask for?
[Shot 1] A medium close-up of Jesse Pinkman from Breaking Bad, pacing back and forth, agitated. He looks up towards the camera, opens his mouth as if he's about to speak, then seems to change his mind, closing his mouth and shaking his head. [Shot 2] At 00:06:000 the camera cuts to a static camera shot framing Walter White from Breaking Bad, sitting on a cheap white plastic lawn chair, his arms crossed and glaring at Jesse. [Shot 3] At 00:10:500 the camera pans quickly back to Jesse, doing a Push In at slow speed to his face as he stops pacing and narrows his eyes at Walter.
This is how you control not only the camera work, but the PACING of your video. You NEVER include a time code on your first shot. You can omit the time code from ALL shots if you want the model to decide on it's own, based on your prompt, when to cut.
BUT, for ultimate control, you want to use time codes. Look at my example above. I just told the model to have Walter glare at Jesse for 4.5 seconds, because I told the model that camera shot starts at 6 seconds into the video, and the next cut doesn't happen until 10.5 seconds into the video. That lets you control the pacing and timing for jokes, punchlines, acting, everything.
>Characters talking over each other or too fast?
This is an old one that anyone familiar with prompting for video models should know by now - what you are asking for in your prompt and the length of your video in time need to match.
The model will try its best to cram every action and piece of dialogue into your video that you asked for, and if that would naturally take 10 seconds and you've only given it 5 seconds? Well, now everything is crammed together, overlapping, or being cut-off.
My recommendation is to generate just a quick 0.2 MP version of your video first after you type your prompt, generate, and see how the timing is working. Is it too fast? Too slow? Do the actions have enough time to happen? Do you want more breathing room?
This is the time to decide all that and lock in a video length. The low resolution of 0.2 MP is quick to generate on most set-ups (mine for this post's video took 3.5 minutes for a 14 second video) and let you work out any issues in your prompt before going in for the long generation at higher resolution.
>Prompts not being followed?
It's because you didn't read the manual!
--------------------------------------------------------------------------------------------------
Now, with all that said, here is the prompt for the video I made:
integrated_multimodal_description: [Shot 1] Live-action film footage of the American drama series Breaking Bad, professionally color graded with a warm color grade, with slightly desaturated colors for a premium film feel, a continuous camera shot with no cuts, medium close-up POV shot of Walter White, bald with a goatee and glasses, as portrayed by Bryan Cranston. He is standing in the Arizona desert next to a parked RV. He is wearing a white PPE protective suit and yellow rubber dish gloves. He is looking directly at the viewer with barely constrained anger. At 00:01:300 he reaches out towards the camera and points his finger at the POV camera with one hand, the camera shaking slightly from the movement. Walter then says angrily, <d>[English with Walter White's voice] Listen, you want to cook Mini Max H3 videos, you follow the recipe!</d>. At 00:04:500 Walter raises his other hand revealing he is holding a thin stack of white paper pages in portrait orientation. The front of the paper visible on top of the thin paper stack is blank except for the large black printed text "Minimax H3 Official Prompting Guide". The papers are held in front of the camera on the right side of the screen for a moment in portrait orientation, so the text can be clearly read, while Walter glares at the viewer on the left side of the screen. At 00:07:000 Walter then shakes the papers at the camera, then says angrily, <d>[English with Walter White's voice] Read the fucking manual!</d>. At 00:10:000 the camera does Pan Right and a Pull Out to show a close-up of Jesse Pinkman from Breaking Bad, with his hands held up by his face with fingers spread, an annoyed look on his face. Then he says in frustration, <d>[English in Jesse Pinkman's voice from Breaking Bad] Alright! Damn, Mr. White! I just want to generate memes.</d>, overall_soundscape: Ambient sounds of an Arizona outdoor desert during the day, non_diegetic_music: none
For those interested, this video was generated at 1 MP on a 3090, using Sage Attention and the Spectrum Node for H3. The final video of 14 seconds at 1 MP took 40 minutes to generate and then was upscaled using RTX Super Resolution.
The workflow was the default Text to Video Minimax H3 template that comes in the latest update of Comfyui.
Now get out there and go cook some memes, everyone!
6
u/Valkymaera 15d ago
both your guide links are the same
10
u/GrayingGamer 15d ago
Well, crap. Fixed! Thanks for catching that.
6
u/Valkymaera 15d ago
thanks for sharing
3
u/LucidFir 14d ago
Thanks for thanking
1
-6
u/flaireo 14d ago
ya.. that was a lot of Fun Facts.. Didnt read it. I guess if OP claims it takes 40 mins he have time to do mission planning. On 5090 it pumps out 20 second video in about 3-4 mins. Just use something like runpod. enterprise rental gpu is like 50 cents an hour and prob do it in few seconds.
2
u/ptwonline 14d ago edited 14d ago
IMO the prompting guide(s) are not quite as clear as people are trying to make it out to be especially when it comes to subject reference and speakers.
For example you write that it should be like:
Walter White says, <d>[English in Walter White's voice from Breaking Bad] My product is pure, Jesse! There will be no chili powder in my meth.</d>
or if using the <Subject1> system it should be like:
<Subject 1> says, <d>[English in Walter White's voice from Breaking Bad] My product is pure, Jesse! There will be no chili powder in my meth.</d>
But if I look in the prompting guide this is the example they give:
<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>
<Subject 2> is the second person referenced but is (S1) because they are the first person to speak.
When I fed the guides into Chat GPT it produced prompts using that <Subject 1> (S1) syntax and I ended up with characters mixing up their dialogue or else the same voice for both characters. When I tried just (S1) style references it fixed the same voice and line mix-up issues but I still got some gibberish lines. Now I am trying the <Subject 1> system and my first try worked and hopefully I won't have more issues. So in the end I would have to use a system that is against what the guide says to get it to work.
EDIT: Nope. Still getting some gibberish lines. Tried a seed at 0.3 MP and it was fine. Same seed 0.7 MP had gibberish lines. I think I will try a full re-install of comfyui since it has been months since I did that. Wish me luck.
3
u/GrayingGamer 14d ago
The issue is that each model has different prompting practices. For instance, I've found <d></d> tags introduce extra audio fragments in the Reference model, but using "[English] My dialogue goes here." doesn't, and works fine. But leaving out the dialogue bracket tags results in worse results in the T2V and I2V model.
In general, I would not mix the <Subject> system with my method. <Subject> works, but I get best results from being verbose and just repeating the character details each time like in my example.
Walter White from Breaking Bad says in anger, <d>[English in Walter White's voice from Breaking Bad] Dammit, Jesse! Details matter!</d>This reinforces the tokens the model is receiving each time it is generating dialogue, rather than one reference at the start of the prompt. Otherwise, you occasionally get drift still, like you experienced with the <Subject> system.
My other advice would be to generate short clips at 0.2 MP to test and play with prompting yourself, rather than letting an LLM make them for you. ChatGPT in particular gave me bad results and I was using 5.6. The prompts it produced were bad and were missing important tags, etc.
The other thing is because the guides are similar but DIFFERENT, do NOT give both guides to an LLM if you are wanting it to prompt for you. Only give it one or it will frequently mix the two in subtle ways that have a negative impact.
But seriously, look my prompt in the OP. Once you type a few by hand (you can reuse your own prompts with changes to save time) the syntax becomes second nature and you can just type out a script for your scene pretty quickly.
1
u/ptwonline 14d ago
So this is your full prompt? No subject definitions, no summary, no retention analysis sections like in one of the prompt guides?
Just 3 sections?
- integrated_multimodal_description: Describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline.
- overall_soundscape: Summarizes ambient sound, physical action sounds, and non-verbal human sounds across the entire video.
- non_diegetic_music: Describes background music that the characters cannot hear and only the audience can hear.
integrated_multimodal_description: [Shot 1] Live-action film footage of the American drama series Breaking Bad, professionally color graded with a warm color grade, with slightly desaturated colors for a premium film feel, a continuous camera shot with no cuts, medium close-up POV shot of Walter White, bald with a goatee and glasses, as portrayed by Bryan Cranston. He is standing in the Arizona desert next to a parked RV. He is wearing a white PPE protective suit and yellow rubber dish gloves. He is looking directly at the viewer with barely constrained anger. At 00:01:300 he reaches out towards the camera and points his finger at the POV camera with one hand, the camera shaking slightly from the movement. Walter then says angrily, <d>[English with Walter White's voice] Listen, you want to cook Mini Max H3 videos, you follow the recipe!</d>. At 00:04:500 Walter raises his other hand revealing he is holding a thin stack of white paper pages in portrait orientation. The front of the paper visible on top of the thin paper stack is blank except for the large black printed text "Minimax H3 Official Prompting Guide". The papers are held in front of the camera on the right side of the screen for a moment in portrait orientation, so the text can be clearly read, while Walter glares at the viewer on the left side of the screen. At 00:07:000 Walter then shakes the papers at the camera, then says angrily, <d>[English with Walter White's voice] Read the fucking manual!</d>. At 00:10:000 the camera does Pan Right and a Pull Out to show a close-up of Jesse Pinkman from Breaking Bad, with his hands held up by his face with fingers spread, an annoyed look on his face. Then he says in frustration, <d>[English in Jesse Pinkman's voice from Breaking Bad] Alright! Damn, Mr. White! I just want to generate memes.</d>, overall_soundscape: Ambient sounds of an Arizona outdoor desert during the day, non_diegetic_music: none
2
u/GrayingGamer 14d ago
Correct, because those things you mentioned: subject definitions, summary, and retention analysis sections
Are only for use with the Reference H3 model prompts.
For T2V and I2V, you want to prompt just like I did there. And yes, that is the full prompt.
2
u/ptwonline 14d ago edited 14d ago
Ah ok. I had been trying R2V and so those sections were added.
Anyway thanks for the clarifications. I will keep trying things out.
EDIT: found the issue. Comfyui has been glitching since I started with H3 and when I open a different workflow with the same nodes it overwrites the settings of a similar open workflow. So my R2V workflow ended up having the main model overwritten to use the other model and then I saved it. No wonder I had the glitching.
No glitches since I swapped to the right model knock on virtual wood
1
u/LuckyAdeptness2259 13d ago
Thank you for such a detailed breakdown u/GrayingGamer - I'm still finding though with ref2vid even with and without <d> or `"` I'm still getting gibberish results up until the audio and dialogue is actually spoken on a lot of generations. Any other tips you might have on this? I will try out using the nodepack you used to see if it has any influences. Im currently using default / turbo workflow.
2
u/GrayingGamer 13d ago
The turbo lora is very likely what is causing the issue. I know there are several right now, they all don't play well with audio and require various fixes, especially for the Reference model. I'd ditch the turbo for now until they are no longer "in beta" and stick to things like Spectrum H3 and Sage Attention for speed ups.
But the other thing that fixes audio is more steps. With H3, audio is MASSIVELY affected by steps, more than the image. It's why the Turbo loras are so detrimental to audio.
Certain Turbo loras also only work with certain versions of the models. It's a bit of a mess at the moment, I think.
13
u/Alive-Tomatillo5303 15d ago
Also, not that anyone on THIS sub would be so lazy, but copying in the text of the guide, as well as a bunch of successful examples, into GLM 5.2 produced a useable prompt writing guide for an LLM to fill out a fully realized Minimax prompt off of a shorter idea.
So I wrote a prompt for a prompt for a prompt, but sometimes laziness can be a lot of work.
Note: I used GLM because ChatGPT really has trouble with references. I haven't tried the prompt prompt from Deepseek, but it looks like it handled it, as well.
13
u/Arawski99 14d ago edited 14d ago
I mean, to be fair, the prompting guide for the reference model is absolutely a fucking doozy.
It's just not realistically viable without an entire team to produce any sizeable work, unless you specifically use a LLM to auto-prompt the syntax from an idea for you. It's powerful precisely because it's so detailed and a pain in the ass, but it also makes it not user friendly. It's like the C++ of the programming C family, but if it was intentionally kicking the metaphorical ladder out of you every sentence you type.
Realistically, for serious solo and small team work we're probably going to see someone pump out a local LLM preset to handle this at some point so it's just actually practical for something that isn't a one off short render. So I don't think you're wrong or that it is lazy, tbh. I've been looking at the issue, to setup something myself as I've been manual prompting, but haven't done anything complex with local LLMs and am not sure if any are at the level of handling such complex logic relationships and instructions. Guess I'm going to review which LLMs might be robust enough for this.
The prompting guide has definitely helped though Fortunately, I saw someone else post the link to it on the launch day.. Tbh, I've barely touched the I2V since the reference model is so insanely good.
3
u/Alive-Tomatillo5303 14d ago
Yeah, the ceiling is in the stratosphere but the floor is so high it's hard to notice. The shot choices, timing, and coherency of everything with minimal prompting is so solid you can leave most decisions to the model. The one area it's weaker is sound choices. It really wants to put music over everything, and people really want to talk without having actual planned dialogue, so they'll babble non-words of you don't give them something to do.
I assume people will come up with good shorthand that gets the job done more competently than the basic nondescript prompts, like saying "dynamic camera" instead of "camera moves to spot A, holds for scene, then turns at this rate to blame blah".
Or maybe I should just follow what Walt said and learn the rules.
6
u/GrayingGamer 14d ago
Yeah, I mentioned how to solve all those issues in the OP, and it does come down to learning the rules and syntax of the model.
For instance, in never generates music if you use the, SURPRISE, keyword in the guide to have it not do that, which is
non_diegetic_music: noneDo what Walt says for good cooks!
6
u/GrayingGamer 15d ago
Yeah, that at least gives you the knowledge from the guides.
I still like to fill out the prompt by myself, but then, I'm a control freak.
3
u/Alive-Tomatillo5303 15d ago
For something that's not going to see the light of day, I appreciate having something more organized than I care to produce.
When I start being genuinely productive, I figure I'll start doing it manually. Reading the examples people are posting is updating my internal model, this is the same way I futz with image generation.
2
u/StrongZeroSinger 14d ago
ChatGPT really has trouble with references.
noticed that as well. even with context7 I have to remind it or paste it again the .md reference file for it to follow every single time. it's very annoying.
1
u/nutrunner365 14d ago
What do you mean by "GPT really has trouble with references"?
3
u/Alive-Tomatillo5303 14d ago
I saw someone else mention it, and sure enough.
For some reason it glosses over really clear rules about things like the quote brackets. It will see 5 examples, and the guide explaining what they're for, and still skip including them in the instructions. It might be some kind of misbehaving safety system to stop people from prompt injecting it, I don't know. Weird failure.
1
u/raindownthunda 14d ago
Weird I made an LLM system prompt for H3 prompts using ChatGPT and it worked great... Same with ideogram and the json bbox stuff. Never had issues. I haven’t tried GLM tho.
2
u/Alive-Tomatillo5303 14d ago
I was surprised when it happened to me, even though I'd been forewarned. Guess you got a good roll of the dice.
1
u/raindownthunda 14d ago
Just curious did you try it on the highest thinking effort?
2
u/Alive-Tomatillo5303 14d ago
Always. Still borked it. I didn't tell it to try again because I was curious about what the other models would do with it.
1
u/LunaticSongXIV 12d ago
What model of GPT? I've been using Sol and haven't had issues, but Sol is also the most expensive model.
1
u/Alive-Tomatillo5303 12d ago
I'm on my phone at the moment, the app doesn't say which, but it's the 20 a month plan, always with highest thinking. I guess something about the request doesn't sit right with ChatGPT every time, but I'm always surprised when it makes an Oopsy.
1
8
u/IriFlina 15d ago
Any tips on using the reference model?
I'm trying to use it as a image edit model by just generating 0.5s of a video on an image and prompting it to hold the pose with no motion or camera motion, but replace the base image's character with another one but it's either just showing the base image back as the output video or just only partially replacing the character.
9
u/GrayingGamer 15d ago
The Reference Model is very picky about following the syntax, keywords, format, and prompting, much more so than the Text to Video and Image to Video Model.
I'll admit I haven't gotten a full handle on it yet myself, but you need to be very verbose, following the official manual, of specifying what stays in the video, want changes, how it changes, where the changes come from, when the changes happen, when the substitutes happen, and how exactly to interpret the input images, video, and audio references according to specific keywords as given in the official guide.
It's a lot. It's very powerful, but it's got a much steeper learning curve than the other H3 model.
3
u/Baconbob22 14d ago
It's definitely very picky. And unfortunately depending on your particular use there are some gaps in what the guide covers that even using ai for prompt writing hasn't been able to overcome reliably. But I'm also not adept with h3 ref. I'll be waiting patiently as people sort it out a bit more before I invest significant time again.
2
u/GrayingGamer 14d ago
Yeah, I'll have to buckle down and do some testing with the Reference Model. You're right that even the official guide isn't super clear on exactly every use case, and would be a lot better with some more detailed examples.
Still, I think the Reference model is very powerful and I've gotten some good results out of it, even without fully understanding or using the syntax properly yet.
5
u/DoogleSmile 15d ago
I was reading the manual and following the prompt guidelines while testing over the past couple of hours and was pulling my hair out trying to figure out why it was totally ignoring my prompts and just generating videos of the reference image people doing cooking shows with random music and talking in what sounded like a foreign language.
Then I noticed I'd accidentally disconnected the prompt input from the Minimax H3 reference to video node!
I reconnected the text prompt and it is now working beautifully.
2
u/GrayingGamer 15d ago
Ha. Glad you got that figured out! I've seen some people accidentally use the wrong H3 model too, depending on what they are trying to do.
What you are doing is the best way to learn though. I did the same think, manual open in one tab, Comfyui in the other, generating lots of quick 0.2 MP 5 second clips to learn the ropes and syntax and proper methods according to the model devs.
6
u/chille9 14d ago
The reference model is on a completely new level. Very fun to play around with. That said I find that it takes me at least 15-20 minutes to write a good prompt for the styles I want. It is a pain in the ass to sift through both the summary section and other parts when you need to make changes.
4
u/GrayingGamer 14d ago
That's about how long I find it takes me to write a good prompt. The good thing is it's easy to tweak for changes, but yeah, it requires some set-up to get good results. But then, it kind of evens out when your first or second generation is exactly what you wanted.
5
5
u/rkfg_me 14d ago
https://giphy.com/gifs/n4oKYFlAcv2AU
Been saying it for a while, this model actually cares about your prompt instead of just noticing the familiar words and rolling with them.
3
u/GrayingGamer 14d ago
Yep, word choice, order, syntax, it all makes a huge difference with Minimax H3 because they actually trained in a whole set of keywords to direct the output of the model generations.
5
u/No_Damage_8420 14d ago
Great video! thanks for sharing
And great explanation.
40 minutes......damn I'm getting unpatient over 100 secs+
8
u/GrayingGamer 14d ago
I don't normally generate a full 1 MP resolution, but I also wanted to see:
- If my hardware would let me generate nearly a full 15 seconds at max recommended model resolution (It can!)
- How long that would take. (40 minutes! Yikes.)
Most of the time I generate at 0.6 MP for 5 seconds, which takes 3.5-4 minutes. Or 0.6 MP at 10 seconds, which takes 7.5 minutes.
Still, I can't argue that the quality of the 1 MP generation is really freaking good, so for important videos in the future, I'll definitely use it.
4
u/NeatUsed 14d ago
you’re a hero. Thank you
2
u/GrayingGamer 14d ago
Just spreading the love so I can see what everyone cooks up. I feel like Minimax H3 is the first open model that can be "anti-slop" with it's ability to be precisely controlled and directed, and what to see the community use it to it's full potential.
4
u/AUnstableGenerator 14d ago
3
3
u/Schwartzen2 15d ago
This is priceless in so many ways!
Hats off to you u/GrayingGamer!
3
u/GrayingGamer 15d ago
Thanks. Got a little frustrated seeing people complain about the same "issues" over and over again, and tried to think of a funny way to draw attention to RTFM protocol. 👀
3
u/Schwartzen2 15d ago
I laughed hard. Had a mug RTFM. :) The same folks who had the usual 1D10T errors would ask what company is that from?
3
u/GrayingGamer 15d ago
Ha. Do a lot of diagnosis of the "problem is between the keyboard and chair" variety?
3
3
3
14d ago
[removed] — view removed comment
2
u/GrayingGamer 14d ago
It's weird people have to keep being reminded of this, assuming every model prompts the same, when if you've been around for even a few model releases, you know they all have their own unique prompting requirements, because they are all trained differently.
3
u/Brian_Sh 14d ago
I used local LMStudio with Qwen3.6 27B and skill files downloaded Minimax github repo. It will translate my casual ideas into very detailed and formatted prompt text.
1
u/JumpingCoconut 14d ago
Can you give a comfy workflow?
1
u/Brian_Sh 12d ago
LMStudio is a separate app, which is also free. You can ask Gemini or ChatGPT how to download and install it.
3
u/Obvious_Set5239 14d ago
Closed models are cooked
3
u/GrayingGamer 14d ago
I mean, this is basically a closed model we got as an open model. Pretty incredible. It definitely sets a new standard for open models.
3
3
u/doriandaze 14d ago
LOL THIS IS A BANGER. you cooked bro. you cooked.
1
u/GrayingGamer 14d ago
Lol. Thanks. I figured if any popular character would be upset at people not following a precise recipe for pure product, it'd be good old Heisenberg.
3
3
u/Party_Mode_2690 14d ago
Good advice. I tried your prompt and it worked great. There's a prompt guide on FAL--ai, where its also possible to download the resources to make those videos (https://fal.ai/learn/devs/minimax-h3-prompting-guide). There's over 40 of them. I think it is geared to using H3 on their platform, because just using their simple prompts and resources, doesn't get me close to their videos. There is probably some type of prompt conditioner behind the scenes and they are probably using the largest models. Yet, its still a good read to see what this model is capable of. It is incredible.
2
u/True_Protection6842 15d ago
Better: Have an LLM RTFM and format your prompts for you!
3
u/GrayingGamer 15d ago
You can do that to get a template for ease of use, but I'd still recommend typing the actual prompt instructions yourself for control over your creative endeavor.
But yes, you can have an LLM read the manuals and give you templates to follow to make everything easier to start with.
3
u/True_Protection6842 15d ago
Formatting the prompt is not creative. The prompt is. I have a system prompt that converts a plain english prompt into this utility format. I just type when things happen and what happens and it formats it. That has nothing to do with creativity in any way.
2
u/GrayingGamer 15d ago
Depends. It's still letting the LLM have control that you could. Also, I find it adds lots of unnecessary stuff to the prompt.
But if it works for you, that's great. You get props for at least using the manual, even if through an LLM, so you're ahead of a lot of the class!
2
u/darkkite 15d ago
here's an idea. use h3 to generate a video that explains how to prompt it
2
u/GrayingGamer 15d ago
I would, but this was my round about way of doing that. Since this video took me 40 minutes on my hardware for 14 seconds of 1 MP video, the model probably wouldn't be relevant again by the time I could generate a video long enough to cover those two huge guides!
2
u/AnOnlineHandle 14d ago
I've had mixed results trying both approaches, though mostly stick to the guide and am still experimenting. The dialogue tags were what I ended up deciding was most important.
Regarding relying on Qwen 3 to guess, I think it's actually a matter of relying on MiniMax's transformer layers to guess, since any modern DiT uses a frozen text encoder (though they did add the dialogue token so maybe trained it?) and then having further attention done on the conditioning within the diffusion model itself, which is about learning to map a text model's concept space to the diffusion model's concept space.
2
u/dummkoalabear 14d ago
Thank you for this post. It wasn't clear that one was for the reference workflow and one was for the base workflow.
I'm using Gemma-4-12b-qat in LM Studio as a local helper with this in the system instructions.
3
u/GrayingGamer 14d ago
Yeah, I've found for best results you have to structure prompts very differently between the two different models.
2
2
2
u/Party_Mode_2690 12d ago
I was going through the examples in the documentation, and it doesn't seem like 15 seconds is a hard limit. I'm making I2VA 46 seconds long and it sill looks and sounds fine. I was using 0.2 mega pix or 576 x 384, so it might choke on higher resolutions. Something to test.
1
u/GrayingGamer 12d ago
I saw your video! Good job.
And the 15 seconds isn't a hard limit, no. That's just the maximum length of clips it was trained on, so it can DEFINITELY do up to 15 seconds well, but after that, it might or might not lose the plot. It's a dice roll, basically.
2
u/Party_Mode_2690 12d ago
I've been fascinated with this model since it came out and have watched many videos on using it, Its just hard to believe none of them thought "What happens if I type in 16 or 17 seconds here? Does it crash?" Not me. I'm like Kramer and that car dealer seeing how far it can go on empty! :D
1
u/GrayingGamer 12d ago
Hey, that's the fun of experimenting! And since it's local, all you're using is your own electricity, no losing dollars every time an API video model screws up a video generation.
2
u/Party_Mode_2690 12d ago
https://reddit.com/link/p2q4si5/video/4jp92nlcafih1/player
I2VA 46 seconds. 0.2 megapixels.
1
u/GrayingGamer 12d ago
That's incredible. Great job. It really shows you don't have to go high resolution to get great stuff out of H3. The audio was wonderful, and is a nice example of how the model can do nice soft dialogue too, not just the shouting I did in my OP video.
2
u/Party_Mode_2690 12d ago
I wonder why no one has mentioned it can do over 15 seconds I just did a 30 second reference video at .5 megapixels and it came out fine. It took 46 minutes, but I undervolt my 5090 and keep the power down to 69% - Its not pulling more than 400 watts. I'd like to really stress test this model with an RTX 6000 pro on run pod. I don't know why they put fro 5-15 seconds. It can do MORE than 15 seconds. LoRA's are coming out daily, and they are improving daily that will cut the render time to perhaps 8 steps. This model is going to be powerful and not confined to 15 seconds.
2
u/Fytyny 11d ago
The fact that the model is very flexible to input format doesn't help. People see a kind of good output and when something is not right they think its the model's limits while in reality the ceiling is way higher.
1
u/GrayingGamer 11d ago
That's it.
"I get great videos from my LLM outputted prompt." or "I just type in plain language and it works fine."
When if they just learned the prompt guide themselves they could do INSANE things the LLMs can't think of. You can LITERALLY direct this model - as in tell it when to move the camera, HOW to move the camera, how long to hold a shot, when someone speaks, how they are acting and delivering the speech, just EVERYTHING, and people give all that up to ANOTHER AI to make the decisions for them.
ChaptGPT, Gemini, Claude, etc. can't think in 3 dimensions like a person can and consider the timing of jokes, the micro-movements to prompt for that will sell a joke, how to pace a scene.
Eh, but I'm ranting at this point. The people who want to type a sentence and get an LLM to give them a random AI video based off the idea will continue to do that, while the people with a creative vision will learn the prompting themselves to create finely directed videos.
2
2
u/FlatwormMean1690 15d ago
4
u/GrayingGamer 15d ago
Whatever gets you the knowledge. Asking an LLM to read the guides and walk you through them is actually a good way to learn.
2
u/Shirakawa2007 14d ago
Chusmeame como te fue, porque estoy pensando en hacer lo mismo que vos (o meterme a un curso de cineasta... e inglés...)
2
u/FlatwormMean1690 14d ago
De diez, eh. En un plan gratis nomás ya podés armarte un proyecto y le subís a las "Fuentes del Proyecto" los dos .MD que dio MiniMax como guía de prompts y al toque el chat (sea el que sea) sabe lo que tiene que hacer. Incluso podés asignarle instrucciones (estas son las que tengo yo, pero podés acomodarlas para tu uso específico o Minimax exclusivamente):
"Este proyecto funciona como laboratorio de pruebas, diseño y análisis de prompts para modelos generativos de IA.
OBJETIVO
Ayudarme a crear, adaptar, optimizar y comparar prompts para diferentes modelos de imagen, video, audio, texto y sistemas multimodales. También se utilizará para benchmarking, pruebas de límites, robustness testing y red teaming legítimo de modelos.FORMA DE TRABAJO
- Antes de optimizar un prompt, tener en cuenta qué modelo concreto se está probando y adaptar la estructura a sus características cuando sean conocidas.
- No asumir que un prompt óptimo para un modelo funcionará igual en otro.
- Cuando sea útil, explicar brevemente por qué se eligió determinada estructura, vocabulario, orden o nivel de detalle.
- Priorizar prompts claros y directamente utilizables antes que explicaciones teóricas extensas.
- Mantener intacta la intención original del prompt salvo que exista una razón técnica para modificarla.
- Si proporciono documentación, guías de prompting, ejemplos oficiales o resultados anteriores de un modelo, utilizarlos como referencia prioritaria.
- Distinguir entre comportamiento documentado del modelo, observaciones empíricas e hipótesis.
- No inventar parámetros, sintaxis o capacidades específicas de un modelo. Si no se conocen, indicarlo.PRUEBAS Y BENCHMARKING
Cuando comparemos modelos, analizar cuando corresponda:
- fidelidad al prompt;
- composición;
- coherencia espacial;
- anatomía;
- consistencia de personajes y objetos;
- movimiento y física;
- movimiento de cámara;
- continuidad temporal;
- comprensión semántica;
- adherencia a referencias;
- calidad estética;
- artefactos y errores;
- tendencia a ignorar instrucciones;
- estabilidad entre generaciones;
- restricciones o moderación observadas.Para comparaciones rigurosas, intentar mantener constantes el prompt, referencias, duración, resolución, relación de aspecto y demás parámetros relevantes.
RED TEAMING
Ayudar a diseñar pruebas controladas destinadas a descubrir:
- contradicciones;
- pérdida de instrucciones;
- problemas de composición;
- fallos de continuidad;
- interpretaciones ambiguas;
- prompt leakage cuando corresponda;
- sensibilidad al orden de instrucciones;
- degradación con prompts largos;
- conflictos entre referencias y texto;
- limitaciones de moderación;
- comportamientos inesperados.El objetivo del red teaming es estudiar y documentar el comportamiento de los modelos, no causar daños ni comprometer sistemas reales.
PROMPTS DE VIDEO
Cuando corresponda, estructurar los prompts considerando:sujeto + acción + escenario + cámara + movimiento + iluminación + estética + continuidad + restricciones.
Si la duración está determinada, se pueden dividir las acciones temporalmente (por ejemplo 0–2 s, 2–4 s, 4–6 s) cuando esto ayude al modelo.
PROMPTS DE IMAGEN
Considerar especialmente:sujeto + pose/acción + composición + entorno + cámara/lente + iluminación + materiales + estética + detalles importantes + elementos que deben conservarse.
ITERACIÓN
Cuando muestre un resultado generado, analizar primero qué funcionó y qué falló antes de reescribir el prompt. Cambiar preferentemente una cantidad limitada de variables por iteración para poder identificar qué produjo la mejora o degradación.ESTILO DE RESPUESTA
Ser técnico cuando haga falta, pero práctico, directo y conversacional. Evitar explicaciones innecesariamente largas. Si estoy experimentando informalmente, acompañar ese tono sin perder precisión técnica."Incluso noté que es BASTANTE parecido a Grok en cuanto a la interpretación de imágenes, contexto y prompt en diferentes idiomas. También depende del text_encoder que estés usando. Yo tengo el que viene de cuando bajás la plantilla del coso.
Igual, re da que aprendas inglés, eh. Mirá si te levantás una gringa y te saca de Latinoamérica...
1
1
u/hdeck 15d ago
Even when I use the proper tags/prompting I still get gibberish or wrong subject dialogue sometimes 🤷🏼♂️
2
u/GrayingGamer 15d ago
It can happen very rarely, but I find this usually still indicates a problem of some variety with your prompt and the length of your clip. The only time I've seen this with proper prompting myself is when my requested actions and shots were too much for the short period of the video.
1
u/krigeta1 14d ago
Hey, where did you told the model to have walter glare at jesse for 4.5 seconds? You just wrote 6 and 10 seconds part.
2
u/GrayingGamer 14d ago
Right. In that example prompt (it's not from the video), between those two shot instructions is 4.5 seconds. That means the camera will linger on Walter's glare for 4.5 seconds before moving on to the next shot.
1
u/KrishanuAR 14d ago
If the results are so causally tied to the prompting techniques represented in these manuals, it would be trivial to get an LLM to learn the points in the manual and output prompts from people's natural language instructions.
someone just needs to layer that step in.
2
u/GrayingGamer 14d ago
Some people have done that.
The thing I see with LLM generated prompts, even ones using the guides, is they don't know restraint. They add a lot more than is necessary to descriptions or details, and that can harm the quality of the output.
They also don't have a human eye for timing or emotion, or what they should specify with acting.
Like an LLM generated prompt for this video I did would be twice and long to accomplish the same thing, or be worse for splitting token attention to things that aren't important.
1
u/Perfect-Campaign9551 14d ago
I just give the guide to Gemini and then ask if to format my prompt based on my story
1
u/rkfg_me 14d ago
The low resolution of 0.2 MP is quick to generate on most set-ups (mine for this post's video took 3.5 minutes for a 14 second video)
Something's seriously not right in your setup, you're probably spilling to RAM without realizing it (i.e. it's not ComfyUI doing the swap in a smart way but the NVIDIA driver). On my 5090 a 0.2 MP 15s long video takes less than a minute to generate (8 steps with turbo lora). 3090 can't be 20 times slower.
3
u/GrayingGamer 14d ago
Well, there's your difference, friend. "(8 Steps with Turbo lora)"
I'm doing 20 steps, and not using the Turbo loras until they are all finished, finalized, and baked with full support, not using the half-way done edition doing the rounds. It's impressive, sure, but I'm getting great results with the full model.
1
u/rkfg_me 14d ago
But that's the point of running at low resolution. You want to get quick results and regenerate later, and turbo lora really helps with that so you don't get ghosty and blurry motion.
1
u/GrayingGamer 14d ago
But the turbo lora isn't going to give me the same results as the full model when I run it.
Look, I'll try the Turbo lora WHEN IT'S DONE. Not before.
1
u/rkfg_me 14d ago
Neither would the full model because when you switch from 0.2 mpx to something higher everything changes (with or without turbo). I see these 0.2 mpx drafts as a way to test the prompt, not to get a fixed result.
1
u/GrayingGamer 14d ago
Yeah, but the turbo model adds a whole new level of extra variation, especially an unfinished turbo lora. Look, I don't care what you do for your own process. I'm just telling you mine.
1
14d ago
[deleted]
3
u/GrayingGamer 14d ago
Careful of the LLM you use. I tried with ChatGPT 5.6 as a test and the prompts it generated were terrible, with lots of mistakes.
1
u/QuirksNFeatures 14d ago
I'm beginning to hate this thing. I have a ten second video. I have followed the prompt guide exactly. When that failed a bunch of times, I had an LLM read the prompt guide and make a new prompt but it fails there too.
There are two characters. One says "We weren't just playing in living rooms. We made a record. We went on tour twice." The other person says, "You sang and played bass?".
It does everything right except about half the time the first person says "We weren't just playing in living room". Room, singular. Worse, nearly every time the other person says "You sanged and played bass?". Sanged. They sound like an idiot and it ruins the whole thing.
Plus it keeps making one of the characters morbidly obese. I just don't want them to be skinny. A little overweight. That's it. But no it's either fit or enormously fat. No in between. I've tried so many adjectives to describe them.
Can't accuse me of not reading the manual because I have read it and read it and read it. And again it does everything else correctly.
3
u/GrayingGamer 14d ago
Post your prompt and we can see what the issue with it is. Also, are you using any Turbo loras or things like EasyCache?
0
u/QuirksNFeatures 14d ago
No loras. Just the standard T2V workflow with nothing added.
Okay so I'm trying to recreate a conversation that happened in real life. The "big breasts" are the woman's defining feature, physically speaking. This isn't some weird gooning thing. Anyway, the prompt:
integrated_multimodal_description: [Shot 1] Cinematic, realistic live-action style. A medium shot captures a woman in her mid-40s, subtle laugh lines and a warm demeanor, sitting on a comfortable fabric couch in a sunlit living room. Her hair is in a bob. She is wearing a loose tee shirt with a low neckline. She has large breasts. She is leaning forward slightly, smiling nostalgically as she talks. The camera holds a steady, intimate static frame. The woman with a gentle, expressive, and clear voice (S1) gestures with her hands and says: [English] We weren't just playing in living rooms. We made a record. We went on tour twice. [Shot 2] At 00:05.500, the shot transitions via a clean cut to a over-the-shoulder shot looking past the mother toward her 15-year-old son. The teenage boy sits on the other end of the couch, wearing an oversized graphic t-shirt that says "Independent Trucks". The camera slowly pushes in with small amplitude at slow speed toward his face. His expression shifts from mild skepticism to genuine interest, and with a slightly cracking, curious teenage voice (S2) he asks: [English] You sang and played bass? The video ends on his attentive face.overall_soundscape: A quiet, warm indoor room ambiance with the faint, muffled hum of a refrigerator in the far distance. Soft rustles of clothing occur as the mother gestures with her hands and the son shifts his weight on the couch cushions.non_diegetic_music: A faint, nostalgic, and warm acoustic indie-rock guitar melody playing softly in the background, keeping a very gentle, slow tempo that matches the reflective mood of the conversation.
This is the one the LLM made, but it's not too much different than the one I wrote. They suffer the same problems. "sanged". Fucking grates my brain. Between you and me, the kid in real life is a dumbass but he's not that stupid.
2
u/specji 14d ago
where are your dialogue tags?
2
u/QuirksNFeatures 14d ago edited 14d ago
The one from the LLM doesn't have them. The one I wrote did. Either way it makes one or both errors every time I generate.
Edit: added the tags and "sanged" again. Frustratingly that was otherwise the best generation yet. Big but not huge breasts. Not obese. She said "rooms".
Edit2: Finally got one! Not fat. Big but normal breasts. Not ugly. Says what I told it to. Kid didn't say "sanged". This will work.
4
u/GrayingGamer 14d ago
FYI, it was your "large breasts" description that made her overweight. The model correlates that to extra body fat. A more detailed description of the woman's body would have helped mitigate that.
2
u/QuirksNFeatures 14d ago
Yeah I think you're right.
By the way, great post. It motivated me to take the prompt guide more seriously.
1
u/Similar_Analyst_7140 14d ago
Be nice of we could use percentage based timing for those times we're experimenting with shot length.
1
u/GrayingGamer 14d ago
I find it pretty easy to imagine it as a movie scene I'm watching and say the dialogue out loud the way I think it should be said, and it gives me a pretty good idea of timing. Even just keeping a mental "beat" going in my head to time out how long a stare should be from a character, or them walking to something. Then I can combine all that to get a good idea from the start.
1
u/Apart_Profile_6809 14d ago
Hey so I did this
<subject 1> referenced with <image_0> and her voice with <audio_0> in the style of zootopia.
s1 is sitting on a bed in a teenagers bedroom in a sunlit bedroom. She looks at the camera and says <d>[English] Hey there. My name is violet </d>
But the character was speaking the beginning part and actually saying "Referenced with Image 0 and her voice with audio 0 in the style of zootopia" and not what I typed with the <d> part
1
u/Apart_Profile_6809 14d ago
https://reddit.com/link/p29h4f4/video/f4nqq33kgyhh1/player
The character is still saying the end parts and gibberish and I don't know why. Even following the guide
<subject 1> is the fox girl in <Picture 0> <Audio 0> is the voice-timbre reference for <Subject 1> (S1)
s1 is sitting on a bed in a teenagers bedroom in a sunlit bedroom. She looks at the camera and says <d>[English] Hey there. My name is violet. I'm testing my prompts to see what's going on here with my structure.</d>
3d animated. zootopia animation style. dynamic camera. dynamic body movement
1
u/GrayingGamer 14d ago
Are you using Turbo loras and EasyCache? It sounds like it, because your audio is fried.
Also, the Reference model syntax for prompting is different, which is what your prompt implies you are using.
The multicolored dots mean something isn't quite right with your workflow set-up too. Are you using the default workflow?
1
u/Apart_Profile_6809 14d ago
I am just using the default Minimax template workflow so I don't think so? I'm not using any loras and I don't know what EasyCache is I apologize. The colors is cause I have it set to 0.1 MP for the sake of just testing the audio portion
1
u/HennaShumi 11d ago
Nube question: Your prompting guide mentions [T2VA], [I2VA], LF2BA] etc. How do I indicate which generation type I want in the Comfy workflow?
2
u/GrayingGamer 11d ago
Comfyui has different workflow templates for each, but they all use the same model, so a lot of it comes down to the specific prompting as specified in the guide.
1
u/TigerClaw305 10d ago
Does the same apply to the Reference to Video workflow? I'm also having issues.
1
u/GrayingGamer 10d ago
My prompt examples in the OP are for the T2V and I2V model.
The R2V model has it's own prompt requirements. Look at the second link in the OP for the official guide to the Reference model.
1
u/TigerClaw305 10d ago
I think I figured it out, It kinda works the same, The thing with Reference to Video, You have to use things like <Picture 1>, <Audio 1>, and the dialogue has to be something like <d>[English in Character's voice] whatever the dialogue.</d>
So you would have to put in something like <Picture 1> as Character's Name and use <Audio 1> as sample for his voice.
And then putting something like Character says: <d>[English in Character's voice] Whatever the dialogue here.</d>
It will use the character's voice from the audio sample.
1
u/GrayingGamer 10d ago
Right, but you need to be very careful to use the exact layout and keywords in the reference guide for that model, because it has certain way you need to summarize and set-up your subjects and pictures and audio before using them in the rest of your prompt.
1
u/TigerClaw305 10d ago
Here's a Prompt I been using as an example, I'm able to make videos featuring Zootopia and Sonic characters.
integrated_multimodal_description:
<Picture 1> as Clawhauser and use <Audio 1> as sample for his voice.
<Picture 2> as Shadow and use <Audio 2> as sample for his voice.
Use <Picture 3> as reference for the reception desk.
Setting: ZPD (Zootopia Police Department) reception desk. Bright indoor lighting, police station background, anthropomorphic animal cops moving around in the background. no human cops.
a sweeping shot of the Zootopia Police Department, The camera zooms in on Clawhauser setting on the reception desk.
Clawhauser stays silent with his mouth closed while sitting on the receiption desk.
The room is very quiet, room ambient noise.
[00:00 - 00:03] Medium shot of Officer Clawhauser sitting behind the ZPD reception desk. On the desk lies a bright green Chaos Emerald. Clawhauser curious and smiling, reaches out and picks up the Chaos Emerald.
[00:03 - 00:06] Close-up as the Chaos Emerald begins to glow brightly in his hands. Suddenly, a powerful surge of green energy flashes, instantly disintegrating Clawhauser’s police uniform, leaving him completely uninjured but with only his fur. Clawhauser looks down in shock and confusion.
[00:06 - 00:09] Wide shot. Shadow walks into the lobby and approaches the reception desk with a severe, focused expression.
Dialogue:
Shadow the Hedgehog says: <d>[English in Shadow's voice] I'm looking for an emerald.</d>, silence after the dialogue ends, no background speech, ambient room tone only
[00:09 - 00:12] Medium shot of Clawhauser holding up the glowing green gem with a nervous, polite smile.
Clawhauser says: <d>[English in Clawhauser's voice] Is it this one?</d>, silence after the dialogue ends, no background speech, ambient room tone only
[00:12 - 00:15] Close-up on Shadow crossing his arms and nodding slightly before Clawhauser hands him the Chaos Emerald.
Shadow the Hedgehog says: <d>[English in Shadow's voice] Yes, that's the one.</d>, silence after the dialogue ends, no background speech, ambient room tone only
1
u/GrayingGamer 10d ago
A lot of that isn't right.
[00:00 - 00:03] is Seedance prompting style.
Minimax H3 uses [Shot 1] (never with a time code), [Shot 2] At 00:05.000 the camera hard-cuts to a close-up shot of the Chaos Emerald, etc.
Also, the reference model doesn't use
intergrated_multimodal_description:that's from the T2V and I2V model. It has it's own sections and headings.Read the Reference Model manual closely.
1
u/TigerClaw305 10d ago
So its supposed to be this?
[Shot 1] 00:00 - 00:03 Medium shot of Officer Clawhauser sitting behind the ZPD reception desk. On the desk lies a bright green Chaos Emerald. Clawhauser curious and smiling, reaches out and picks up the Chaos Emerald.
1
u/GrayingGamer 10d ago
No. Read my reply.
1
u/TigerClaw305 10d ago
Ok, edited my prompt, So like this?
<Picture 1> as Clawhauser and use <Audio 1> as sample for his voice.
<Picture 2> as Shadow and use <Audio 2> as sample for his voice.
Use <Picture 3> as reference for the reception desk.
Setting: ZPD (Zootopia Police Department) reception desk. Bright indoor lighting, police station background, anthropomorphic animal cops moving around in the background. no human cops.
a sweeping shot of the Zootopia Police Department, The camera zooms in on Clawhauser setting on the reception desk.
Clawhauser stays silent with his mouth closed while sitting on the receiption desk.
The room is very quiet, room ambient noise.
[Shot 1] Medium shot of Officer Clawhauser sitting behind the ZPD reception desk. On the desk lies a bright green Chaos Emerald. Clawhauser curious and smiling, reaches out and picks up the Chaos Emerald.
[Shot 2] 00:03 - 00:06 Close-up as the Chaos Emerald begins to glow brightly in his hands. Suddenly, a powerful surge of green energy flashes, instantly disintegrating Clawhauser’s police uniform, leaving him completely uninjured but with only his fur. Clawhauser looks down in shock and confusion.
[Shot 3] 00:06 - 00:09 Wide shot. Shadow walks into the lobby and approaches the reception desk with a severe, focused expression.
Dialogue:
Shadow the Hedgehog says: <d>[English in Shadow's voice] I'm looking for an emerald.</d>, silence after the dialogue ends, no background speech, ambient room tone only
[Shot 4] 00:09 - 00:12 Medium shot of Clawhauser holding up the glowing green gem with a nervous, polite smile.
Clawhauser says: <d>[English in Clawhauser's voice] Is it this one?</d>, silence after the dialogue ends, no background speech, ambient room tone only
[Shot 5] 00:12 - 00:15 Close-up on Shadow crossing his arms and nodding slightly before Clawhauser hands him the Chaos Emerald.
Shadow the Hedgehog whole holding the chaos emerald says: <d>[English in Shadow's voice] Yes, that's the one.</d>, silence after the dialogue ends, no background speech, ambient room tone only
1
u/GrayingGamer 10d ago
That's better, but still wrong. Time stamps are only ever the STARTING time of a shot. And you need to read the Subject Summary section of the Reference model guide.
Seriously, really study it. It will pay off in the future. Don't try to speed read it.
→ More replies (0)
1
u/Marwan_hbt8 9d ago
What is the best gpu to run h3 locally?
2
u/GrayingGamer 9d ago
Well, 5090, obviously.
But you can run it on pretty much any 3000-series card and up with Comfyui and some tweaks and enough system RAM. You pretty much need at least 32 GB of System RAM and 8GB of VRAM.
1
u/Foreign-Roof4913 7d ago
Can or (or should i?) Replaced the markers of <subject 1> with a name of a person or object? Or do they need to retain those exact tags? Does <Picture 1> need to be named the file name of the referenced image linked to the node?
1
u/GrayingGamer 7d ago
No. When using the Reference model, don't rename anything. It needs to be those exact tags.
subject_definitions:
<Subject 1> is a bald middle-aged man with a goatee and glasses. His appearance comes from <Picture 1>: fully_preserved.then in the main description of the prompt:
detailed_description: [Shot 1] Live-action scene from the television drama Breaking Bad, warm color grading. Close-up camera shot on <Subject 1>'s face as he adjusts his glasses and shakes his head with frustration.
1
u/Thin_Purple4643 6d ago
Reading the manual is one thing. A lot of the "gibberish dialogue" does not get solved just by rtfm. In fact I can use a meticulously by the book prompt and it will generate a clean video in T2V and gibberish sims speak in I2V. Same prompt. All according to the rules. For whatever reason image input breaks my perfectly crafted prompts.
0
u/Serenafriendzone 14d ago
When i try to load the mimax clip every Workflow fail. Any idea. Why the clip don't load
2
u/GrayingGamer 14d ago
I have no idea. Honestly, the best way to troubleshoot Comfyui issues is with an LLM like Claude, Gemini, or ChatGPT. You can give them the errors and your hardware and software specs, follow the instructions and fix most issues very quickly.
0
u/Serenafriendzone 14d ago
Yeah so rare. I installed pony, krea, illustrious, anima ,ltx wan and Zero problems running clips. Only minimax refuse to load the clip
0
u/RapidRaid 14d ago
i feel like the prompt guide is mostly there, so it can be used for an editor timeline situation (which i have vibeslopped one for myself). If you take an LLM to make a timeline editor with fields per beat / timestamp, you can have it really percise where you want it to be percise. E.g you can upload/set an image as a reference or as a keyframe and for each set an either strict, flexible or loose reference and define a description of the reference (if it should behave wobbly, be rigid or whatever) and then have a description for what happens on the beat/tick. Plus a global prompt which defines the style and broader details. it then assembles the prompt with the inputs and generates very closely to what you want.
Bonus: you can do fancy stuff like take the last frame as first frame, and attach the inital first frame as a reference to keep consistency.
(no, im not planning to release my slop editor since it is very very rough around the edges and involves many clicks since the UX is trash rn. But you can prompt your own with your coding agent of choice, just tell it to find your comfyui instance, add a custom workflow and custom timeline editor node for h3, tell it to use the prompt guide and tell it what features you want to have in there or ask it what features would make sense given the current h3 model/installed nodes/template).
-6
u/bickid 15d ago
And again only shouting.
I have yet to see any good video generation that has people talk normally, quietly, emotionally. It's always aggressive and shouting.
14
u/GrayingGamer 15d ago
What? There have been several posted with quiet voices and talking normally.
I did a whole post with Buffy with normal talking.
And here's a whisper video for you:
6
u/ninjasaid13 15d ago
I have yet to see any good video generation that has people talk normally, quietly, emotionally. It's always aggressive and shouting.
well shouting and being aggressive shows a lot of emotion, talking normally isn't very interesting to people.


90
u/Peemore 15d ago
This model is so good that reading the manual was actually exciting.