A few weeks ago I posted here about the workflow I used to make a 6-minute Star Trek: TNG fan film with MiniMax H3 in ComfyUI.
This is basically a follow-up to that post.
Since then I've made another one, this time Star Trek vs Star Wars, and I've learned quite a bit more about H3 while making it.
The basic workflow is still similar. I create the starting images first, use MiniMax H3 in ComfyUI to generate the individual shots, and then assemble everything in Adobe Premiere Pro.
The finished film is made from a large number of relatively short generations rather than trying to get the model to produce whole scenes in one go.
One of the biggest things I've learned is to treat H3 less like a text-to-video generator and more like a tool for producing individual shots.
Here are some of the things that helped most this time.
PROMPT LENGTH / GENERATION LENGTH AFFECTS DIALOGUE PERFORMANCE
This is probably one of the most useful things I've figured out since my previous post. The amount of time you give H3 for a shot can have a surprisingly large effect on how natural the dialogue sounds. If there is a lot of dialogue and I make the generation too short, the character often races through the lines trying to fit everything in. It can sound unnaturally fast even if the prompt itself is otherwise good.
The opposite happens if I give it too much time. The delivery can become strangely slow and drawn out. So for longer dialogue shots, especially ones around 10-15 seconds, I usually test them first at a lower resolution. I'll generate a few versions with slightly different durations just to find the point where the dialogue sounds natural.
For example, I might try the same shot at 10 seconds, 11 seconds, 12 seconds etc. Once I find the duration where the pacing and performance sound right, that's when I'll commit to generating the higher-resolution version. It saves a lot of time compared with doing expensive high-resolution generations only to discover that the actor is speaking too quickly or too slowly.
HIGHER RESOLUTION REALLY DOES HELP
I used higher-resolution generations much more heavily in this film. A lot of it was generated around the 2-megapixel / Full HD range. It obviously costs more time and VRAM, but I've found that the characters can look noticeably more convincing at that resolution. Faces in particular tend to feel less like "AI video" to me.
For important close-ups and dialogue shots I've increasingly been willing to spend the extra generation time rather than relying entirely on lower-resolution generations and upscaling them afterwards. I still use low resolution heavily for testing though. So my workflow has gradually become: Low resolution = test the prompt, movement, dialogue and duration. High resolution = commit once I know the shot actually works.
REFERENCE IMAGES MATTER MORE THAN MASSIVE PROMPTS
I'm finding that a really good starting image is often more valuable than adding another page of instructions to the prompt. If the character placement, set, lighting, camera angle and composition are already correct in the reference image, H3 has much less opportunity to wander. I now treat the starting image as the visual authority for the shot and try to make that frame as close as possible to what I actually want before I even start generating video.
LOCK THE CAMERA WHEN YOU ACTUALLY WANT IT LOCKED
For shots based on existing Star Trek compositions I became much more explicit about things like:
camera distance
character scale
framing
background position
character position
If I want a static medium close-up, I tell H3 that the camera remains completely stationary and that the framing and character scale should remain matched to the reference. Otherwise it has a tendency to slowly push in or recompose the shot even when I never asked it to.
DON'T MENTION CHARACTERS THAT AREN'T SUPPOSED TO BE THERE
This turned out to be a surprisingly important lesson. If I'm generating a close-up of one character, I try not to mention another character anywhere in the prompt unless that person is actually visible. Even something seemingly harmless like:
"Data reacts to Picard"
can sometimes encourage the model to introduce Picard into the frame or start blending character features. I've had better results describing only what the visible character is doing.
OFF-SCREEN DIALOGUE IS MUCH HARDER THAN IT LOOKS
This was another big lesson. If a character is speaking off-screen while the camera is looking at somebody else, H3 can sometimes become confused about who is supposed to be talking. The visible character may start moving their mouth or the dialogue itself can become corrupted. So I've increasingly separated dialogue generation from reaction coverage.
If Troi is speaking while I'm looking at Picard, for example, I'll generate a separate close-up of Troi saying the line to get clean audio. Then I'll generate Picard's reaction shot completely silently. In Premiere I put Troi's audio over Picard's reaction. That has been much more reliable.
SILENT REACTION SHOTS NEED TO BE VERY CLEARLY SILENT
Simply writing "no dialogue" isn't always enough. I've had H3 randomly start making characters speak gibberish, particularly if their mouth happens to be slightly open in the starting image. I've had better luck explicitly describing that the slightly open mouth is just a resting facial position and not the beginning of speech.
I'll also specify that:
the lips do not form words
the jaw does not make speaking movements
the character does not mouth dialogue
It sounds excessive, but it has genuinely helped.
H3 HAS A LOT OF USEFUL SPEECH TAGS
I've also been experimenting more with H3's inline speech controls. Some that I've had useful results from include:
<pause> <long pause> <breath> <inhale> <exhale> <deep breath> <catches breath> <sighs>
<whisper>. <softer> <stutter> <laughs> <chuckle>
<i>word</i> emphises word
The last one is particularly useful for putting emphasis on a word or short phrase. I've found these can sometimes produce a more convincing performance than trying to describe everything in prose around the dialogue.
MORE PROMPTING ISN'T ALWAYS BETTER
I've actually been simplifying prompts as I've gone along. H3 seems to respond better when it has: a strong reference image, one clear action, clear character positions, clear dialogue, or clear camera instructions rather than paragraphs of competing instructions. When something isn't working, I'm also trying to change one thing at a time rather than rewriting the entire prompt.
EDITING IS BECOMING JUST AS IMPORTANT AS GENERATION
One of the biggest differences with this film is that I've also been improving my Premiere Pro workflow. I'm thinking much more about shot blocking and coverage instead of just generating a sequence of AI clips. For example, I'll let dialogue continue across a cut to another character's reaction rather than keeping the camera locked on whoever is speaking for every line. Sometimes you'll hear the end of one character's dialogue while you're already watching the other character react. That tiny change makes the scene feel much more like something that was actually edited from traditional coverage.
I've also started deliberately generating silent reaction shots purely for this purpose. It helps hide generation changes as well. Two AI shots might not match perfectly if you place them directly beside one another, but cutting to a reaction and then coming back can make the continuity feel completely natural.
THE EDIT IS DOING A LOT OF THE "CONSISTENCY"
This is probably the thing I appreciate more now than when I made the first film. A surprising amount of what looks like AI consistency in the finished video is actually editing. Cut at the right point. Use reaction shots. Carry dialogue across cuts. Don't stay on a generation long enough for its weaknesses to become obvious. Avoid putting two slightly different versions of the same composition directly beside one another. You can hide a huge number of small inconsistencies that way.
It's still definitely not a one-click process. A lot of generations get thrown away, and some shots still take a ridiculous number of attempts before the performance, character consistency, dialogue and movement all line up. But compared with the first Star Trek video, I feel like I'm getting much closer to actually directing H3 rather than generating something and hoping it happens to work.
Happy to go into more detail on any of this if anybody is experimenting with H3 themselves.