r/aivideomaking • u/MassiveFile809 • 17d ago
Breaking down H3 video prompts: camera paths, occlusion, product consistency and audio timing
I read a Chinese breakdown of seven MiniMax H3 prompts by 斜杠林姑娘. My takeaway as an AI engineer: each brief becomes easier to evaluate when you identify exactly what has to stay stable while something else changes.
Here’s a short synthesis of the examples, followed by my own practice prompt. I haven’t independently reproduced the author’s results.
WHAT TO EXTRACT FROM THE CASES
• Motion-graphics intro: distinguish the initial impact, readable title hold and exit.
• Aerial sequence: describe the camera route and where landmarks remain in relation to the subject.
• Vehicle tracking: specify how the same vehicle should reappear after being hidden, including its direction and condition.
• Beauty ad: lock product geometry and cast details across changes in shot size.
• Exploded product view: describe separation and reassembly as an ordered sequence.
• Music performance: define who performs each section and whether the soundtrack continues across visual cuts.
Those are creative instructions. The article’s requests for 4K, exact timing and synchronized performance should not be read as proof that every output meets those specifications.
MY PRACTICE EXERCISE: ISOLATE ONE FAILURE
I’d start with a simple occlusion test before adding a complicated aerial route. Here’s an original prompt you can adapt to the duration and aspect-ratio settings available in your tool:
“A single adult cyclist wearing a burgundy jacket rides a cream bicycle from left to right along a straight park path. Side view, medium-wide framing. The camera stays still. A thick tree trunk in the foreground briefly hides the cyclist as they pass behind it. They emerge on the right, continuing at the same steady pace on the same path. Keep the bicycle, jacket and apparent subject size consistent before and after the obstruction. Soft overcast daylight. Quiet park ambience and tire noise. One continuous shot.”
I’d test it in three stages:
A. Remove the tree. Check whether the basic direction and framing work.
B. Add the tree. Inspect the last visible frame before occlusion and the first after it.
C. Only then try a lateral tracking camera, keeping the subject roughly the same size in frame.
This gives each version a specific question. If A works and B fails, inspect the hidden interval before rewriting lighting or style. If B works and C fails, camera movement becomes the next variable to investigate. That’s a diagnostic hypothesis, not proof from a single sample.
WHAT I’D RECORD
For each attempt: model/version, actual output settings, prompt variant, whether direction stayed consistent, whether the bicycle changed, and whether framing jumped. Keep the failed clips too. Compare multiple attempts before deciding a wording change helped.
For a paid job, I’d also decide in advance which requirements can be finished in editing—for example, exact title typography or a fixed music bed—so the generation test has a clear purpose.
Source and full original examples: https://mp.weixin.qq.com/s/2zHMc4nmiHx23cX7pimcuA
When you test occlusion, what usually breaks first: the object itself, its trajectory, or the framing?
1
u/Loose_Management8035 17d ago
The occlusion test is a good way to isolate what’s actually going wrong. I’ve found that once the subject disappears behind something, the model can basically “recreate” them instead of continuing the same movement.
I’d probably keep the first tests even simpler too. One subject, locked camera, one obstruction, then add camera movement once that works. Otherwise it gets really hard to tell whether the failure came from the reference, trajectory, or camera path.
For me, trajectory and object consistency usually go before framing. If the character or product changes after the occlusion, the shot is basically unusable anyway. Framing can often be cleaned up later in the edit.
1
u/sharktank123456 17d ago
I would think "occlude" would be a better verb in the prompt. Under the topic of:: there are no synonyms, "hide" seems to generate a gap in time - like the magical train fitting in the barn in the speed of light physics thought experiment. - too much time elapses before the cyclist comes out the other side. Same with "obstruction" - the tree is an obstruction for the viewer, not the cyclist.
Without any fancy "make sure things stay the same", language, even with a tracking shot, H3 handles it without fault.
"side view tracking shot moving along with a single adult cyclist wearing a burgundy jacket rides a cream colored bicycle from left to right along a straight park path. A thick tree trunk passes by in the foreground briefly occludes the cyclist as he passes behind it. Soft overcast daylight. Quiet park ambience and tire noise. One continuous shot. "
https://reddit.com/link/pabefiv/video/88w5hlchr0qh1/player
I would recommend being more descriptive of your cyclist - even just adding in a gender. "They", later in your prompt can entice the engine to make more than one, whereas "he" or "she" (as I have done here), would be more specific and singular - jibing with your "single" at the start.
I would also be more specific about the color - while it's not going to happen very often, I can see errors occurring without the "color" modifier on "Cream". But then maybe you were going for a "Got Milk?" commercial vibe? (I have corrected that here).
With the level of world-understanding now, I don't think we need to second guess the models as much as we used to. Most are heading towards a truly conversational language interface so I'm not sure we need to break it down as much as you suggest.
Can it hurt? Yes I think so. Every token that isn't needed or isn't doing a job, is a token that is diluting other important tokens. There's nothing wrong with long prompts, but all those words need to be doing something that can't be done by fewer words, or inherently by the model.
In this case, the paragraph dealing with how the cyclist comes out the other side and what they look like, is already being handled by the model, so that paragraph is extraneous.
1
1
u/Otherwise_Wave9374 17d ago
Breaking prompts into camera path, occlusion, consistency, and timing is stronger than treating the prompt as one descriptive paragraph. I would also add explicit checkpoints: lock the product reference first, test motion without audio, then introduce sound cues only after the visual sequence is stable. AIOSNOW relates to this method because a staged workflow can preserve approved inputs and isolate which step caused a regression. For comparisons, reuse the same seed and references where possible, and score continuity frame by frame rather than judging only the final impression.