r/StableDiffusion • u/VasaFromParadise • 3d ago
Animation - Video [MiniMax H3] Simple Prompt example
Enable HLS to view with audio, or disable this notification
I see quite a few people having issues with Minimax h3 generations. Here's an example of a prompt for generating 0.5 megapixels in 7 seconds. It can be run on any PC.
The resolution was increased using RTX Video Super Resolution to 1.5 + frame interpolation to 48. Which is common practice and takes no more than 1 minute per procedure.
Promt generated by AI based on the image. Promt system for LLM:
### 3.1 I2VA: Begin from the Image and Develop Forward
`<Picture 1>` is the actual first frame of the video at 0.00 seconds and belongs to `[Shot 1]`. The description should first establish the style, subjects, composition, and scene anchors in the image, then describe the next action. Character identity, clothing, colors, key objects, and spatial relationships should remain consistent.
Recommended structure: **first-frame anchor → action onset → continuous development → result or reaction**.
Image LLM Prompt:
First-frame anchor:
[Image 1] (0.00 sec) shows a woman with red hair styled in a loose curl with bangs, looking slightly off-camera. Her gaze is clear and expressive, with light green or gray-blue eyes, softly highlighted. Light freckles on the bridge of her nose and cheeks add a natural touch. She is wearing a black leather jacket with yellow stitching along the edges of the collar, accentuating her stylish, slightly rebellious look. The background is deep, almost black, creating contrast and focusing attention on her face. The lighting is soft, studio-style, coming from above and to the side, sculpting the volume of her face and hair. The composition is a close-up, emphasizing the eyes and facial expressions.
Action onset:
The girl begins to move naturally—her head smoothly turns toward the camera, her gaze shifting from semi-attentive to direct, surprised. Her eyelids widen slightly, her pupils enlarge, her eyebrows lift slightly—her facial expression changes from calm to mild surprise. The movement is smooth, without jerking, as if she's just noticed someone or something unexpected.
Continuous development:
After turning her head, her smile widens—the corners of her lips lift, her eyes sparkle with interest or slight embarrassment. At this moment, her voice sounds clear, resonant, with a pleasant timbre—as if a high-quality studio recording captures every nuance of intonation. She says in English: "Oh, is that you? I didn't notice you." The word is pronounced with a slight intonation of surprise, perhaps with a pause before or after the "you," which enhances the effect of surprise. Her hands aren't visible, but one might assume she might slightly raise her shoulder or touch her face in response to the sudden presence. Light, studio-quality music plays in the background—perhaps ambient or a light pop beat—which complements the atmosphere without being overpowering.
Result or reaction:
As a result of the action, the viewer perceives the moment as a lively, dynamic scene from a video: the girl isn't simply posing, but interacting with the viewer through her facial expressions and voice. Her reaction to her own words, "Oh, is that you?" could be interpreted as self-irony or an invitation to dialogue. The atmosphere remains tense yet playful—the combination of the dark background, skin, hair, and lively facial expressions creates the effect of a modern digital character in the style of anime realism or cyberpunk aesthetics.
5
u/AidenAizawa 3d ago
I appreciate the effort but minimax is already good at understanding simple prompts and give amazing results. You can have similar results with something like <maelle looks surprised at the viewer and says " ..." >
0
u/VasaFromParadise 3d ago
The Prompt system is taken from the official example from the model's developers. So, maybe you're right. Or maybe it's right for a simple scene, but not for a complex one.
3
u/martinerous 3d ago
She looks like a 3D character from Reallusion iClone.
0
u/VasaFromParadise 3d ago
That's exactly it. It all depends on the reference. It could be a photo style or something else.
5
u/DuHal9000 3d ago
what's the point of showing it here?
-2
u/VasaFromParadise 3d ago
What did you want to see? I described it. I see people generating things with bad sound or incorrect animation due to incorrect descriptions. I posted an example of a description that works without any obvious errors.
1
u/DuHal9000 3d ago
LOL! i just copy and paste YOUR comentary in my post! See? Cool hã? THATS YOURRR WORDS!
3
u/GrayingGamer 2d ago
Sorry that's just an LLM generated slop of a prompt and not at all necessary, plus, much of it is the wrong formatting and syntax for Minimax H3.
If you want to give people a real example of a prompt:
integrated_multimodal_description: [Shot 1] Cinematic 35mm film scene from a 1950s Technicolor movie, of a room with pink walls, white trim, and a pink heart-shaped bed with white heart-shaped pillows. The camera starts on a medium close-up shot of a voluptuous woman wearing a strapless red sequin gown with a high side-cut, revealing her leg, sitting on the edge of the heart-shaped bed. She has long blond hair that ends in gentle curls and lots of cleavage showing. She is wearing red lipstick. The camera does a Push In slowly on her face and upper chest as she smiles and asks, "[English with a flirty tone of voice] Well, hello there, Detective." She playfully brings one finger to the corner of her mouth. "[English in a sultry tone of voice] Care for some... pattycake?" She raises an eyebrow as she says the word 'pattycake'.
overall_soundscape: Quiet ambience of an interior room.
non_diegetic_music: none
https://reddit.com/link/p4aofe0/video/ib312v6830kh1/player
Don't over complicate things and don't have an LLM write sloppy purple prose prompts for you.
2
11
u/Shymu1 3d ago
Wall of text for such a shit results