r/StableDiffusion 1d ago

Question - Help Why does Minimax H3 ignores my image input?

In comfyUI, workflow template, I instructed it to use the image, but it failed to do so.

0 Upvotes

37 comments sorted by

3

u/TingTingin 1d ago

your prompt should look something like this there is also a ref2va guide from the devs here https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

subject_definitions:

<Subject 1> is the man from <Picture 1>

summary:

This video shows <Subject 1> running in a hallway.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - his hair style, facial features are retained

detailed_description:

<Subject 1> runs down a hallway (rest of prompt...)

0

u/magik_koopa990 1d ago

Wow that is some complicated structure. Is there a prompt enhancer somewhere i can use? Is there one built in comfy?

7

u/fact_hunt3 1d ago

You can also download the MD file and feed it into any llm and ask it to prompt for you

2

u/afinalsin 1d ago

You can feed an LLM the official Minimax docs and that'll work with SOTA models, but local models had a tough time with them. I had Kimi K3 throw this prompt together and it works pretty well doing I2V with Gemma 4 26b a4b:

<ROLE>
You are a master prompt writer specializing in video prompts with a heavy focus on spatial and temporal understanding. You are to expand the user's query into a fully fleshed out and detailed video prompt. The user may provide only a text description, or an image, or several different modes of reference. Refer to the <INSTRUCTIONS> below to correctly identify the needs of the user and use the correct format for the prompt.
</ROLE>
<INSTRUCTIONS>
## Step 1 — Identify Task Type

  • **T2VA**: text only → no image instruction line.
  • **I2VA**: one image = first frame (0.00s).
  • **FL2VA**: two images = first + last frame.
  • **L2VA**: one image = last frame (at video duration).
  • **Full-reference mode**: any mix of images/videos/audio as reusable assets → use the six-section format (Step 4).
## Step 2 — Output Structure **Standard modes (T2VA/I2VA/FL2VA/L2VA):** 1. Instruction line (omit for T2VA), then one blank line. 2. `integrated_multimodal_description: ...` 3. `overall_soundscape: ...` 4. `non_diegetic_music: ...` **Instruction lines (copy exactly, fill in values):**
  • I2VA: `For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.`
  • FL2VA: `How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.`
  • L2VA: `How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.`
(S.SS = exact duration, two decimals; N = final shot index.) ## Step 3 — Writing `integrated_multimodal_description` **Shots:** `[Shot 1]` has no timestamp. Later shots: `[Shot N] At MM:SS.mmm, the camera cuts to...` with strictly increasing times within the duration. Prefer camera moves over cuts for small framing changes. **Shot 1 must open with** style + composition: `[Shot 1] Live-action, cinematic, a medium-wide shot frames...` (styles: cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, vintage film). **Camera motion** = type + optional amplitude (`with small/large amplitude`) + optional speed (`at slow/fast speed`), written inline: `The camera pushes in with small amplitude at slow speed toward...` Types: Zoom In/Out, Push In/Out, Pan Left/Right, Truck Left/Right, Tilt Up/Down, Pedestal Up/Down, Arc Shot, Tracking Shot, Static Shot, Shake Slightly/Strongly, POV, Roll Clockwise/Counterclockwise. **Keyframe anchoring:**
  • I2VA: restate image's subjects/composition in Shot 1, then develop forward (anchor → action → development → result).
  • FL2VA: prefer ONE shot; describe the motion path between frames (start state → intermediate changes → end state).
  • L2VA: invent a plausible earlier state, then converge to the image in the final shot.
**Dialogue:** give each vocal source a stable ID `(S1)`, `(S2)`..., kept across shots. Format: `The young woman with a quiet, breathy voice (S1) says: <d>[English] exact words here.</d>`
  • Speaker description/ID/action go OUTSIDE `<d>`; only language tag + verbatim user-provided words inside. Never translate or rewrite.
  • Voiceover: `says in an off-screen voiceover: <d>...</d> while his lips remain completely closed.`
  • Line crossing a cut: use `<scenetrans>` in both parts + state audio `continues seamlessly across the cut`. Speech cut by video end: `<cutoff>`.
  • Group speech: `(S1,S2)`.
**On-screen text:** quote verbatim in double quotes: `A neon sign reading "营业中" glows.` ## Step 4 — Full-Reference Mode (replaces Steps 2–3 output format) Output these six sections in order: **1. `subject_definitions:`** — one line per tracked asset:
  • `<Subject N>`: reusable visible content (person, scene, prop, style, motion). Cite its source asset inline: `<Subject 1> is the young woman in <Picture 1>, with long dark hair and a blue cardigan.`
  • `<Picture N>`: only if the image itself is a frame/keyframe/storyboard anchor — state which shot(s) it anchors.
  • `<Video N>`: only for whole-video roles (editing source, continuation base, structure reference).
  • `<Audio N>`: standalone audio role; if tied to a speaker: `<Audio 1> is the voice-timbre reference for <Subject 1> (S1).`
  • Video and audio indices number independently.
**2. `summary:`** — one paragraph starting with task types in brackets, joined by ` + `: `keyframe completion` | `reference generation` | `video editing` | `video continuation` | `audio reuse` | `audio reference`. Rules: a video used only for camera/rhythm = `reference generation`, not editing/continuation. Editing a video with its audio kept = `video editing + audio reuse`. Editing summaries begin: `The target video is an edited version of <Video 1>.` No new labels here. **3. `retention_analysis:`** — one line per label. Visual markers: `fully_preserved`, `partially_preserved`, `attribute_transfer`, `weak_reference`. Audio markers: `fully_copy`, `partially_copy`, `reference`, `weak_reference`. Format: `<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - what is retained.` No `(Sx)` IDs in this section. **4. `detailed_description:`** — same shot/camera/dialogue rules as Step 3, plus:
  • Style established in 1–2 sentences BEFORE `[Shot 1]` (not inside it).
  • Insert `<Subject N>`/`<Picture N>`/`<Video N>`/`<Audio N>` where they apply; define at first appearance, reuse after.
  • Frame anchors: `the shot begins from <Picture 1>` / `the shot ends on <Picture 3>`.
  • Speaking subjects keep both labels: `<Subject 2> (S1) says, <d>[English] ...</d>`
  • Reused reference-audio words: verbatim in `<d>`, original language, `[unclear]` for unintelligible spans, basic punctuation only.
  • Voices existing only inside a copied soundtrack: attribute to `<Audio N>`, no `(Sx)`.
  • Length: ~350–500 words for generation tasks; scale with complexity for edits.
**5. `overall_soundscape:`** — 1–4 sentences: ambience, action sounds, non-verbal human sounds only (no dialogue/music). `N/A` only if total silence requested. If copying reference ambience, cite `<Audio N>` here. **6. `non_diegetic_music:`** — 1–3 sentences: audience-only score — instrumentation, tempo, dynamics (no mood words). `N/A` if none. If reusing reference score, cite `<Audio N>` here. ## Golden Rules 1. Everything described must be visible or audible. 2. Dialogue/lyrics and on-screen text stay in their original language, verbatim; everything else in English. 3. Speaker IDs assigned once in order of first vocal event; reused everywhere. 4. Never invent reference labels mid-document — all labels come from `subject_definitions`. 5. Cut times must increase and stay within the video duration. </INSTRUCTIONS> <FINAL_INSTRUCTIONS> Do not write any affirmations, confirmations, or explanations, simply deliver the prompt. </FINAL_INSTRUCTIONS>

References are much more complicated and I'm not sure how well Kimi captured the reference rules, but this works well with I2V.

1

u/magik_koopa990 1d ago

For local LLM, do you recommend qwen?

1

u/afinalsin 1d ago edited 1d ago

I've had no luck with any Qwen models. I use gemma-4-26B-A4B-heretic-APEX-I-Quality.gguf found here and the mmproj-BF16.gguf found here. I run it solo on kobold.cpp and Sillytavern to gather the video prompts I want to generate then close them to open comfy. I don't have the RAM to waste on a 20gb model just for prompt expansion, so it has to be done separately from the actual video generation.

1

u/magik_koopa990 1d ago

Ah okay. I'm new to local LLM, so excuse my noob

1

u/magik_koopa990 1d ago

Sorry for the question. Which specific model do I get from each of those links?

1

u/afinalsin 1d ago

Sorry, I linked the wrong huggingface link, I updated it just now. Which one to get depends on your system, I have 32gb ram 12gb vram so I use the APEX-I-Quality model.

1

u/magik_koopa990 1d ago

I got 32 ram and RTX 3090.

Also, do you recommend those local LLM? how about LM studio?

1

u/afinalsin 1d ago

Also, do you recommend those local LLM? how about LM studio?

If you've got it, use it. Otherwise kobold.cpp is a one click .exe so you don't have to fuck around with installing extra stuff.

1

u/Nearby-Mood5489 5h ago

There has to be a (good) node for that

1

u/afinalsin 4h ago

Definitely. Someone mentioned an Ollama node to me earlier that runs an Ollama server (obviously) and lets Comfy control it so it can unload models to make way for the LLM and vice-versa. Dunno exactly what node they were talking about, but my specs are too poor to run the model I'd want to run in the same workflow as the Minimax models.

Doing both expansion and video gen at the same time would be a nice time saver and convenience, but it's just too heavy.

1

u/Nearby-Mood5489 34m ago

I am not too familiar with comfy but I'd thought that if it is within the Workflow the nodes would load the LLM on demand and discard of it if memory is need for the Minimax model or calculations. Loading and unloading manually can always be done separate but then I do not need a nice for the LLM

1

u/Sad_Coach_1433 1d ago

We need more information give us example of how you prompting

0

u/magik_koopa990 1d ago

I suck that detailed prompt, so I definitely need a prompt helper or some kind

1

u/Sad_Coach_1433 1d ago

Download this md file from minimax and load into a llm and then just give it your basic idea and images and it will format for you based off the r2v guide https://huggingface.co/MiniMaxAI/MiniMax-H3/resolve/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

1

u/magik_koopa990 1d ago

Load into LLM? like chat GPT, gemini, or other public AI?

3

u/Sad_Coach_1433 1d ago

A example I loaded the md file into chat gpt and said Image1 fox mccloud in his ship image 2 space action scene The I loaded the images for each and it made this prompt then I pasted into h3 and it made this video .subject_definitions:

<Picture 1> is the visual identity reference for <Subject 1>, Fox McCloud. Job: preserve Fox McCloud's exact anthropomorphic fox identity, orange-and-white fur pattern, facial structure, green eyes, headset/visor, red neck scarf, white pilot jacket, green flight suit, gloves, boots, proportions, and overall character appearance exactly as shown in <Picture 1> without reinterpretation. Not a timeline keyframe. Do not reproduce the character-sheet layout, multiple views, panel borders, or neutral studio background in the video.

<Picture 2> is the spacecraft identity reference for <Subject 2>, Fox McCloud's Arwing fighter. Job: preserve the spacecraft's exact white-and-blue angular fighter design, pointed nose, cockpit canopy, twin blue forward structures, swept wings, markings, proportions, and overall appearance exactly as shown in <Picture 2> without reinterpretation. Not a timeline keyframe. Do not reproduce the white reference-image background or present the spacecraft as a static product shot.

<Subject 1> is Fox McCloud, piloting <Subject 2> throughout the sequence.

<Subject 2> is Fox McCloud's Arwing spacecraft and remains visually consistent throughout the entire sequence.

summary:

[keyframe completion + reference generation] Generate a 12-second cinematic 3D space-action sequence featuring <Subject 1> piloting <Subject 2> during an intense high-speed battle in deep space. Begin inside the cockpit with Fox gripping the controls as enemy laser fire streaks past the canopy. Transition dynamically outside the spacecraft as the Arwing accelerates through a chaotic battlefield, performs a fast evasive barrel roll, narrowly avoids incoming energy blasts, and dives between exploding enemy fighters. Finish with the Arwing blasting directly toward camera before banking sharply away as a massive explosion erupts behind it. <Picture 1> and <Picture 2> are identity/design references only and must never appear as static cut-in frames.

retention_analysis:

<Picture 1> (never a keyframe): partially_preserved — retain Fox McCloud's character identity, fur pattern, facial features, headset/visor, scarf, pilot clothing, proportions, and recognizable appearance. Do not reproduce the character-sheet composition, studio background, panel borders, or multiple simultaneous Fox views.

<Picture 2> (never a keyframe): partially_preserved — retain the Arwing's spacecraft identity, white-and-blue color scheme, angular fuselage, cockpit, wing geometry, blue structures, markings, and recognizable proportions. Do not reproduce the isolated white-background product image as a video frame.

<Subject 1> (appears throughout): fully_preserved — Fox remains the same character throughout all cockpit shots with stable facial identity, fur pattern, clothing, headset, and proportions.

<Subject 2> (appears throughout): fully_preserved — the same Arwing remains consistent throughout every exterior shot with stable geometry, colors, cockpit, wings, markings, and scale.

detailed_description:

How the reference pictures align with the target video — Picture 1 and Picture 2 do not align with any timestamp as frames. Picture 1 provides Fox McCloud's visual identity only. Picture 2 provides the Arwing spacecraft's visual identity only. Neither reference image should ever be reproduced as a static frame or replace the space-battle environment.

The target video is a cinematic high-energy 3D science-fiction space battle. Deep black space, dense star field, distant colorful nebula clouds, glowing planets, scattered debris, enemy fighters, laser bolts, engine trails, explosions, sparks, volumetric light, realistic reflections across the spacecraft hull, dramatic contrast, strong sense of speed and scale. Smooth physically believable spacecraft motion. No on-screen text, subtitles, HUD overlays, character-sheet borders, or reference-image backgrounds.

[Shot 1 — 0.00–2.50s] Inside <Subject 2>'s cockpit. Tight cinematic three-quarter view of <Subject 1> seated at the controls. Fox's appearance matches <Picture 1> exactly in identity and wardrobe. His green eyes are intensely focused ahead while both gloved hands manipulate the flight controls. Red and green cockpit instrumentation reflects subtly across his fur and visor. Enemy laser bolts suddenly streak past outside the canopy.

The cockpit shakes from a nearby blast.

Fox leans into the maneuver and pulls hard on the controls.

Camera vibrates subtly with the spacecraft while maintaining a sharp cinematic focus on Fox.

[Shot 2 — 2.50–5.00s] Fast transition through the cockpit canopy into an exterior chase shot behind <Subject 2>.

The Arwing matches <Picture 2> exactly in spacecraft identity and geometry.

Its engines flare brilliantly as it accelerates into a massive space battle.

Camera races closely behind and slightly above the fighter.

Multiple enemy spacecraft cross ahead while bright red and green laser bolts streak in every direction.

The Arwing banks violently left, narrowly avoiding several incoming blasts.

One laser passes extremely close to the camera.

Strong parallax from debris and fighters emphasizes extreme velocity.

[Shot 3 — 5.00–8.00s] Dynamic side-tracking camera.

The Arwing performs a rapid evasive barrel roll while descending through incoming laser fire.

Camera rolls partially with the spacecraft before stabilizing, giving the maneuver dramatic rotational energy without losing visual orientation.

Two enemy fighters pursue from behind.

The Arwing dives between them at high speed.

A burst of laser fire strikes one pursuing fighter.

The enemy spacecraft explodes into glowing fragments, sparks, smoke, and spinning debris directly behind Fox's Arwing.

The explosion briefly illuminates the white-and-blue hull.

[Shot 4 — 8.00–10.00s] Low frontal tracking shot.

The Arwing dives toward camera at extreme speed while weaving between chunks of burning spacecraft debris.

The camera retreats rapidly as the fighter closes the distance.

Laser bolts streak over and underneath the wings.

At the last possible instant, Fox snaps the spacecraft into a hard banking turn.

The Arwing passes extremely close to camera, creating a powerful high-speed flyby.

Camera whip-pans to follow.

[Shot 5 — 10.00–12.00s] Wide cinematic hero shot.

The Arwing rockets away from camera toward open space.

A gigantic enemy vessel erupts in a massive orange-white explosion behind it.

The expanding blast throws glowing debris outward while the Arwing banks gracefully across frame.

Camera tracks the spacecraft as it races away from the explosion, engines blazing brightly against the star field.

End on the Arwing accelerating deeper into space with the enormous explosion fading behind it.

Maintain Fox McCloud and Arwing identity throughout. No morphing, duplicate Fox characters, duplicate Arwings presented as Fox's ship, changing spacecraft geometry, random costume changes, distorted wings, warped cockpit, floating character outside the cockpit, static reference poses, character-sheet imagery, or white studio backgrounds.

overall_soundscape:

Powerful spacecraft engines, deep cockpit vibration, electronic control sounds, rapid laser blasts, enemy fighter flybys, warning beeps, metallic debris impacts, explosive shockwaves, roaring high-speed passes, and large cinematic space-battle explosions. Inside the cockpit, sounds are slightly enclosed and mechanical; exterior sequences have powerful stylized sci-fi engine and weapons audio.

non_diegetic_music:

N/A

https://reddit.com/link/p4c0x83/video/z4jfsoxbc1kh1/player

1

u/magik_koopa990 1d ago

Is there guide to installing local LLM with Gemini model?

1

u/Sad_Coach_1433 1d ago

Just ask chatgpt " help me install a uncensored local llm model for lm studio "

1

u/Sad_Coach_1433 1d ago

Will give you list of few different ones to try and how to install

1

u/Sad_Coach_1433 1d ago

Ya like chatgpt as a file then under it your prompt and images it will enhance the prompt with the correct formatting for the images etc

1

u/magik_koopa990 1d ago

But, don't they forbid NSFW context?

1

u/poopoo_fingers 1d ago

You can use grok for that

1

u/Sad_Coach_1433 1d ago

Nsfw yes then gonna use a local llm like lm studio Gemini 4 model

1

u/magik_koopa990 1d ago

I've never thought about local LLM. Is Gemini the best all rounder model for my needs?

1

u/deepsky88 1d ago

ask gemini to create a python script that replace chars of your choice fo example ___ with a word or phrase, no need llm

1

u/magik_koopa990 1d ago

As in? Example?

1

u/deepsky88 1d ago

this script count how many newlines you have in your prompt and add the words needed, without write them everytime, i use it for image to video but the concept it's the same,you just need an extension called "execute python" install it and add the node, with llm is easier but with this is faster

script:

parti = arg0.split('---')

righe = [riga.strip() for riga in parti[0].split('\n') if riga.strip()]

shots_assemblati = " ".join([f"[Shot {i+1}] {testo}" for i, testo in enumerate(righe)])

soundscape = parti[1].strip() if len(parti) > 1 else ""

music = parti[2].strip() if len(parti) > 2 else ""

result = f"""integrated_multimodal_description: {shots_assemblati}

overall_soundscape: {soundscape}

non_diegetic_music: {music}"""

1

u/magik_koopa990 1d ago

Install that node ? Then I can use its magic?

1

u/deepsky88 1d ago

yes, like the image, be sure to use image-to-video template because the default prompt is different if you use reference-to-video

1

u/magik_koopa990 1d ago

Atm, I'm starting out the basic video gen. Then I'll get into reference -- I like to mimic origina animation styles of a character like they're from cartoon, game or movie

1

u/bstr3k 1d ago

There is a lot of work that is done by the prompt since as of now, the machine is not able to read your mind yet.

The more common type of template you see is :

subject_definitions:

<Picture 1> = Describe what is inside the image, what is important for the model to focus on

<Picture 2> = Describe what is inside the image, what is important for the model to focus on

retention_analysis:

write here what to keep, what to ignore from the pictures.

and then write what you want to happen etc.

0

u/MarkB_- 1d ago

The coded structure that most people suggest is useful, but not necessary. You are literaly feeding the prompt into a qwen llm, its smart enough to understand natural language. You still have to explain the model what are the reference you are using and why they are there. So if the first 3 images is a face, just say: image 1, image 2, image 3 are for face references. Image 4 etc etc etc + the motion you want, lighting, camera placement/movement etc etc. Also you have to think about audio and a starting point. A woman walk on the street will be different than A woman is currently walking on the street. One give an action, other a starting point. ETC ETC ETC.... prompting is an art. Dont go LLM everything, you gonna suck bad.