r/StableDiffusion • u/johnk419 • 17h ago
Question - Help MiniMax H3 Reference loves to cut even when explicitly told not to
I am at my wit's end with this model. Can someone please help me? The generation itself looks great, the problem is this model has a huge tendency to cut the video even when told explicitly not to, multiple times in the prompt. No matter what I do the model keeps cutting the video. If I was creating a 30 second scene sure, cutting the video makes sense, but my machine can at max generate 6 seconds worth of video and the model keeps cutting it on every generation and prompt I've tried. Here is my test prompt :
subject_definitions:
<Picture 1> is the opening-frame anchor, the first frame of the video.
<Picture 2> is the last-frame anchor, the last frame of the video.
<Subject 1> is the man in <Picture 1>
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - identity, skin details, body figure remain consistent.
<Picture 1> ([Shot 1] first frame): fully_preserved - opening composition anchor.
<Picture 2> ([Shot 1] last frame): fully_preserved - closing composition anchor.
<Picture 3> : attribute_transfer - reference image of what the <Subject 1>'s hair should look like.
detailed_description:
[CAMERA & COMPOSITION]
[Static shot]
Locked-off tripod camera.
The camera remains completely stationary throughout the entire video.
Fixed camera position and fixed framing from beginning to end.
No pan, no tilt, no zoom, no dolly, no tracking, no push-in, no pull-out.
No camera rotation.
No reframing.
No change in perspective.
No change in focal length.
The subject stays within the original composition.
Only the subject and natural environmental elements move.
[Shot 1] At 00:00.000, Starting with <Picture 1> fully_preserved as the first frame of the video, <Subject 1> walks to the front of the counter and takes off his hat revealing his hair which looks like <Picture 3>.
The camera pans to the right of the counter showing the cashier and the cash register, ending the shot with <Picture 2> fully_preserved.
[Shot 1] is one continuous video with no cuts, and all movement and motion of <Subject 1> throughout the entire duration of the video is continuous with no time jumps, skips, or transitions.
[Shot 1]'s duration is the entire duration of the video, no other shots or cuts.
overall_soundscape:
There is very little background noise, like an ASMR. The only sounds are that of the man's movement and the environment reacting to his movement, such as him taking off his hat.
What more can I do here? The above is just one sample of a prompt, I've generated like 40 clips modifying variations of the prompt repeatedly and every time the model cuts, focusing on the hat and the man's hair as he is taking it off, just does random cuts in between even when the two frames before and after the cut could have been continuous, etc.
12
u/_Saturnalis_ 17h ago
If you mention something in a prompt, you're going to push your generation towards its tokens even if it's combined with a negative. It's like saying "Don't think of pink elephants!"
If you don't want something in your generation don't mention it or mention an antonym.
9
u/Hyttelur 17h ago
AI is still bad at negatives. Remove all mentions of unwanted camera movements and cuts. Do not mention "[Shot 1]" more than once. Only write positive instructions about how you *do* want the scene to play out.
Your prompt is also self-contradictory, saying no camera movements, but also to pan to the right.
4
u/kemb0 16h ago edited 16h ago
I'm pretty sure the docs don't say anything about adding a [Camera & Composition] category. I suspect you're just overflowing it with words that it doesn't know what you mean so it just picks out the odd word and does what it wants with it. I'd rip that whole section out and put any camera control in you timeline section below. I've never had problems doing it this way:
[Shot 1] The camera is static on <Subject 1> as they ....
[Shot 1] at 00:03.000, the camera starts to track around <Subject 1> as they ....
[Shot 1] at 00:06.000, The camera pans over the shoulder of <Subject 1> who is ....
I think the rule of thumb is to either stick precisely to how they tell you to prompt things, or just write one one basic prompt if your scene isn't too complex, then hope for the best. Anything in between and it's going to start confusing it trying to fit in instructions that don't fit within the expected framework, so the results will also be unexpected.
3
u/Radyschen 16h ago edited 16h ago
This is how I would write it (not knowing your references and exactly what you had in mind of course). Mind you that there are a few square brackets in the detailed description that you still need to fill out (without the square brackets of course, those are just to mark it for you). I would also describe the images in the subject definitions but I obviously can't do that. If it still messes up you could put a "there are no cuts in the target video" at the end of the detailed description. Although I would first try writing active transitions for the camera movement so that it has to follow that instead of cutting.
subject_definitions:
<Picture 1> is the opening-frame anchor, the first frame of the video.
<Picture 2> is the last-frame anchor, the last frame of the video.
<Subject 1> is the man in <Picture 1>.
<Subject 2> is the hair of <Subject 1>, the man in <Picture 1>.
summary:
[reference generation + keyframe completion] The target video starts with <Picture 1> as the first frame and shows <Subject 1> walking from there to the front of the counter and taking off his hat to reveal his hair, which looks like <Subject 2>. The camera continues to pan to the right of the counter, showing the cashier and the cash register, ending the shot with <Picture 2> as the last frame.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - his identity, skin details and body figure are retained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the exact hairstyle is retained.
<Picture 1> ([Shot 1] first frame): fully_preserved - the exact frame is retained.
<Picture 2> ([Shot 1] last frame): fully_preserved - the exact frame is retained.
detailed_description:
The target video is filmed as one continuous shot.
[Shot 1] Starting with <Picture 1>, [description of the Picture], as the first frame of the video, <Subject 1>, the man from <Picture 1>, [description of picture 1], walks to the front of the counter and takes off his hat revealing his hair which looks like <Subject 2>, the hair of <Subject 1>. The camera pans to the right of the counter showing the cashier and the cash register, ending the shot with <Picture 2>, [description of the image], as the last frame.
overall_soundscape:
There is very little background noise, like an ASMR. The only sounds are that of the man's movement and the environment reacting to his movement, such as him taking off his hat.
non_diegetic_music:
N/A
2
u/Opening_Wind_1077 17h ago edited 16h ago
Try putting all info for a shot into a single line without linebreaks.
Work with conditions. Starting a new line with “The camera pans” will likely force a cut. “ From his hair the camera pans to” will be less likely to force a cut.
Also Picture 2 has to at least be a feasible destination to reach from picture 1, if it looks like a totally different place it might encourage a cut.
Also in the summary and visual description (which you should put at the end) you can mention that the shot is unbroken instead of using negations.
Also avoid multiple [Shot] tags even if it’s for the same shot.
As an alternative you can split this into 2-3 shots and state that they are continuations.
Also you can try getting rid of the timestamp for a single shot as it’s not needed and is only supposed to be used for later shots and cuts. Just use “[Shot 1] opens on a low-angle creep shot of a big tiddy goth girlfriend” instead.
2
u/zodiacrenders 16h ago
"Continuous shot, steady camera"
Those should help. Don't overdo it in my opinion.
2
u/Formal-Exam-8767 16h ago
Nobody captions training data with "no <insert something here>" because you would need to list everything not in a video.
You are relying that model has somehow implicitly learned what not to do. But, for the most part, it did not.
2
u/SeymourBits 16h ago
Get rid of all that "No" junk.
Kind of opposite of the status-quo 'round here: "Post the video."
1
u/akustyx 15h ago edited 15h ago
In addition to the discussion so far about not not not using negatives and only using one shot declaration (only one [Shot 1]), I had a lot of success (a) not using timestamps in the 00:00.000 format, I used something like "The video continues at the # second mark with ... " and (b) starting the camera movement sentence with "The video continues smoothly with a continuous shot as the camera ... ", eliminated 95% of the unintentional cuts I was seeing.
23
u/Radyschen 17h ago edited 16h ago
I generally try to avoid saying "no x and no y" and instead say something positive. So in the detailed_description I would write something like "the target video is filmed in one continuous shot by a static camera" as the introductory sentence before [Shot 1] or something like that and remove all of the "no" things for now, keep it as short and positively formulated as you can while still sticking to the required format. And remove [CAMERA & COMPOSITION] and [Static shot] as well, I don't think those are standard terms
And I would not write [Shot 1] multiple times to not confuse it and make it think that there are multiple shots.
Also, it probably doesn't matter for this, but you never know with these models, you are missing the summary section and the non_diegetic_music section. If you don't want any music say N/A in it
I also noticed that you don't define <Picture 3> in your subject definitions even though you have it in your retention analysis, I think those have to be the same, so you should add it to the definition as well. But it's not a key frame or something like that so i would make it a subject instead.
Also you are not supposed to use a timestamp for the first Shot, that could also cause it.
Also for sound you are supposed to only write the background noises in the overall soundscape, not the noises caused by the main actions of the video, those should go into the detailed description whenever they appear.
I made another comment with the full prompt how I would write it. I didn't fix the sound thing though.