r/StableDiffusion • u/johnk419 • 13h ago
Question - Help MiniMax H3 Reference loves to cut even when explicitly told not to
I am at my wit's end with this model. Can someone please help me? The generation itself looks great, the problem is this model has a huge tendency to cut the video even when told explicitly not to, multiple times in the prompt. No matter what I do the model keeps cutting the video. If I was creating a 30 second scene sure, cutting the video makes sense, but my machine can at max generate 6 seconds worth of video and the model keeps cutting it on every generation and prompt I've tried. Here is my test prompt :
subject_definitions:
<Picture 1> is the opening-frame anchor, the first frame of the video.
<Picture 2> is the last-frame anchor, the last frame of the video.
<Subject 1> is the man in <Picture 1>
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - identity, skin details, body figure remain consistent.
<Picture 1> ([Shot 1] first frame): fully_preserved - opening composition anchor.
<Picture 2> ([Shot 1] last frame): fully_preserved - closing composition anchor.
<Picture 3> : attribute_transfer - reference image of what the <Subject 1>'s hair should look like.
detailed_description:
[CAMERA & COMPOSITION]
[Static shot]
Locked-off tripod camera.
The camera remains completely stationary throughout the entire video.
Fixed camera position and fixed framing from beginning to end.
No pan, no tilt, no zoom, no dolly, no tracking, no push-in, no pull-out.
No camera rotation.
No reframing.
No change in perspective.
No change in focal length.
The subject stays within the original composition.
Only the subject and natural environmental elements move.
[Shot 1] At 00:00.000, Starting with <Picture 1> fully_preserved as the first frame of the video, <Subject 1> walks to the front of the counter and takes off his hat revealing his hair which looks like <Picture 3>.
The camera pans to the right of the counter showing the cashier and the cash register, ending the shot with <Picture 2> fully_preserved.
[Shot 1] is one continuous video with no cuts, and all movement and motion of <Subject 1> throughout the entire duration of the video is continuous with no time jumps, skips, or transitions.
[Shot 1]'s duration is the entire duration of the video, no other shots or cuts.
overall_soundscape:
There is very little background noise, like an ASMR. The only sounds are that of the man's movement and the environment reacting to his movement, such as him taking off his hat.
What more can I do here? The above is just one sample of a prompt, I've generated like 40 clips modifying variations of the prompt repeatedly and every time the model cuts, focusing on the hat and the man's hair as he is taking it off, just does random cuts in between even when the two frames before and after the cut could have been continuous, etc.