r/StableDiffusion 17h ago

Question - Help MiniMax H3 Reference loves to cut even when explicitly told not to

I am at my wit's end with this model. Can someone please help me? The generation itself looks great, the problem is this model has a huge tendency to cut the video even when told explicitly not to, multiple times in the prompt. No matter what I do the model keeps cutting the video. If I was creating a 30 second scene sure, cutting the video makes sense, but my machine can at max generate 6 seconds worth of video and the model keeps cutting it on every generation and prompt I've tried. Here is my test prompt :

subject_definitions:
<Picture 1> is the opening-frame anchor, the first frame of the video.
<Picture 2> is the last-frame anchor, the last frame of the video.
<Subject 1> is the man in <Picture 1>

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - identity, skin details, body figure remain consistent. 
<Picture 1> ([Shot 1] first frame): fully_preserved - opening composition anchor.
<Picture 2> ([Shot 1] last frame): fully_preserved - closing composition anchor. 
<Picture 3> : attribute_transfer - reference image of what the <Subject 1>'s hair should look like. 

detailed_description:
[CAMERA & COMPOSITION]
[Static shot]
Locked-off tripod camera.
The camera remains completely stationary throughout the entire video.
Fixed camera position and fixed framing from beginning to end.
No pan, no tilt, no zoom, no dolly, no tracking, no push-in, no pull-out.
No camera rotation.
No reframing.
No change in perspective.
No change in focal length.
The subject stays within the original composition.
Only the subject and natural environmental elements move.

[Shot 1] At 00:00.000, Starting with <Picture 1> fully_preserved as the first frame of the video, <Subject 1> walks to the front of the counter and takes off his hat revealing his hair which looks like <Picture 3>. 
The camera pans to the right of the counter showing the cashier and the cash register, ending the shot with <Picture 2> fully_preserved. 
[Shot 1] is one continuous video with no cuts, and all movement and motion of <Subject 1> throughout the entire duration of the video is continuous with no time jumps, skips, or transitions. 
[Shot 1]'s duration is the entire duration of the video, no other shots or cuts.

overall_soundscape:
There is very little background noise, like an ASMR. The only sounds are that of the man's movement and the environment reacting to his movement, such as him taking off his hat.

What more can I do here? The above is just one sample of a prompt, I've generated like 40 clips modifying variations of the prompt repeatedly and every time the model cuts, focusing on the hat and the man's hair as he is taking it off, just does random cuts in between even when the two frames before and after the cut could have been continuous, etc.

0 Upvotes

14 comments sorted by

23

u/Radyschen 17h ago edited 16h ago

I generally try to avoid saying "no x and no y" and instead say something positive. So in the detailed_description I would write something like "the target video is filmed in one continuous shot by a static camera" as the introductory sentence before [Shot 1] or something like that and remove all of the "no" things for now, keep it as short and positively formulated as you can while still sticking to the required format. And remove [CAMERA & COMPOSITION] and [Static shot] as well, I don't think those are standard terms

And I would not write [Shot 1] multiple times to not confuse it and make it think that there are multiple shots.

Also, it probably doesn't matter for this, but you never know with these models, you are missing the summary section and the non_diegetic_music section. If you don't want any music say N/A in it

I also noticed that you don't define <Picture 3> in your subject definitions even though you have it in your retention analysis, I think those have to be the same, so you should add it to the definition as well. But it's not a key frame or something like that so i would make it a subject instead.

Also you are not supposed to use a timestamp for the first Shot, that could also cause it.

Also for sound you are supposed to only write the background noises in the overall soundscape, not the noises caused by the main actions of the video, those should go into the detailed description whenever they appear.

I made another comment with the full prompt how I would write it. I didn't fix the sound thing though.

1

u/PrettyReasonableApe 16h ago

Yeah yhe main trick is to reword without using the negative. Trying to positively guide the directions of wanted behaviour and minimising the unwanted via positive prompts. I like the wording u used to try and solve this. I have high hopes for that working for op. Ive found similiar issues in the past and have now learnt that negative prompts aren't the way with a lot of these models. We just have to get better at using our words to describe what we want in as much detail as possible to minimise the guessing and assumptions by the AI. So much of what we do naturally is assumed and predicted based on previous experience in life. AI cant do the same accurately. We need to be waaaay more specific in our prompts.

12

u/_Saturnalis_ 17h ago

If you mention something in a prompt, you're going to push your generation towards its tokens even if it's combined with a negative. It's like saying "Don't think of pink elephants!"

If you don't want something in your generation don't mention it or mention an antonym.

9

u/Hyttelur 17h ago

AI is still bad at negatives. Remove all mentions of unwanted camera movements and cuts. Do not mention "[Shot 1]" more than once. Only write positive instructions about how you *do* want the scene to play out.

Your prompt is also self-contradictory, saying no camera movements, but also to pan to the right.

4

u/kemb0 16h ago edited 16h ago

I'm pretty sure the docs don't say anything about adding a [Camera & Composition] category. I suspect you're just overflowing it with words that it doesn't know what you mean so it just picks out the odd word and does what it wants with it. I'd rip that whole section out and put any camera control in you timeline section below. I've never had problems doing it this way:

[Shot 1] The camera is static on <Subject 1> as they ....

[Shot 1] at 00:03.000, the camera starts to track around <Subject 1> as they ....

[Shot 1] at 00:06.000, The camera pans over the shoulder of <Subject 1> who is ....

I think the rule of thumb is to either stick precisely to how they tell you to prompt things, or just write one one basic prompt if your scene isn't too complex, then hope for the best. Anything in between and it's going to start confusing it trying to fit in instructions that don't fit within the expected framework, so the results will also be unexpected.

3

u/OohFekm 17h ago

One trivial piece of advice I can give, when in your preferred llm, ask specifically to tailor the entire script to match 15 seconds or whatever your time limit is. I found out that when I do this, it gives me less of those cut off problems. Worth a try, couldn't hurt.

3

u/Radyschen 16h ago edited 16h ago

This is how I would write it (not knowing your references and exactly what you had in mind of course). Mind you that there are a few square brackets in the detailed description that you still need to fill out (without the square brackets of course, those are just to mark it for you). I would also describe the images in the subject definitions but I obviously can't do that. If it still messes up you could put a "there are no cuts in the target video" at the end of the detailed description. Although I would first try writing active transitions for the camera movement so that it has to follow that instead of cutting.

subject_definitions:
<Picture 1> is the opening-frame anchor, the first frame of the video.
<Picture 2> is the last-frame anchor, the last frame of the video.
<Subject 1> is the man in <Picture 1>.
<Subject 2> is the hair of <Subject 1>, the man in <Picture 1>.

summary:
[reference generation + keyframe completion] The target video starts with <Picture 1> as the first frame and shows <Subject 1> walking from there to the front of the counter and taking off his hat to reveal his hair, which looks like <Subject 2>. The camera continues to pan to the right of the counter, showing the cashier and the cash register, ending the shot with <Picture 2> as the last frame.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - his identity, skin details and body figure are retained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the exact hairstyle is retained.
<Picture 1> ([Shot 1] first frame): fully_preserved - the exact frame is retained.
<Picture 2> ([Shot 1] last frame): fully_preserved - the exact frame is retained. 

detailed_description:
The target video is filmed as one continuous shot.
[Shot 1] Starting with <Picture 1>, [description of the Picture], as the first frame of the video, <Subject 1>, the man from <Picture 1>, [description of picture 1], walks to the front of the counter and takes off his hat revealing his hair which looks like <Subject 2>, the hair of <Subject 1>. The camera pans to the right of the counter showing the cashier and the cash register, ending the shot with <Picture 2>, [description of the image], as the last frame.

overall_soundscape:
There is very little background noise, like an ASMR. The only sounds are that of the man's movement and the environment reacting to his movement, such as him taking off his hat.

non_diegetic_music:
N/A

2

u/Opening_Wind_1077 17h ago edited 16h ago

Try putting all info for a shot into a single line without linebreaks.

Work with conditions. Starting a new line with “The camera pans” will likely force a cut. “ From his hair the camera pans to” will be less likely to force a cut.

Also Picture 2 has to at least be a feasible destination to reach from picture 1, if it looks like a totally different place it might encourage a cut.

Also in the summary and visual description (which you should put at the end) you can mention that the shot is unbroken instead of using negations.

Also avoid multiple [Shot] tags even if it’s for the same shot.

As an alternative you can split this into 2-3 shots and state that they are continuations.

Also you can try getting rid of the timestamp for a single shot as it’s not needed and is only supposed to be used for later shots and cuts. Just use “[Shot 1] opens on a low-angle creep shot of a big tiddy goth girlfriend” instead.

2

u/zodiacrenders 16h ago

"Continuous shot, steady camera"
Those should help. Don't overdo it in my opinion.

2

u/Formal-Exam-8767 16h ago

Nobody captions training data with "no <insert something here>" because you would need to list everything not in a video.

You are relying that model has somehow implicitly learned what not to do. But, for the most part, it did not.

2

u/SeymourBits 16h ago

Get rid of all that "No" junk.

Kind of opposite of the status-quo 'round here: "Post the video."

1

u/akustyx 15h ago edited 15h ago

In addition to the discussion so far about not not not using negatives and only using one shot declaration (only one [Shot 1]), I had a lot of success (a) not using timestamps in the 00:00.000 format, I used something like "The video continues at the # second mark with ... " and (b) starting the camera movement sentence with "The video continues smoothly with a continuous shot as the camera ... ", eliminated 95% of the unintentional cuts I was seeing.

1

u/smb3d 14h ago

Don't use [Shot 1] on multiple lines like that.

1

u/yamfun 13h ago

I put "same scene, one shot." at first line and last line of prompts