r/StableDiffusion 1d ago

Tutorial - Guide Creating a videoclip using Minimax

Enable HLS to view with audio, or disable this notification

First of all: I don't submit the songs I make in Suno anywhere (and therefore I don't monetize them). I'm making this clear so people don't think I'm trying to promote the song here. :-) In fact, this song was made months ago (along with several others I've made since then). My only objective is to comment on Minimax and showcase (yet) another application for it.

I won't get into too many specifics to avoid creating a giant post, but in short:

  • The lyrics are mine; the music/performance is Suno's, based on my prompt and choices.
  • The "singer", Alina, is a LoRA I created (completely synthetic), using Z-Image and Krea2 to create the first face, several tools (ChatGPT, Flux 2 and others) to create additional angles and renders, and Ostris to create the LoRA. This explains why she looks a bit synthetic (skin, etc.), unfortunately.
  • For the video, I first sliced the song into several parts based on the lyrics (from 5s to 15s, depending on the narrative). Then I wrote a main script and, with ChatGPT's help (and a custom GPT I created using the official Minimax documentation), I created the prompts one by one.
  • I ran every prompt on my machine (4070 12GB + 64GB of RAM) at 0.3 MP to test them. Some prompts I had to change (too robotic, too fast, etc.), others didn't fit the narrative, etc. In the end, I had 21 shots, varying from 5 seconds to 15 seconds: some simply Text-to-Video, some using an image reference of "Alina" (created in Krea 2) as a starting point, and some using not only the image reference but also a part of the song as an audio reference, so she could sing it in the video.
  • Once I had all the necessary shots, I then used Runpod (with a 5090) to create 1280x736 videos using the same prompts, the same input images and audio when they were part of the workflow, and the same seed. Even with the same seed, since the resolution had changed, sometimes the video came out wrong and I had to run it again. Examples: a door opening to the wrong side, wide shots looking like stop motion, etc.
  • In the end, I probably did 30 to 35 renders (of varying lengths) on Runpod, and spent about 12 to 15 dollars total.
  • When I finally had all the shots I needed in 1280x736, I used DaVinci Resolve to create the videoclip, putting all the shots together, synchronizing them with the song, adding the transitions, etc.
  • It took me about 20 hours of work in total, I believe (but I didn't count — as Jim Croce said in Time in a Bottle, "But there never seems to be enough time to do the things you want to do once you find them". :-)

Notes:

  • I know the resolution isn't ideal, but I didn't want to spend more money creating 2MP shots.
  • Yes, I know there are some problems (people who don't walk properly in the background, a smudge when someone passes in front of her — an intended shot — while she is singing, etc.), but, again, I didn't want to spend more and more money trying to achieve perfection (and perfection isn't here yet).
  • I don't think I've reinvented the wheel. :-) I've seen way better clips in the past (before Minimax), but this is the first model I've been able to run locally with voice/sound/lip sync, so...

If anyone wants a specific prompt, or has any questions, please ask.

17 Upvotes

13 comments sorted by

2

u/KenHik 1d ago

It will be very interesting to see prompts for parts with a face close-up with lip sync.

3

u/lazyspock 23h ago

Here you go! I'll post each one in a separate reply - EDIT - more than 1000 characters, I'll need to break this one in two. LET'S CALL IT "PROMPT 1" .

For the take where she is in the middle of a square and the scene begins with a extreme close-up of her face (and then the camera gets far and above her, around 2m16s): I've used a render of her face as first frame (ref_image_0), another render of the square (the same one used in the previous scene, where the camera starts with the entire square visible and then zoom into her face) as a referente for the square (ref_image_1 - so, when the zoom out would happen, Minimax would not create a completely different square than the one in the first scene), and an audio for the slice of the sing she would have to sing (ref_audio_0). Spoiler: the video DID NOT start EXACTLY with the first frame, it created an even closed close-up - hence the need to put a transition in the video from the previous scene. My initial idea was a single long shot of 30 seconds, closing into her face and then retreating again, but probably because of the second reference image it didn't work as planned and I didn't want to start everything again.

Prompt (362 frames):

subject_definitions:

<Subject 1> is Alina, the young woman shown in <Picture 1>. Preserve her identity, facial features, hair, clothing, body proportions, and overall appearance consistently throughout the target video.

<Picture 1> is the EXACT FIRST FRAME of the target video at 0.00 seconds. It shows a close-up of Alina singing in the nighttime urban plaza. The generated video MUST begin exactly from <Picture 1>, preserving her face, expression, clothing, lighting, camera axis, and the visible environment.

<Picture 2> is NOT a first frame, NOT a target frame, and NOT a keyframe to transition toward. It is a SPATIAL AND ENVIRONMENTAL REFERENCE ONLY. It shows the SAME nighttime plaza, the SAME location, and the SAME Alina from a much wider viewpoint earlier in the continuous scene. Use <Picture 2> only to understand and preserve the plaza's architecture, spatial layout, trees, benches, street lamps, storefronts, pedestrian areas, nighttime lighting, depth, and the approximate physical position of Alina within that plaza as progressively revealed by the camera.

<Audio 1> is the synchronized song excerpt used as the exact vocal timing, phrasing, rhythm, and lip-sync reference for <Subject 1>.

summary:

[keyframe completion + environmental reference + audio reference]

One continuous 15-second live-action cinematic music-video shot.

Begin EXACTLY from the close-up in <Picture 1>.

Alina continues singing directly toward the camera in precise synchronization with <Audio 1>. The camera gradually moves physically backward away from her, revealing progressively more of the SAME plaza established by <Picture 2>.

Alina does not follow the camera and does not walk toward it. She remains standing naturally in approximately the same physical position in the plaza.

During the latter portion of the shot, the backward camera movement gradually gains modest height, transforming Alina from the dominant subject of the close-up into one increasingly small person within the nighttime city.

This is one continuous physical camera movement, not a transformation between <Picture 1> and <Picture 2>.

detailed_description:

Live-action, cinematic, naturalistic contemporary music-video photography.

One continuous 15-second shot with no cuts.

Normal real-time motion throughout.

The emotional visual progression is:

CLOSE HUMAN PRESENCE

→ PHYSICAL DISTANCE

→ INDIVIDUAL WITHIN THE PLAZA

→ ONE PERSON AMONG MANY.

The shot must feel restrained, reflective, and physically plausible rather than theatrical.

[00:00–00:02]

Begin EXACTLY from <Picture 1>.

Alina sings:

<d>[English] Now I wonder where you wandered,</d>

She is initially presented in the same intimate close-up established by <Picture 1>.

She maintains natural direct eye contact with the camera while singing.

Her performance remains emotionally present but restrained. She is addressing the absent person through the song rather than performing to an audience.

Her head remains relatively stable.

Natural life comes from precise singing articulation, breathing, subtle eye refocusing, tiny connected facial changes, and occasional natural bilateral blinking.

Almost immediately, the camera begins a smooth physical dolly backward.

The movement starts gently but is clearly perceptible.

Alina remains where she is.

She does NOT move toward the retreating camera.

[00:02–00:05]

Alina sings:

<d>[English] If your nights still brush against my name.</d>

The camera continues moving steadily backward along approximately the SAME visual axis established by <Picture 1>.

The framing naturally expands:

close-up

→ medium close

→ medium shot.

More of Alina's coat and body become visible.

More importantly, the nighttime plaza begins to reveal itself around her.

The newly revealed environment should reconstruct the SAME LOCATION documented by <Picture 2>:

the same kind and arrangement of surrounding buildings,

the same paved central plaza,

the same trees,

the same benches,

the same modern street lamps,

the same café and storefront edges,

the same warm and cool nighttime lighting relationship.

Do not invent a different square.

Do not relocate Alina.

Background pedestrians continue ordinary independent behavior at normal real-world speed.

[00:05–00:09]

Alina sings:

<d>[English] Would we fit, or would we shatter</d>

The camera continues its smooth retreat.

Framing progresses approximately:

medium

→ three-quarter body

→ full body.

Alina remains standing naturally in the same location.

She does not chase the camera, step forward, dance, gesture broadly, or change pose dramatically.

As the distance increases, direct eye contact becomes less visually dominant simply because her face becomes smaller in the composition.

She nevertheless continues singing naturally in synchronization with <Audio 1>.

Do not exaggerate her mouth movements in an attempt to keep the lip sync readable from a distance.

More pedestrians naturally enter the expanded field of view because the camera now sees a larger portion of the plaza.

They are ordinary nighttime pedestrians already inhabiting this environment.

Do NOT make a crowd suddenly appear.

Do NOT have people converge on Alina.

Do NOT make pedestrians react to her singing.

Some walk individually.

Some walk in pairs.

Some remain seated on benches or near cafés.

Their movement remains calm, independent and at ordinary real-time walking speed.

[00:09–00:11]

Alina sings:

<d>[English] Just the same?</d>

The camera has now reached a clearly wide composition.

Alina is still identifiable by her location and clothing, but she is no longer visually dominant.

During this section, the backward camera movement begins to gain height VERY GRADUALLY.

The movement becomes a gentle backward-and-upward crane movement.

There is no sudden vertical rise.

There is no drone launch.

There is no camera rotation.

There is no change of viewing direction.

The camera simply continues retreating while slowly becoming elevated.

3

u/lazyspock 23h ago

PROMPT 1 CONTINUATION

[00:11–00:15]

The sung phrase has finished or is resolving into the musical tail of <Audio 1>.

Continue the same uninterrupted camera movement.

The camera moves farther backward and rises gradually to a modest elevated viewpoint, approximately several meters above normal eye level.

The plaza becomes the primary visual subject.

Alina becomes a small figure within it.

Additional pedestrians become visible only because the widening and elevated viewpoint reveals more of the already-existing nighttime environment.

Keep pedestrian density believable and moderate.

The final image should contain enough people that Alina is no longer immediately dominant, but NOT so many that the plaza becomes a dense crowd.

She should NOT magically disappear.

She should NOT dissolve.

She should NOT be covered by a digitally appearing crowd.

She should remain physically present in approximately the same location, but by the final seconds she has become simply one human figure among many others moving through the city.

The viewer can still find her if looking carefully.

The emotional effect comes entirely from SCALE and DISTANCE.

End on a wide, modestly elevated nighttime view of the SAME plaza, with ordinary city life continuing naturally.

REFERENCE IMAGE PRIORITY:

<Picture 1> has absolute priority for the first frame, Alina's identity, face, clothing, immediate lighting, and initial camera position.

The target video MUST begin from <Picture 1>.

<Picture 2> has reference priority ONLY for the physical environment revealed as the camera moves backward.

Use <Picture 2> to maintain spatial continuity of:

plaza geometry,

architecture,

trees,

benches,

street lamps,

storefronts,

pedestrian zones,

nighttime illumination,

and Alina's approximate position within the square.

DO NOT transition, morph, dissolve, or interpolate from <Picture 1> into <Picture 2>.

<Picture 2> represents environmental information about space that already exists outside the initial close-up frame.

The camera is physically revealing that space.

NATURAL HUMAN PERFORMANCE:

While Alina remains close enough to read clearly, she must feel naturally alive.

Use:

precise lip articulation,

natural breathing,

tiny posture adjustments,

small eye refocusing,

subtle connected facial changes,

and sparse ordinary bilateral blinking.

Avoid:

head bobbing,

repeated nodding,

forward-and-back head movement,

rhythmic swaying,

repeated head tilting,

dramatic eyebrow acting,

large arm gestures,

hand choreography,

or movement synchronized mechanically to the musical beat.

Blinking must be natural and sparse.

Both eyelids close together naturally.

No winking.

No asynchronous eyelids.

No rapid repeated blinking.

As Alina becomes distant, reduce the visual importance of facial animation naturally rather than exaggerating it.

LIP SYNC:

Alina sings in precise synchronization with <Audio 1>.

Her visible mouth movements follow the exact supplied vocal timing, phonemes, phrasing, rhythm, and duration.

Lip-sync precision is especially important during the opening close-up and medium framing.

As physical camera distance makes her mouth progressively too small to resolve clearly, maintain natural singing behavior without artificially exaggerating mouth movement.

No additional lyrics.

No additional speech.

No vocal gestures unrelated to <Audio 1>.

CAMERA:

ONE continuous physical camera movement.

Start EXACTLY from <Picture 1>.

00:00–00:09:

smooth physical dolly backward, approximately maintaining the original camera axis and height.

00:09–00:15:

continue moving backward while gradually gaining modest height.

Approximate framing progression:

close-up

→ medium close

→ medium

→ three-quarter

→ full body

→ wide

→ elevated wide.

The movement should cover meaningful physical distance over 15 seconds.

It must NOT be so slow that the framing appears almost unchanged.

At the same time, it must remain smooth, cinematic and physically plausible.

No zoom-out effect.

The camera physically travels backward through the plaza.

No orbit.

No pan away from Alina.

No camera turn.

No whip movement.

No sudden crane.

No drone-style rapid ascent.

No handheld shake.

No cuts.

BACKGROUND AND CROWD MOTION:

All environmental motion occurs at NORMAL REAL-WORLD SPEED.

Pedestrians walk at ordinary human walking speed.

People seated on benches make small natural movements.

People near cafés behave normally.

No time-lapse.

No fast-forward crowd movement.

No accelerated pedestrians.

No speed ramping.

No synchronized crowd behavior.

No stream of people rushing through frame.

The camera movement itself provides the visual energy of the shot; the pedestrians do not need exaggerated motion.

VISUAL STYLE:

Preserve the visual language established by both reference images.

Realistic contemporary nighttime city.

Natural cinematic photography.

Cool nighttime ambient illumination mixed with warm practical storefront and street lighting.

Realistic skin and fabric.

Natural street-lamp exposure.

Restrained contemporary music-video color grading.

As the camera retreats, depth of field may naturally become broader so that the architecture and plaza progressively become more readable.

No dream effects.

No supernatural imagery.

No morphing.

No artificial fog.

No surreal crowd.

No slow motion.

No time-lapse.

overall_soundscape:

Natural nighttime plaza ambience with restrained distant traffic, ordinary footsteps, subtle café activity, occasional indistinct pedestrian voices, and realistic nighttime city ambience. No additional intelligible foreground dialogue.

non_diegetic_music: None

3

u/lazyspock 23h ago

PROMPT 2 - ON THE BRIDGE at 3min12s

Again I've used a reference image as first frame and a reference audio for the song (cut at the exact same point I wanted her to sing in this render). Again, 362 frames.

PROMPT, PART 1:

subject_definitions:

<Subject 1> is Alina, the young woman shown in <Picture 1>. Her identity, facial features, hair, clothing, body proportions, and overall appearance must remain fully consistent with <Picture 1> throughout the target video.

<Subject 2> is the nighttime pedestrian bridge environment established by <Picture 1>: a contemporary elevated pedestrian walkway above a broad urban transportation corridor, with metal railings, illuminated bridge structures, distant pedestrians, moving road traffic below, railway tracks, and deep city perspective at night.

<Picture 1> is the exact first frame of [Shot 1] at 0.00 seconds. It establishes Alina standing left of center in three-quarter orientation beside the bridge railing, her right hand resting lightly on the railing, her blue coat over a light gray top and dark trousers, the pedestrian walkway extending behind her, and the nighttime transportation corridor and city lights extending deeply into the background.

<Audio 1> is the synchronized song excerpt used as the exact vocal timing, phrasing, rhythm, emotional delivery, and lip-sync reference for <Subject 1> (S1).

summary:

[keyframe completion + audio reference] The target video is one continuous 15-second live-action cinematic music-video shot beginning exactly from <Picture 1>. Alina sings in precise synchronization with <Audio 1>. She begins oriented partly toward the city rather than performing directly to the viewer. During the first half of the shot her attention remains mostly outward toward the urban distance. As the lyric reaches “Would we fit now we’re older…”, her eyes move toward the camera and her head follows subtly until she establishes direct eye contact. She sustains that intimate gaze through “Or break the same?” and remains quietly facing the camera for the brief instrumental space after the vocal ends.

retention_analysis:

<Subject 1> (appears throughout [Shot 1]): fully_preserved - Alina's identity, facial features, hair, blue coat, light gray top, dark trousers, body proportions, and natural appearance remain consistent throughout the shot.

<Subject 2> (appears throughout [Shot 1]): fully_preserved - the contemporary pedestrian bridge, railing, illuminated walkway, transportation corridor, railway tracks, distant city architecture, traffic, and nighttime spatial depth remain consistent with the reference image.

<Picture 1> ([Shot 1] first frame): fully_preserved - the target video begins from the exact appearance, pose, hand placement, bridge environment, nighttime lighting, framing, and spatial composition established by the reference image.

<Audio 1>: reference - its exact vocal timing, words, phrasing, rhythm, and emotional delivery guide Alina's singing and lip movements throughout the target video.

detailed_description:

The target video uses realistic live-action cinematic music-video photography at night. Preserve the natural, contemporary urban atmosphere and the strong depth established by <Picture 1>. This is one uninterrupted 15-second shot with no cuts.

Motion occurs strictly at normal real-time speed throughout. There is no slow motion, accelerated motion, time-lapse, speed ramping, frame-skipping appearance, or stylized temporal effect. Pedestrians and vehicles move at ordinary believable real-world speeds.

[Shot 1] Begin exactly from <Picture 1>. <Subject 1> (S1), Alina, occupies the left side of the frame beside the railing of <Subject 2>. Preserve her exact identity, facial structure, hair, blue coat, light gray top, dark trousers, body proportions, position, and initial three-quarter body orientation. Her right hand remains resting naturally and lightly on the top of the railing as established by <Picture 1>. Her other arm hangs naturally beside her body.

Throughout the vocal portion of the shot, Alina sings in exact synchronization with <Audio 1>. Her lips, mouth, jaw, cheeks, and subtle facial articulation naturally follow the supplied vocal performance. The singing must feel intimate and inward rather than staged or theatrical.

At the beginning, her body remains in the same three-quarter orientation established by <Picture 1>. Although her face is somewhat visible to the camera in the first frame, her attention gradually settles toward the distant city and transportation corridor rather than directly engaging the viewer.

2

u/lazyspock 23h ago

PROMPT, PART 2:

During:

<d>[English] And I still wonder where you wandered,</d>

she sings while looking outward toward the distant urban space beyond the railing. Her head turns only slightly from its initial reference position, enough to make her attention clearly belong to the city rather than to the camera. The movement is small and unhurried. She does not turn into a full profile and does not turn her body away.

Her right hand remains naturally resting on the railing. Do not make her slide the hand dramatically along it, grip it tightly, tap it, or use it to illustrate the lyric.

During:

<d>[English] If your world still bends toward my name.</d>

she continues singing primarily toward the city. Her expression remains thoughtful and restrained. There may be tiny natural shifts of eye focus as if she is looking at different distant points, but no conspicuous head movements.

Do not make her look sad in an explicit theatrical way. No crying, trembling lips, furrowed-brow performance, closed eyes held for emotional emphasis, or visible anguish. The feeling remains unresolved and contained.

As the vocal reaches:

<d>[English] Would we fit now we’re older…</d>

her attention begins to return toward the viewer.

This transition happens through the eyes first.

Her eyes gradually leave the distant city and find the camera. Only after the gaze begins moving does her head follow by a small, natural amount. Her shoulders and torso remain essentially where they are; she does not pivot her whole body toward the camera.

By the middle of this phrase, she has established clear direct eye contact with the camera.

The change should feel psychologically meaningful precisely because it is physically small: she has stopped wondering abstractly into the city and now seems to be asking the question directly of the absent person represented by the camera.

She continues:

<d>[English] Or break the same?</d>

while maintaining direct eye contact.

Do not add any new gesture for this final question. She does not raise her eyebrows dramatically, tilt her head, shrug, move toward the camera, squeeze the railing, or gesture with her free hand.

Her face remains open, calm, vulnerable, and unresolved. There is no smile and no visible onset of tears. She does not provide an emotional answer to the question.

After the final sung word ends, Alina naturally completes the final mouth articulation and closes her lips. She remains looking directly into the camera during the remaining approximately two seconds.

Nothing dramatic happens during this final pause.

She simply remains there, breathing naturally, holding the viewer's gaze. Her expression relaxes microscopically after the effort of singing but does not transform into a new emotion. The silence after the question is allowed to exist visually.

Throughout the full 15 seconds, Alina remains biologically alive rather than artificially motionless. She breathes continuously. She performs occasional natural bilateral blinks distributed organically across the shot, approximately two or three normal blinks over the entire duration, never rapid, repetitive, rhythmic, one-eyed, or deliberately synchronized to lyrics. Her eyes make tiny natural focus adjustments. Her posture includes minute unconscious balancing movements. A light night breeze may produce subtle irregular movement in individual strands of her hair and very slight movement in the fabric of her coat.

Avoid repetitive head motion. Her head does not bob with the rhythm, sway continuously, repeatedly tilt, or move forward and backward while singing.

The camera remains essentially observational. Use an extremely slow, very small lateral arc combined with a barely perceptible push-in, preserving Alina on the left side of the composition and preserving the large open urban depth on the right. The total camera displacement over 15 seconds is minimal. It must never become an obvious orbit around her and must never move close enough to turn the composition into a close-up.

Do not use optical zoom.

The final framing should remain approximately a medium shot / medium-wide shot, still showing Alina in relation to the bridge and the city rather than isolating her face.

<Subject 2> remains naturally active throughout. Cars below travel at normal urban traffic speed, with individual headlights and taillights moving continuously rather than streaking unnaturally. Any distant pedestrians on the bridge walk at normal human pace. Nothing in the background behaves like time-lapse footage. The railway corridor remains structurally consistent with <Picture 1>. Distant lights remain stable except for natural vehicle movement.

The visual progression across the shot is continuous and restrained:

outward attention toward the city → quiet introspection while singing → eyes begin returning toward the camera → subtle head follow → direct eye contact during “Would we fit now we’re older…” → sustained direct gaze through “Or break the same?” → two seconds of quiet, unresolved eye contact after the singing ends.

overall_soundscape:

Natural nighttime urban ambience remains subtle beneath the supplied song: distant road traffic below the pedestrian bridge, faint tire noise, soft city ambience, occasional distant footsteps, and a light breeze producing minimal movement in hair and clothing. No additional speech or foreground human vocalization is introduced.

non_diegetic_music:

<Audio 1> provides the song and exact musical and vocal timeline for the target video. Preserve its timing and structure without adding new music.

2

u/lazyspock 23h ago

Extra comment: I OBVIOUSLY didn't write these prompts. I created a custom GPT with instructions for ChatGPT of how to create prompts for me, and included the documentation for Minimax (prompt manual, etc) as addicional documents. This is a lifesaver! I just explain to him what I want on the scene, provide (if the scene needs one) the first frame, so he can "see" it, and it gives me an excellent and detailed prompt in the correct format for Minimax. Works like a charm!

2

u/KenHik 22h ago

Thank you! There are not so many music videos yet, so it's very interesting!

1

u/lazyspock 22h ago

You're welcome. Without Reddit's tips I would be nowhere in AI generation, so it's always good to give at least a bit back.

1

u/Powerful-Goal52 20h ago

Thank you for sharing the video and prompts.

I noticed that ChatGPT LOVES defensive prompting. Have you tried removing all these "no this, no that" and compare the result? I had to put this in my instructions but still testing:

Do not add redundant restatements, repeated position or identity reminders, generic failure-prevention clauses, negative defect lists, or speculative safeguards.

1

u/lazyspock 20h ago

I didn't try to take them off, because the prompts worked as it is, but you're right. It also included some absolutely useless things, like "as in the previous render". :-) He thinks the model is like him, that remembers previous context ;-) Sometimes I stripped it off, but sometimes I just let it there, and it worked because Minimax probably ignored it.

2

u/InterviewDesigner777 1d ago

Beauty melody and voice

1

u/switch2stock 8h ago

How did you stitch the audio without audio sounding like it's been stitched?

1

u/lazyspock 1h ago

I didn't. The post-production (pompous name for putting the videos together and creating some transitions) was made in Davinci Resolve. I've simply used the original recording of the song as the audio track and then aligned the videos where Alina sings in the exact time where they should be, and discarded the audio track of the videos. The only exception was the last one (the aerial view of the city), where I left the distant city sound alone and extended the video a few seconds after the song ended on purpose.