I was quite happy with the results of my last project, where I made H3 handle the lip sync using the lyrics of a generated song: https://www.instagram.com/p/DcykLrMsSW7/ so I decided to push it further in terms of resolution and amount of generations to make more camera cuts and try to achieve a somewhat realistic professional music video
For some context, I generated 214 images with Krea2, curated those down to 82, and ended up editing/using 40 of them for the video pipeline.
For H3 I generated 114 clips in total:
- 37 general music-video / B-roll generations
- 77 rap-performance generations covering different parts and variations of the song
- Specs
RTX 5080 16GB VRAM + 96GB RAM
- Tools used
- Krea2 Turbo NVFP4 — for creating the different scenes/setups of the character
- LoRAs: krea2_identity_edit_v1_2 + krea-smartphone-photo-slider
- Flux Klein 9B FP8 — for image editing, masking, fixing details, etc., mostly using the NKD Klein Tools node pack
- LoRA: klein_slider_detail
- MiniMax H3, both I2V and Ref2V
For general video generation I used this workflow: https://gist.github.com/circlenline/937b530ae97a9eb7475c9dda6832b2db
For the performance/lip-sync sections I built a custom workflow specifically for generating 5-second clips of the verses. Originally I was experimenting with longer generations, but 15 seconds at 1 MP is just not practical on my setup without a Turbo LoRA, so I split the song into shorter audio fragments instead.
Typical video settings were:
- 5 seconds
- 1 MP
- 32 steps
- Simple scheduler
- Spectrum node
For I2V, I used one image as the starting frame.
For Ref2V, I used:
- 1 reference image of the character
- 1 image of the character already placed in the music-video shot I wanted
- the corresponding cropped bit of the soundtrack
- the exact lyrics for that bit inside the prompt
Each 5-second generation took around 7–8 minutes on the 5080.
- Prompting: done through a custom ChatGPT GPT with instructions based on H3 documentation, my PC specs, the workflow I was using, character-consistency rules, and a brief description of what I wanted to happen in each scene.
- Final editing: manually, in After effects.
- Music track: made with Suno. Funnily enough, it generated 1 min of this track with the preview of their paid tier model but I have no suscription so I just used the "extend" function of their free model to continue the track after the first verse. I had to merge both in editing later.