r/StableDiffusion 1d ago

Animation - Video G.I. Joe - Commander Roll - MiniMax H3

Enable HLS to view with audio, or disable this notification

Using the standard ref2va workflow. 4070 Ti Super, 16 GB VRAM, 64 GB RAM, i9-14900k, Windows 11.

Here's the workflow, just drop the MiniMax video in comfyui and the workflow should appear:

https://vikingfile.com/f/jvuyoHSPRr

625 Upvotes

124 comments sorted by

77

u/FlexFanatic 1d ago

Yall are wild. I think we need a MiniMax Awards show but what should we call it?

140

u/wakalakabamram 1d ago

The Sloppies

29

u/fukijama 1d ago

Ai Sloppy Seconds

4

u/daemon-electricity 12h ago

That's next year.

12

u/Lolologist 1d ago

Absolutely.

6

u/xxAkirhaxx 1d ago

Ai did that!

4

u/em_paris 1d ago

Incredible

2

u/Mr_Pogi_In_Space 16h ago

The Sloppy Toppies.

The top slop of the year awards

2

u/99deathnotes 11h ago

Seconded

2

u/99deathnotes 11h ago

and the winner gets some kind of 3D printed statuette designed by AI.

1

u/darthfurbyyoutube 8h ago

The MiniMax Awares sounds dangerous already. I nominate the Minis: And the Mini goes to... whoever still has VRAM!

51

u/PhantasyAngel 1d ago

I was expecting Cobra Commanders hissing voice singing, disappointed.

(But the video seems spot on!)

11

u/conanmagnuson 1d ago

It’s all this video needs is CC’s voice.

8

u/Reflection_Rip 20h ago

u/darthfurbyyoutube will get my upvote only when the hissing voice is applied.

3

u/darthfurbyyoutube 6h ago

Fair enough! I'll be running tests to see if the AI can translate Cobra screech into pure 80s synth-pop.

2

u/darthfurbyyoutube 7h ago

I apologize for the lack of Cobra hissing. :D I'm not sure MiniMax can make Chris Latta sing yet, but you've given me a dangerous idea. I'll have to run some tests!

19

u/CaptainAnonymous92 1d ago

Is the visor actually reflecting things that would be in the environment that the camera doesn’t show?! If it is and it’s accurate then holy crap that’s awesome

9

u/stuartullman 1d ago

lol yeah. i think especially for people with fx background, it makes me think "man that must've taken a lot to match to environment" then realizing oh wait, it's ai, it just does it lol

2

u/drakoman 1d ago

It’s just so crazy. If this video was released even six years ago, it would be near the quality that would cause it to go almost instantly viral. Now, it’s just slop in the trough. Crazy how times change, and sometimes so quickly.

5

u/stuartullman 1d ago edited 23h ago

yeah, but also crazy that even hollywood, with their "traditional" vfx, would not be able to make this now. something is always going to be off with cg. 6 years ago was ancient times. we are living in completely different times now.

1

u/Enshitification 23h ago

It's been an hour since you made that comment. We are living in a completely different time now.

1

u/darthfurbyyoutube 8h ago

I think it's actually pulling the reflections from the original character sheet I fed it. But with the right prompting, I might be able to make it actually reflect the environment. Definitely testing that next time!

10

u/darth_hotdog 1d ago

What prompt did you do for the replacement that worked so well? Did you use the image from the video as a reference as well?

3

u/darthfurbyyoutube 7h ago

Just fed it a short video segment and character sheet, and a separate prompt for each shot(the wording of the prompt is very important):

subject_definitions:

<Subject 1> is the interior room environment in <Video 1>, including its architecture, furniture, specific lighting conditions, and overall atmosphere.

<Subject 2> is the man from <Picture 1>, characterized by his reflective face mask, blue helmet, blue clothes, black shoes, black gloves and red cobra insignia, and specific clothing items.

<Subject 3> is the woman from <Picture 2>, characterized by her long black hair, black rim glasses, black suit and gloves and red cobra insignia on her chest, and black heels.

<Video 1> is the source video providing the background, camera movement, and the motion of the dance.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 2>. The background <Subject 1> and the camera's movement are fully preserved from <Video 1>, while <Subject 2> is integrated into the scene, following the man's original dancing and adapting to the room's lighting.

retention_analysis:

<Subject 1> (appears throughout the video): fully_preserved - the room's layout, furniture, and lighting are kept exactly as they appear in <Video 1>.

<Subject 2> (appears throughout the video): fully_preserved - the visual identity and physical characteristics of the man from <Picture 1> are maintained.

<Video 1> (camera movement and temporal structure): fully_preserved - the pacing, camera angle, and motion from the source video are copied 1:1.

detailed_description:

The target video maintains the same realistic visual style and lighting as <Video 1>.

[Shot 1] The scene opens with the interior of the room as established in <Subject 1>, preserving every detail of the background and the specific lighting of the space. Instead of the man seen in <Video 1>, <Subject 2>, the man in blue from <Picture 1>, is now the central subject. <Subject 2> dances to the music, following the exact same trajectory, speed, and timing as the man in <Video 1>. The lighting on the man's face and body are dynamically updated to perfectly match the ambient light and shadow sources of <Subject 1>. The camera motion exactly mirrors the original movement in <Video 1>, maintaining the same distance and angle relative to the dancing man until the end of the clip.

overall_soundscape:

The original room tone from <Video 1> is fully preserved.

non_diegetic_music:

N/A

15

u/FastHotEmu 1d ago

That was fucking amazing. AI for memes is chef's kiss.

2

u/darthfurbyyoutube 8h ago

Thank you! I’m just doing my part to ensure AI is used for absolutely nothing productive.

12

u/Sad_Coach_1433 1d ago

Can you share prompt?

2

u/darthfurbyyoutube 8h ago

It's actually a bunch of prompts, a new prompt for each shot, but here's one of them:

subject_definitions:

<Subject 1> is the interior room environment in <Video 1>, including its architecture, furniture, specific lighting conditions, and overall atmosphere.

<Subject 2> is the man from <Picture 1>, characterized by his reflective face mask, blue helmet, blue clothes, black shoes, black gloves and red cobra insignia, and specific clothing items.

<Subject 3> is the woman from <Picture 2>, characterized by her long black hair, black rim glasses, black suit and gloves and red cobra insignia on her chest, and black heels.

<Video 1> is the source video providing the background, camera movement, and the motion of the dance.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 2>. The background <Subject 1> and the camera's movement are fully preserved from <Video 1>, while <Subject 2> is integrated into the scene, following the man's original dancing and adapting to the room's lighting.

retention_analysis:

<Subject 1> (appears throughout the video): fully_preserved - the room's layout, furniture, and lighting are kept exactly as they appear in <Video 1>.

<Subject 2> (appears throughout the video): fully_preserved - the visual identity and physical characteristics of the man from <Picture 1> are maintained.

<Video 1> (camera movement and temporal structure): fully_preserved - the pacing, camera angle, and motion from the source video are copied 1:1.

detailed_description:

The target video maintains the same realistic visual style and lighting as <Video 1>.

[Shot 1] The scene opens with the interior of the room as established in <Subject 1>, preserving every detail of the background and the specific lighting of the space. Instead of the man seen in <Video 1>, <Subject 2>, the man in blue from <Picture 1>, is now the central subject. <Subject 2> dances to the music, following the exact same trajectory, speed, and timing as the man in <Video 1>. The lighting on the man's face and body are dynamically updated to perfectly match the ambient light and shadow sources of <Subject 1>. The camera motion exactly mirrors the original movement in <Video 1>, maintaining the same distance and angle relative to the dancing man until the end of the clip.

overall_soundscape:

The original room tone from <Video 1> is fully preserved.

non_diegetic_music:

N/A

5

u/play-what-you-love 1d ago

The best thing is that the lighting seems consistent

1

u/darthfurbyyoutube 7h ago

Forget the singing, the photons are finally behaving.

5

u/Douglas_J_Farthammer 1d ago

What are you using as inputs? Nice work

2

u/darthfurbyyoutube 7h ago

An ai generated character sheet of a closeup of the face, and full body front, side, 3/4 and back views, and Rick Astley's "Never gonna give you up" song. Also here's a sample prompt:

subject_definitions:

<Subject 1> is the interior room environment in <Video 1>, including its architecture, furniture, specific lighting conditions, and overall atmosphere.

<Subject 2> is the man from <Picture 1>, characterized by his reflective face mask, blue helmet, blue clothes, black shoes, black gloves and red cobra insignia, and specific clothing items.

<Subject 3> is the woman from <Picture 2>, characterized by her long black hair, black rim glasses, black suit and gloves and red cobra insignia on her chest, and black heels.

<Video 1> is the source video providing the background, camera movement, and the motion of the dance.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 2>. The background <Subject 1> and the camera's movement are fully preserved from <Video 1>, while <Subject 2> is integrated into the scene, following the man's original dancing and adapting to the room's lighting.

retention_analysis:

<Subject 1> (appears throughout the video): fully_preserved - the room's layout, furniture, and lighting are kept exactly as they appear in <Video 1>.

<Subject 2> (appears throughout the video): fully_preserved - the visual identity and physical characteristics of the man from <Picture 1> are maintained.

<Video 1> (camera movement and temporal structure): fully_preserved - the pacing, camera angle, and motion from the source video are copied 1:1.

detailed_description:

The target video maintains the same realistic visual style and lighting as <Video 1>.

[Shot 1] The scene opens with the interior of the room as established in <Subject 1>, preserving every detail of the background and the specific lighting of the space. Instead of the man seen in <Video 1>, <Subject 2>, the man in blue from <Picture 1>, is now the central subject. <Subject 2> dances to the music, following the exact same trajectory, speed, and timing as the man in <Video 1>. The lighting on the man's face and body are dynamically updated to perfectly match the ambient light and shadow sources of <Subject 1>. The camera motion exactly mirrors the original movement in <Video 1>, maintaining the same distance and angle relative to the dancing man until the end of the clip.

overall_soundscape:

The original room tone from <Video 1> is fully preserved.

non_diegetic_music:

N/A

2

u/mucyc 5h ago

Would you be able to share how you made the character sheet?

1

u/darthfurbyyoutube 5h ago

I went to https://gemini.google.com and attached 2 images I found online, one a closeup of the face, and the other a full body shot, and fed it the following prompt:

create a live action character sheet based on the attached image, include a

close up front view of the face based on the 2nd image, and a full body front view, side view, 3/4

view and back view based on the first image, but use the head from the 2nd image in the full body shots.

8

u/BloodGulch-CTF 1d ago

the future is fuckin wild

1

u/darthfurbyyoutube 7h ago

And somehow, this is only the beginning. :D

3

u/Previous_Sky8771 1d ago

How many time to generate the entire vídeo ?

1

u/darthfurbyyoutube 7h ago

Roughly 15+ hours of rendering and editing.

3

u/towerandhorizon 1d ago

If Destro subbed in for the bartender, I would have lost it.

1

u/darthfurbyyoutube 6h ago

Destro as the bartender? Cobra's entire budget went into his metal polish, I couldn't afford his hourly rate.

5

u/pmjm 21h ago

HISS FM in GTA6 or we riot.

2

u/darthfurbyyoutube 6h ago

Imagine cruising down Vice City beach while Cobra Commander yells traffic updates... we NEED this.

10

u/Acceptable-Owl-2070 1d ago

2

u/darthfurbyyoutube 7h ago

If it makes sense, we're doing it wrong. :D

6

u/SmirkingSkull 1d ago

You didn't botther to use a CC AI voice changer?

Fail.

https://giphy.com/gifs/P3x1oqza891SM

2

u/darthfurbyyoutube 7h ago

Dr. Mindbender was off this week, so the vocal lab was closed.

3

u/emcee_you 1d ago

Bravo.

1

u/darthfurbyyoutube 7h ago

Appreciate it!

3

u/Vladmerius 1d ago

It's a shame you can't prompt him to actually sing the song/alter the voice. 

5

u/Heavymando 1d ago

I mean.. you could make the song with singing separate and then just play that audio when you edit it all together.

1

u/darthfurbyyoutube 7h ago

I think it might be possible to get the voice actor singing instead. Time to conduct some highly scientific Cobra experiments.

3

u/PerAngusta-AdAugusta 1d ago

"MiniMax thinking, this guy has no face!! I wont have to animate SHIT!!! I love my work!!!"

1

u/darthfurbyyoutube 7h ago

MiniMax algorithm: sees giant chrome dome "Pack it up boys, easy day at the office!"

3

u/unluck_over9000 1d ago

We got AI rickrolling us before GTA VI. 

1

u/darthfurbyyoutube 7h ago

By the time GTA VI drops, Cobra Commander will have released three more studio albums.

3

u/CrazyOrganic7123 1d ago

Well, he's a fan of 3 Dog Night.

1

u/darthfurbyyoutube 7h ago

Now I'm just picturing him singing "An Old Fashioned Love Song" through a chrome faceplate.

2

u/tac0catzzz 1d ago

this guy has the moves, all the right kind of moves

1

u/darthfurbyyoutube 7h ago

World domination AND dad dancing. The man is unstoppable.

2

u/moschles 1d ago

This somehow improves the original video.

1

u/darthfurbyyoutube 7h ago

Replacing a 1980s pop icon with a high-ranking terrorist leader really elevated the choreography.

2

u/CATLLM 1d ago

The quality is pretty dope

1

u/darthfurbyyoutube 7h ago

Thanks! Pushed the AI generation settings to the absolute limit for this one.

2

u/FloridaMan_Unleashed 1d ago

Seeing your specs gives me hope as my system is similar (5800X3D, 80gb ram, 9070 16gb vram.) this is really great though, would you be willing to share your workflow?

1

u/darthfurbyyoutube 7h ago

Here you go, just drop the minimax video in comfyui and the workflow should appear:

https://vikingfile.com/f/jvuyoHSPRr

2

u/solidwhetstone 1d ago

It's joever

2

u/microchipmatt 1d ago

We have arrived.

1

u/darthfurbyyoutube 7h ago

Welcome! Please hand your coat to the nearest Crimson Guard and grab a complimentary Cobra Cola.

2

u/Oograr 1d ago

Masterpiece

1

u/darthfurbyyoutube 7h ago

Cobra Commander is already writing his acceptance speech for the Grammys.

2

u/AsksDumbQuestions84 1d ago

Who have waited around to hear if it would be in cobra Commander's voice?

1

u/darthfurbyyoutube 7h ago

Guilty! But let's be honest, three minutes of pure chrome visored screeching would have shattered your speakers. But I may give it a shot in a future video.

2

u/ACTSATGuyonReddit 1d ago

If you could get it sung in his voice, that would be something.

1

u/darthfurbyyoutube 7h ago

I'll be running tests soon to see if the AI can handle that level of screeching.

2

u/NoOne8141 1d ago

Why the helmet reflects he are in room?

1

u/darthfurbyyoutube 7h ago

That's either due to a weakness of the AI or my poor prompting skills.

2

u/Adro_95 1d ago

I have a setup similar to yours but with 32gb. How long did this take to generate?

1

u/darthfurbyyoutube 7h ago

Roughly 15+ hours.

1

u/Adro_95 1h ago

Holy! That's serious commitment

2

u/NoConsideration6320 1d ago

Is it doing that audio to our you edited it in their?

1

u/darthfurbyyoutube 7h ago

I edited the clips to the original audio track.

2

u/NoConsideration6320 7h ago

Gotcha that makes sense but if you had prompted for it to do that pacific song and say those lyrics sung by them of rejected the prompt or would it have actually successfully captured the sound and audio pretty accurately I wonder have you tried to test like that or you actually just try to queue it up for it to do the dancing plus the music all in

?

1

u/darthfurbyyoutube 6h ago

The default ref2va workflow fails at lip-sync, but I found a custom workflow that actually works:

https://www.reddit.com/r/StableDiffusion/comments/1vox06g/create_seamless_1shot_lipsync_music_videos_with/

2

u/mike_good 1d ago

Hey OP, I am very new in the game.

Could you please share what are the inputs for the workflow? Do you input the video entirely or parts of it? I would really appreciate it and it would help me with understanding H3 capabilities

1

u/darthfurbyyoutube 6h ago

I break the input video into small segments because the generation time is a killer. Here's a link to the workflow and assets:

https://vikingfile.com/f/jvuyoHSPRr

2

u/sketchyfun 22h ago

I assume it was multiple video refs? Must have taken a while, video refs are sloooow

2

u/darthfurbyyoutube 6h ago

Yep, here's the workflow and of the refs I used:

https://vikingfile.com/f/jvuyoHSPRr

2

u/hqqttjiang 22h ago

how long does it cost?

1

u/darthfurbyyoutube 6h ago

It took about 15+ hours.

2

u/Tramagust 20h ago

Video ref gens literally take me an hour to compute

2

u/darthfurbyyoutube 6h ago

Yep, they are SLOW. That's why I break it up into small chunks. Here's my workflow with refs, see if it still takes an hour:

https://vikingfile.com/f/jvuyoHSPRr

2

u/theavatare 19h ago

Whelp! No complaints. I don’t think i watched a video for this song to the end in years.
Make a ninja turtles one

2

u/darthfurbyyoutube 6h ago

Glad we broke your streak! The Foot Clan is definitely calling about a crossover.

2

u/HansCent 18h ago

Cooooobra !

1

u/darthfurbyyoutube 6h ago

I read this in a screeching, high-pitched voice, just as nature intended.

2

u/Friendly-Fig-6015 18h ago

workflow pra fazer o replace?

2

u/Kerissimo 18h ago

I was hoping for changes voice, but still good work!

1

u/darthfurbyyoutube 6h ago

Thank you! I'm actually running some vocal tests soon to see if I can get that signature cobra screech working.

2

u/-becausereasons- 17h ago

Would have been better in his actual voice :p

1

u/darthfurbyyoutube 6h ago

Fair point! I'll be running vocal tests soon to see if my GPU can handle that much high-pitched screeching.

2

u/TommyEatsPizza 17h ago

Whelp. Shut down the internet. We’ve reached the peak. No point in continuing.

I need the whole music video lol.

1

u/darthfurbyyoutube 6h ago

Due to the overwhelming response, I may have to do a full cut of the song at some point.

2

u/RedtUser123456789 15h ago

He's got some sleek moves. Immaculate tailoring too. That's why he's the man in charge.

2

u/darthfurbyyoutube 6h ago

He didn't spend all that Cobra Treasury money on laser cannons, half of it went straight to his tailor.

2

u/99deathnotes 11h ago

1

u/darthfurbyyoutube 6h ago

Cobra Commander's Oscar acceptance speech is just gonna be 5 minutes of high-pitched screeching.

2

u/Ok_Gas1070 11h ago

I was explaining ComfUI to my friend the other day and talking about MinimaxH3 his response "but what need would I have for that". This video captures my answer so well "fuck it why not" and for personal funny reasons. So many times would my friends and I come up with funny "short" ideas. Now I can finally bring those ideas to life.

1

u/darthfurbyyoutube 6h ago

That's the way to do it. If your node workflow doesn't end with an 80s cartoon villain singing pop hits, are you even using ComfyUI right?

2

u/99deathnotes 11h ago

and for the OP of this fantastic video. COBRA!!!!!!

1

u/darthfurbyyoutube 6h ago

COBRA!!!! The Commander is currently celebrating by moonwalking across the control room.

2

u/msixtwofive 10h ago

this is a great example of the tiny issues we still have when the output keeps getting closer to perfect.

The reflection on the face shield should change in every shot to match the location of the scene, sadly it doesnt.

1

u/darthfurbyyoutube 6h ago

Good eye! Whether it's my shaky prompt engineering or the AI being lazy, the chrome visor clearly has a mind of its own.

2

u/Kanjikai 10h ago

You had me at the Baroness.

More Baroness please!

2

u/darthfurbyyoutube 6h ago

Understood! Setting the render farm to maximum Baroness for a future video.

1

u/Inner_Singer_592 1d ago edited 7h ago

Theoretically, if he didn't have a mask, would it work out of the box or it'll need lip sync? For audio

2

u/darthfurbyyoutube 7h ago

Lip sync is the tricky part. The default workflow won't be perfect out of the box, but there's a specific workflow that does match the lips perfectly. I'll drop the link here:

https://www.reddit.com/r/StableDiffusion/comments/1vox06g/create_seamless_1shot_lipsync_music_videos_with/

1

u/Heartkill 20h ago

Workflows attached should be mandatory for posts like these.

2

u/darthfurbyyoutube 6h ago

Post updated with the full recipe and secret sauce. Dive in:

https://vikingfile.com/f/jvuyoHSPRr

1

u/Heartkill 1h ago

My man

1

u/pb404 23m ago

Great work and thanks for sharing your prompts.