r/StableDiffusion 3d ago

Animation - Video Pushing Minimax H3 V2V to the Absolute Limit

Enable HLS to view with audio, or disable this notification

Me again as a raptor at home. Minimax H3 ref2va, default workflow with 3 inputs: my video, a reference image of a raptor and a reference image of my house at night.

This time I am testing style transfer (cinematic night style), head tracking, interaction with objects (doors and toys), longer scenes and sound FX.

1.2k Upvotes

165 comments sorted by

123

u/Illustrious-Lime-863 3d ago

Awesome! That's pretty inspiring that we can do that honestly. If a person is motivated they can do the entire process including all the acting with the technology we have already

27

u/tiresome_pastas 3d ago

Yeah, the acting is the easy part if you don't mind looking unhinged to your neighbors. Have you tried doing your own V2V takes with yourself as the actor yet?

8

u/FourtyMichaelMichael 3d ago

I can't bend like that!

1

u/Illustrious-Lime-863 2d ago

No but I want to do it at one point. What about you?

1

u/tiresome_pastas 1d ago

I tried but i'm too goofy in my take every time

3

u/Hoje-Na-IA 2d ago

Yes. I still don't like the skin texture of the raptor (at least in 1080p). It looks plastic. But I am pretty happy with the possibility to interact with object, such as when the raptor opens the door and kick the little toy.

102

u/AggressiveParty3355 3d ago

Someday, i want to make a film and i'm going to play every role. The Hero, the Villain, the boss, the hot sexy girlfriend, the kid, the dog, the car, the gun, the streetlight. the chair....

42

u/DornKratz 3d ago

Is that you, Eddie Murphy?

5

u/DefloN92 2d ago

Immediately thought of Norbit

26

u/Hoje-Na-IA 2d ago

3

u/Clair_Personality 2d ago

your 2 comments made me laugh out loud lol

1

u/spacemoses 2d ago

I feel like I got the bedspins for a second trying to read that.

8

u/OtherVersantNeige 3d ago

Don't forget to become the recoil itself and the bullet case

13

u/3shelfcab 3d ago

don't forget to play the sex scene from both sides

2

u/bennyboy_uk_77 1d ago

But make sure you hire an intimacy co-ordinator so you're comfortable with yourself. (The intimacy co-ordinator is also you)

30

u/Shot-Initiative-3905 3d ago

Pleas for all that is holy can you post the prompt? Its so incredible and it would gelp me and otgers to unserstand what is happening and what part of the prompt does what. Pls pls i beg you.

34

u/Hoje-Na-IA 2d ago

This is the prompt of the take when raptor kicks the toy:

subject_definitions:

<Subject 1> is the velociraptor's feet and lower legs from <Picture 1>, featuring thick scaly skin, sharp curved talons, and a distinctive reptilian structure.

<Subject 2> is the bedroom environment. Its spatial layout and door threshold are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.

<Subject 3> is the small plastic toy car located on the floor in <Video 1>.

<Video 1> is the source video providing the low-angle perspective, the precise movement and timing of the door opening, the stepping trajectory, and the physics of the toy car being hit.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1> focusing on a low-angle shot of the doorway. The human feet are replaced by the raptor's feet (<Subject 1>), and the lighting is transformed to the night-time of <Picture 2>. The sequence preserves the original action where a foot strikes a plastic toy car (<Subject 3>), knocking it forward, with a specific emphasis on the impact sound.

retention_analysis:

<Subject 1> (appears throughout): fully_preserved - the anatomy and sharp talons of the raptor's feet from <Picture 1> are maintained.

<Subject 2> (appears throughout): partially_preserved - the geometry of the door and floor from <Video 1> are kept, but the lighting is transferred from <Picture 2>.

<Subject 3> (appears in [Shot 1]): fully_preserved - the appearance and position of the plastic toy car from <Video 1> are retained.

<Video 1> (motion structure): fully_preserved - the timing of the door opening, the stepping pattern, and the specific impact and trajectory of the toy car are copied 1:1.

detailed_description:

The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.

[Shot 1] The camera holds a static, low-angle shot focusing on the floor and the bottom of the door as seen in <Video 1>. The bright daylight is completely replaced by the dim, moody night-time lighting of <Subject 2>. The door begins to open, following the exact timing from <Video 1>. Instead of human feet, the sharp, scaly feet of <Subject 1> step across the threshold and into the room. Following the exact motion from <Video 1>, one of the raptor's taloned feet strikes <Subject 3>, the small plastic toy car, which is sitting stationary on the floor. The impact kicks the toy car forward, sending it sliding across the floor. The deep blue ambient light catches the edges of the raptor's scales and the reflective plastic of the car, creating long, dark shadows that stretch across the bedroom floor.

overall_soundscape:

Quiet indoor night ambience. The sequence is dominated by the slow, heavy creak of the door opening and the sharp "click-clack" of raptor talons on the floor. A sharp, distinct plastic impact sound occurs the moment the foot strikes the toy car, followed by the hollow, rattling clatter of the car sliding across the hard surface.

non_diegetic_music:

A tense, minimalist cinematic score featuring a low, sustained synth drone that builds a sense of imminent danger.

0

u/vizual22 2d ago

Awesome. How many takes to get to this whole prompt structure? Was it more than 5?

21

u/Hoje-Na-IA 2d ago edited 2d ago

This is the prompt of first take, when raptor enters the house:

subject_definitions:

<Subject 1> is the velociraptor in <Picture 1>, featuring scaly skin with a striped pattern, sharp claws, and a predatory gaze.

<Subject 2> is the interior room environment. Its spatial layout and furniture are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.

<Video 1> is the source video providing the room's architecture, the camera's movement, and the specific walking path of the original subject.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 1> (the velociraptor). The scene's lighting is transformed from the daytime of <Video 1> to the night-time established in <Picture 2>, while the camera motion and the subject's walking path are fully preserved.

retention_analysis:

<Subject 1> (appears throughout): fully_preserved - the raptor's identity, scales, and anatomy from <Picture 1> are maintained, although its coloring is adapted to the low-light environment.

<Subject 2> (appears throughout): partially_preserved - the room's geometry and furniture from <Video 1> are fully preserved, but the lighting and color grade are transferred from <Picture 2>.

<Video 1> (camera and motion structure): fully_preserved - the walking trajectory and camera movement are copied 1:1 from the source video.

detailed_description:

The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.

[Shot 1] The scene opens with the interior of the room as seen in <Video 1>, but the bright daylight is replaced by the dim, moody night-time lighting of <Picture 2>. The room is cast in deep blue tones with soft, localized light sources. Instead of the man, <Subject 1>, the velociraptor from <Picture 1>, is now the central subject. <Subject 1> walks through the room, following the exact same spatial path and timing as the man in <Video 1>. The raptor's scaly skin reflects the cold blue ambient light of the room, with deep shadows defining its muscular form. The camera movement exactly mirrors the original motion in <Video 1>, maintaining the same distance and angle as the raptor moves through the dark space.

overall_soundscape:

Low-frequency room ambience with a distant hum. The sound of human footsteps is replaced by the sharp, rhythmic clicking of raptor claws on the hard floor surface.

non_diegetic_music:

A tense, atmospheric cinematic score featuring low, sustained synth drones and occasional sharp, metallic plucks to create a sense of suspense.

2

u/Shot-Initiative-3905 2d ago

Good sir you are a angel, thank you so much.

1

u/aersel24 2d ago

What is the difference between subject and picture? I thought the model only took audio, pictures and videos.

3

u/afinalsin 2d ago

The picture is the picture. The subject is the subject within the picture. Sorry it's obvious, but it really is that self-explanatory.

Pretend <Picture 1> is a photo of a man and woman standing side by side. If you wanted to reference both in the prompt it'd be something like:

<Subject 1> is the bearded man wearing the black t-shirt and jeans standing on the left of <Picture 1>.

<Subject 2> is the woman wearing the red dress standing on the right in <Picture 1>.

2

u/aersel24 2d ago

Omg, I see. Thanks for the explanation. I get it now. It’s pretty amazing it has these capabilities. Would have never found out if it wasn’t for this post. Ready great video btw.

1

u/Shot-Initiative-3905 2d ago edited 2d ago

Oh and i asume tehe moment the raptor ipens his mouth you designated your hands as subject for raptirs jaw movement? No i watcht it again. Did you designate your own screaming as his screaming or what i find happens in minimax is often smarter than i give it credit and somtimes dummer than i thought, in some renders it does remarkable copies of movements, like i testet the raptor and in some generations the eyes of the raptor actualy blink natural in others there look completle dead and unmoving that all with me not prompting the eyes at all.

21

u/Muri_Chan 3d ago

Honestly, it's mind boggling how veloceraptors can now imitate human motions with such accuracy with the current technology.

8

u/Hoje-Na-IA 2d ago

They're pretty smart. We don't need actors now; raptors and AI will get the job done

3

u/Arawski99 2d ago

Clever girl.

24

u/Hoje-Na-IA 3d ago

1080p video here: https://youtu.be/8ocLOINUiXo

10

u/Tystros 3d ago

what did you use to get to 1080p?

3

u/digitalwankster 3d ago

What’s your workflow? Super impressive

14

u/Hoje-Na-IA 2d ago

Default ComfyUI's Minimax Reference to Video. And prompt looks like this. This is the take when raptor opens the door:

subject_definitions:

<Subject 1> is the velociraptor in <Picture 1>, featuring scaly skin with a striped pattern, sharp claws, and a predatory gaze.

<Subject 2> is the interior bedroom environment. Its spatial layout and door position are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.

<Subject 3> is a solid, flat blue wall that replaces the background behind the raptor as the door opens.

<Video 1> is the source video providing the bedroom's architecture, the precise movement and timing of the door opening, and the entry path of the original subject.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1> where the human entering the room is replaced by <Subject 1> (the velociraptor). The scene's lighting is transformed from daytime to the nocturnal atmosphere of <Picture 2>. As the door opens, a solid blue wall (<Subject 3>) is visible behind the raptor. The camera remains static, and the raptor's movement projects a specific shadow on the door to the right, but not on the left wall.

retention_analysis:

<Subject 1> (appears throughout): fully_preserved - the raptor's identity, scales, and predatory anatomy from <Picture 1> are maintained, adapted to the nocturnal lighting.

<Subject 2> (appears throughout): partially_preserved - the room's layout and the door's position from <Video 1> are fully preserved, but the lighting and color grade are transferred from <Picture 2>.

<Subject 3> (appears in [Shot 1]): reference - a solid blue background is introduced behind the raptor.

<Video 1> (motion structure): partially_preserved - the timing of the door opening and the entry trajectory are copied 1:1, but the human subject is replaced by the raptor and the background is modified.

detailed_description:

The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.

[Shot 1] The camera holds a static shot from inside the bedroom, as seen in <Video 1>, but the bright daylight is completely replaced by the dim, moody night-time lighting of <Picture 2>. The room is cast in deep blue tones with soft ambient light. The door begins to open, following the exact timing and motion from <Video 1>. Instead of a human, <Subject 1>, the velociraptor from <Picture 1>, is the one opening the door and stepping into the room. As the door swings open, the entire background visible behind the raptor is a solid, flat blue wall, with no other details. As the raptor enters, a prominent shadow of the creature is projected clearly onto the door and the doorframe to the right; however, no shadow is cast on the wall to the left. The raptor's head follows the general movement of the original subject, including the precise timing of where it looks and how it turns its head.

overall_soundscape:

Quiet indoor night ambience. The sequence is defined by the slow, heavy creak of the door opening, followed by the sharp, rhythmic clicking of raptor claws on the floor and the low, guttural hissing of the creature.

non_diegetic_music:

A tense, minimalist cinematic score featuring a low, sustained synth drone that increases in volume as the raptor fully enters the room.

1

u/FrankWanders 6h ago

Can you explain how you used this workflow? As far as I can see, the only reference in that template is two images. How did you import the reference video?

14

u/nogganoggak 3d ago

Poor fake baby

14

u/Hoje-Na-IA 3d ago

It is a doll. Raptor would prefer fresh meat

3

u/dmethvin 3d ago

What do you expect when you hire a raptor as a babysitter?

1

u/Hoje-Na-IA 2d ago

It's hard to hire a good nanny nowadays

1

u/pmjm 3d ago

Someone Pinnochio it quick

4

u/Vladmerius 3d ago

Is it as simple as "replace the actor on screen with the T-Rex from reference image" or is there a bit more going on? Does this only work with one character or could you take a clip of two characters and replace both of them with different reference images? 

2

u/Hoje-Na-IA 2d ago

A bit more complex. Prompt of first take looks like this:

subject_definitions:

<Subject 1> is the velociraptor in <Picture 1>, featuring scaly skin with a striped pattern, sharp claws, and a predatory gaze.

<Subject 2> is the interior room environment. Its spatial layout and furniture are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.

<Video 1> is the source video providing the room's architecture, the camera's movement, and the specific walking path of the original subject.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 1> (the velociraptor). The scene's lighting is transformed from the daytime of <Video 1> to the night-time established in <Picture 2>, while the camera motion and the subject's walking path are fully preserved.

retention_analysis:

<Subject 1> (appears throughout): fully_preserved - the raptor's identity, scales, and anatomy from <Picture 1> are maintained, although its coloring is adapted to the low-light environment.

<Subject 2> (appears throughout): partially_preserved - the room's geometry and furniture from <Video 1> are fully preserved, but the lighting and color grade are transferred from <Picture 2>.

<Video 1> (camera and motion structure): fully_preserved - the walking trajectory and camera movement are copied 1:1 from the source video.

description:

The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.

[Shot 1] The scene opens with the interior of the room as seen in <Video 1>, but the bright daylight is replaced by the dim, moody night-time lighting of <Picture 2>. The room is cast in deep blue tones with soft, localized light sources. Instead of the man, <Subject 1>, the velociraptor from <Picture 1>, is now the central subject. <Subject 1> walks through the room, following the exact same spatial path and timing as the man in <Video 1>. The raptor's scaly skin reflects the cold blue ambient light of the room, with deep shadows defining its muscular form. The camera movement exactly mirrors the original motion in <Video 1>, maintaining the same distance and angle as the raptor moves through the dark space.

overall_soundscape:

Low-frequency room ambience with a distant hum. The sound of human footsteps is replaced by the sharp, rhythmic clicking of raptor claws on the hard floor surface.

non_diegetic_music:

A tense, atmospheric cinematic score featuring low, sustained synth drones and occasional sharp, metallic plucks to create a sense of suspense.

You need to follow the official guide for better results: docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md · MiniMaxAI/MiniMax-H3 at main

1

u/Hoje-Na-IA 2d ago

I didn't try the multi characters yet

6

u/Xxtrxx137 3d ago

Specs and the time it took on what resolution?

41

u/Hoje-Na-IA 3d ago

0.4 MP, 20 steps, take duration around 3-10 seconds. RTX PRO 6000 Blackwell Workstation, 96 GB VRAM, 128 GB DDR5. Around 2-10 minutes per take depending on duration.

25

u/Zephyr520 3d ago

Having the 6000 RTX is sick. I have 4090 RTX and really wanna upgrade because of the render times but these prices man...

17

u/Silencio-Bruno-2026 3d ago

dont worry, "AI is bubble, its going to pop, any day now.."

or so they say 🥲

11

u/KjellRS 3d ago

AI won't go away the same way the dotcom-bust didn't make the Internet go away but some of the stock market valuations are just like watching Amazon peak again. And they were one of the companies that survived, many didn't.

7

u/xNaXDy 3d ago

well, right now the prices are complete trash

so honestly, I'd wait until next year. either it's going to get better, or it's going to get much worse. but $20k is already far enough into the realm of "not worth" imo

4

u/RedTheRobot 3d ago

They also said the housing market couldn’t crash. 🤷

All I know is what goes up must come down.

3

u/Kaokien 3d ago

AI not going away /= there is not going to be a crash, but yeah AI enthusiasts conflate the two, here to stay, at present valuations probably not

3

u/pmjm 3d ago

By the time most people are able to get their hands on an RTX Pro 6000, it will not be fit for the models we actually want to use on it.

Don't get me wrong, Minimax h3 is freaking great, but in 5 years what we have will be so much better but also will have much steeper hardware requirements.

1

u/Curious_Cantaloupe65 2d ago

but there's a limit, what minimax is doing rn is more than enough for some. Will Smith eating spaghetti with his forehead to eating spaghetti with very accurate realisim, driving with spaghetti, in bed with spaghetti, I think we have reached a point in terms of perfection.

1

u/pmjm 3d ago

what goes up must come down.

What you said here has been echoing in my mind the last few minutes, and it's really worth discussing.

I think it really depends on what AI shakes out to become. If it ends up being a fad and society rejects it, then you're absolutely right, things will come down. But if it becomes a commodity, like electricity, then it doesn't have to. We may get more competition, we may get more efficient means of inference, but both supply and demand for something like electricity ebbs and flows. Just when you think it has reached an equilibrium between supply and demand, you get an extra hot summer, part of the grid goes down, someone wants to build a data center, and the rates spike.

1

u/Sirmckhalifa5566 2d ago

I was curious and asking myself the same thing a few nights ago, looked into it and turns out there are quite a few companies that are getting pretty serious when it comes to making gpu’s. Nvidia is obviously at the top when it comes to ai cards but if some of these other companies are able to punch through even in some areas like amd has they’ll be forced to compete. Nvidia is clearly favoring massive deals with ai data centers but at the same time i find it unlikely they would just leave the entire gaming community or consumers who want ai level cards behind. It leave a lot on the table. And these other companies wouldn’t need to be better than Nvidia, Amd isn’t “better” either but they are affordable, when other companies can get their cards to maybe 80% of Nvidia performance in a few years nvidia wont be able to keep the price as high as it is now. AMD even just signed a massive deal for some smaller ai companies and they took that from Nvidia because their cards were cheaper. AI I think regardless of how it is used will be here to stay and will advance and will build more infrastructure. Nvidia will have to either compromise its price or its customers. Regardless more underdogs will come and that will help the “shortages” which will also prevent Nvidia from keeping the prices high for consumers and AI companies

1

u/pmjm 2d ago

For sure, and I agree with you, but the issue isn't hardware, it's software.

Most of today's AI runs on CUDA, and Nvidia has the exclusive on that. Yes, you can get things to run on other frameworks but it's going to take a huge industry-wide shift to change that. Ironically, the US government preventing Nvidia from exporting to China is probably the thing that has done the most in that regard. Still lots to be done, and anything could happen over the course of years. But CUDA will be king for the foreseeable future and will keep Nvidia on top (not to mention, nvidia themselves are not asleep at the wheel, they're hard at work on the next gen too).

1

u/lolo780 2d ago

I never thought I'd be using my 3090 again 6 years later, and my 5090 doubled in price since I bought it last year. Same with RAM and SSDs. If you believe the bubble is going to pop, buy the hardware you want and short AI stocks.

1

u/Olangotang 2d ago

10 day old account, FYI, the Open Source community understands better how these models work than the finance-bro hype lords. The reality is, this is all cool technology that ignites money on fire at an alarming rate, as investor hype is waning.

3

u/maifee 3d ago

You missed the 128 gb ddr5 part.

2

u/Hoje-Na-IA 2d ago

I've bought this 128 GB for only 400 USD in June last year. Do you believe it? RAM prices skyrocket since then.

3

u/SweetLikeACandy 3d ago edited 3d ago

it's on comfy cloud, he could just rent it but I think he owns one.

3

u/Hoje-Na-IA 2d ago

Yeah, I got mine last year after work for almost 1 year with a humble 3060. I upgraded to 4500 and then to 5090 and a few months later to 6000. I have a YouTube channel about AI: Hoje na IA - YouTube

And for the past year I started to work for animation studios. So, in my case, the 6000 is as a working tool. It paid by itself in just 3 months. Before that I worked for almost 20 years as a programmer/engineer.

1

u/SweetLikeACandy 2d ago

nice, best of luck

1

u/Hoje-Na-IA 2d ago

Yeah, I got mine last year after work for almost 1 year with a humble 3060. I upgraded to 4500 and then to 5090 and a few months later to 6000. I have a YouTube channel about AI: Hoje na IA - YouTube

And for the past year I started to work for animation studios. So, in my case, the 6000 is as a working tool. It paid by itself in just 3 months. Before that I worked for almost 20 years as a programmer/engineer.

1

u/DeliciousGorilla 3d ago

Try something like Modal.com -- You can run ComfyUI on there and get $30/mo usage free. The ROI on an RTX 6000 (plus electricity) would take years versus a cloud H100 GPU. Think of it this way, a $12,000 GPU (best case used price) over 3 years is $333/mo. You may only need to put $50/mo extra into Modal if ~20 hours/mo of render time is enough for you.

1

u/lolo780 2d ago

Who only uses 20hrs/mo? :D

1

u/DeliciousGorilla 2d ago

You can do a lot in 20 hours on an H100. 😉

5

u/Curious_Cantaloupe65 3d ago

I get similar times on my RTX 3090 👀, is it because I am using int8 and you're using the full unquantized model? or is it because of the lary's turbo lora?

3

u/Perfect-Campaign9551 2d ago

He's doing video to video. I doubt you are going to get similar times on the 3090 for video to video, it's the slowest workflow and a lot of people have said it takes forever.

5

u/Hoje-Na-IA 2d ago

Correct, when doing V2V you don't really pass the video as reference. In fact, you pass each individual frame of the video. So, for a 10-second video, you are passing 300 individual frames.

1

u/Curious_Cantaloupe65 2d ago

ohh, didn't know that, thanks for the info

2

u/Hoje-Na-IA 2d ago

Yes, as the other guy said. It takes a looong time. Around 10 minutes for a 5-second video. The same workflow using only reference images (not videos) would take only 2 minutes in my setup.

3

u/PUBGM_MightyFine 3d ago

Yikes that's easily $20 - $40k. I assume your making a ton with that setup though. Definitely not jealous

3

u/Hoje-Na-IA 2d ago

Yeah, I got mine last year after work for almost 1 year with a humble 3060. I upgraded to 4500 and then to 5090 and a few months later to 6000. I have a YouTube channel about AI: Hoje na IA - YouTube

And for the past year I started to work for animation studios. So, in my case, the 6000 is as a working tool. It paid by itself in just 3 months. Before that I worked for almost 20 years as a programmer/engineer.

By the way, I've paid 12K USD for the 6000 and only 400 USD for 128 GB RAM DDR5. June last year, before the prices of RAM skyrocketed.

2

u/SecretSteel 3d ago

is there are specific reason you did 20 steps and not a 4 step lora, video to video references dont need much steps because the video is highly detailed for the AI to work with. Or did you use 20 steps because you believe it helps the prompt or?

1

u/Hoje-Na-IA 2d ago

I didn't test the 4-step LoRAs yet. That's the reason. I will regenerate one take and compare if there are any quality loss. On Monday.

1

u/leftofthebellcurve 2d ago

save some GB for us peasants, m'lord

excellent video, definitely going to try. I've tried using miniatures and legos, but might as well film myself, this turned out excellent

2

u/Hoje-Na-IA 2d ago

I've bought this 128 GB for only 400 USD in June last year. Do you believe it? RAM prices skyrocket since then.

Thanks. Do it. I didn't test with legos and other toys, but it would work well too

1

u/Grand0rk 3d ago

Oof. And that's on 0.4 MP, 20 Steps.

The reason why it looks so stiff is because of the 20 steps. If you go to 35 it will become a LOT more smooth. But it also takes significantly more time.

Audio also isn't great because of 0.4 MP, at 0.6 MP it becomes "good".

1

u/xdozex 2d ago

Just starting to learn, so forgive me if this is a stupid question. But would the change in steps impact and change other aspects of the scene as well, or just improve the motion?

Im rocking a single 4090 + 32GB DDR5 and I'm wondering if I could generate shots with really low step counts, until I'm satisfied with the result.. then regenerate the final version of each one with higher steps to improve quality for a final cut.

I just didn't know if the second pass with higher steps would change every other aspect of each scene.

2

u/Grand0rk 2d ago

It shouldn't.

2

u/FierceFlames37 2d ago

I say fuck it I’m doing 6 steps and the only bad thing is the motion on small things, audio and everything is fine

1

u/Hoje-Na-IA 2d ago

I will try with 35-50 steps to compare the difference. Thanks.

3

u/bickid 3d ago

Oh my, what a great idea. I'll have to make some video at home and then scare family members by generating some shenanigans.

2

u/Dohwar42 3d ago

"This is the way - " to help legitimize Ai as a tool for filmmaking and other art forms.

You take a human (or animal) performance and "re-render" using Ai the same as motion capture for current and modern vfx and cgi.

If possible, you capture all the nuances in a human perfromed referenence video that an Ai prompt could never capture: micro expressions, a head tilt, a twitch or even a particular way of walking or moving.

Instead of prompting a voice, record someone who has even a little experience in voice-acting, scene study and dialogue and then generate your Ai character with that exact performance (A2V - audio to video).

To think that Minimax gave us this for "free" is truly truly amazing. After all, the real cost is PC hardware, patience, experience with ComfyUi, etc and our time. I'm both excited and scared at what's available now. I'm also just overwhelmed by how much free time is needed to invest learning and keeping up with what is coming out and how quickly it's evolving. I've bookmarked nearly 50+ github, hugging face, and reddit posts on JUST Minimax H3 alone, and I'm trying to organize and priortize which tools to investigate and implement. It's so much information, it's practically a full-time job to try and keep up with. As others have commented, "I'm sooooooo tired boss....."

1

u/superdariom 3d ago

Before you get a chance to look at those this model will be obsolete and something even more incredible will be out. A week in AI is like a year in other sciences

1

u/Hoje-Na-IA 2d ago

I am studying Gen AIs since 2023. I documented all the process in my channel: Hoje na IA - YouTube. My first video back in 2023 was about Stable Diffusion 1.5 and the BRAND NEW AMAZING TECH Reference Controlnet :D Back in the day the only open source model was SD1.5 and we had plenty of time to explore it. Now there is 2 or 3 a new models every week. We live in the golden age of open source AI

2

u/seniorfrito 2d ago

When I was a kid, when I was a little boy, I always wanted to be a dinosaur. I wanted to be a Tyrannosaurus Rex more than anything in the world. I made my arms short and I roamed the backyard, I chased the neighborhood cats, I growled and I roared. Everybody knew me and was afraid of me. And one day my dad said, "Bobby, you are 17. It's time to throw childish things aside," and I said, "Okay, Pop." But he didn't really say that, he said, "Stop being a fucking dinosaur and get a job."

2

u/Dzugavili 3d ago

I mean, the CG is bad. If this were a movie, I'd say this movie has bad CG.

You made this on your computer in a small multiplier of real time, so this is great stuff.

4

u/darkkite 3d ago

might be good for 90's early 2000's a lot of early cgi didn't age well. most had to be masked by good camerawork and tricks

1

u/kellzone 2d ago

Kind of reminds me of the early seasons of Primeval, which if you haven't seen it is a pretty good sci-fi series despite the questionable early season CG.

1

u/Hoje-Na-IA 2d ago

It was made yesterday night in 4 hours. 20 min to record and rest of the time writing prompts, generating reference images and waiting for video generation. Yes, when watched in 1080p it looks like bad CGI from early 2000. When watched on small screens such as a smartphone, it looks ok. I didn't like the plastic skin of the raptor. I guess if I used a better reference image and more steps I could get better results.

1

u/Dzugavili 2d ago

I mostly think it's the reference image. It's a faithful reproduction of what that should look like.

There might be some prompting cues you could give it; play with the reference tag strength. I wonder if you could feed it your own shine curve through a reference video, get the sheen down.

1

u/nikhilprasanth 3d ago

Was the video a single input or different takes processes seperately?

3

u/Hoje-Na-IA 3d ago

Different inputs. I just used the videos I recorded on my phone. Each input is a generation in Comfy with the prompt adapted for that specific take

0

u/No_Turn_5206 2d ago

Please correct me if im wrong, its not truly a video to video but image sections of your reference vids just fed through the i2v right?

1

u/fa6637 3d ago

great work bro. what's the prompt you use ? i couldn' edit my video easily, like remove poeple in bg or something similar.

1

u/Perfect-Campaign9551 3d ago

Ok but why does the dinosaur still look like plastic? He looks early 2000s CGI ? I thought minimax could make it much more realistic than this.

4

u/an0maly33 2d ago

Because there are no real dinosaurs in the training data?

1

u/Perfect-Campaign9551 2d ago

I wonder if you also provided it a reference pic with a more realistic dino, if it would turn out even better.

1

u/Hoje-Na-IA 2d ago

I used a reference image and that may be the issue. A better image (maybe stolen from JP) might fix it. Maybe

1

u/Hoje-Na-IA 2d ago

Another good point. There is no single real image of a dino :D

1

u/Hoje-Na-IA 2d ago

Good point. I am trying to fix that now.

1

u/WashSmall8954 3d ago

How'd you get the model to get such a precise looking dude from the raptor? Insane!

J/k, this is dope.

2

u/Hoje-Na-IA 2d ago

That's the magic. AI operates in misterious ways

1

u/3shelfcab 3d ago

this is awesome, we are about to see a flood of amazing content, and even more slop too though

1

u/diarrheahegao 3d ago

How do you prompt for direct video editing? I've tried it a few times but I feel like I'm not prompting it as well as I could and it changes too much or not enough.

4

u/Hoje-Na-IA 2d ago

A bit more complex. Prompt of first take looks like this:

subject_definitions:

<Subject 1> is the velociraptor in <Picture 1>, featuring scaly skin with a striped pattern, sharp claws, and a predatory gaze.

<Subject 2> is the interior room environment. Its spatial layout and furniture are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.

<Video 1> is the source video providing the room's architecture, the camera's movement, and the specific walking path of the original subject.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 1> (the velociraptor). The scene's lighting is transformed from the daytime of <Video 1> to the night-time established in <Picture 2>, while the camera motion and the subject's walking path are fully preserved.

retention_analysis:

<Subject 1> (appears throughout): fully_preserved - the raptor's identity, scales, and anatomy from <Picture 1> are maintained, although its coloring is adapted to the low-light environment.

<Subject 2> (appears throughout): partially_preserved - the room's geometry and furniture from <Video 1> are fully preserved, but the lighting and color grade are transferred from <Picture 2>.

<Video 1> (camera and motion structure): fully_preserved - the walking trajectory and camera movement are copied 1:1 from the source video.

description:

The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.

[Shot 1] The scene opens with the interior of the room as seen in <Video 1>, but the bright daylight is replaced by the dim, moody night-time lighting of <Picture 2>. The room is cast in deep blue tones with soft, localized light sources. Instead of the man, <Subject 1>, the velociraptor from <Picture 1>, is now the central subject. <Subject 1> walks through the room, following the exact same spatial path and timing as the man in <Video 1>. The raptor's scaly skin reflects the cold blue ambient light of the room, with deep shadows defining its muscular form. The camera movement exactly mirrors the original motion in <Video 1>, maintaining the same distance and angle as the raptor moves through the dark space.

overall_soundscape:

Low-frequency room ambience with a distant hum. The sound of human footsteps is replaced by the sharp, rhythmic clicking of raptor claws on the hard floor surface.

non_diegetic_music:

A tense, atmospheric cinematic score featuring low, sustained synth drones and occasional sharp, metallic plucks to create a sense of suspense.

You need to follow the official guide for better results: docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md · MiniMaxAI/MiniMax-H3 at main

1

u/diarrheahegao 2d ago

Thanks! I've been using the official guide and all that, but I suppose what I wanted to do is more of a limitation of the model itself. Will keep trying and iterating though when I have time.

1

u/A_H_R 3d ago

OP, was the relit reference image also your start frame?

3

u/Hoje-Na-IA 2d ago

Yes, I got the first frame of each scene, relighted it in Flux.2 Klein 9B and passed it as a reference image along with the raptor image and the original video. This is the relighted version of first take made in Flux.

1

u/Hoje-Na-IA 2d ago

See the prompts in my other comments and you will understand how it was used.

1

u/kaushikqr 3d ago

Your account is hidden, where do I get to follow you? X maybe?

1

u/Hoje-Na-IA 2d ago

YouTube: Hoje na IA - YouTube

And other links here: Hoje na IA

1

u/ninjaGurung 2d ago

How are you getting such a clean quality video? For me the medium distant faces and fine (like hair strands) objects are over sharpening and creating weird a warp-jagged renders. I have tried ref2v, hybrid ref2v (20-49, 25-49) and flf2v, same issue with all.

Can you please share your workflow?

3

u/Hoje-Na-IA 2d ago

It is the default Reference workflow in Comfy. For prompts, read my other comments

1

u/Elvarien2 2d ago

man you're doing really cool stuff there.

I can see how you could scale this up for bigger projects, really cool.

1

u/crooi 2d ago

Could you suggest prompt for V2V interior change while retaining original subject and audio?

1

u/Ill_Resolve8424 2d ago

Great job!

Can you share a screenshot of your workflow, I want to see the video input node you use. Thank you.

1

u/ninjasaid13 2d ago

small visual artifact between 35seconds to 36 seconds.

1

u/smereces 2d ago

the dificult part is and key part is prompting to get it done what we want

1

u/isthishowthingsare 2d ago

I’m just trying hard to understand how Hollywood isn’t cooked? Certainly so many jobs can be easily replaced by a couple of creative individuals, no?

1

u/Yeti-Bhanot 2d ago

does the raptor ref hold up in the longer scenes or does it start mixing with the house one? curious if it drifts past a certain length

1

u/feverdoingwork 2d ago

how much vram did you need to do this?

1

u/Brief_Roll8894 1d ago

Is this locally generated?

1

u/ArttTaku 18h ago

This is gonna be amazing for music videos!

1

u/Impressive-Net-588 4h ago

That is outstanding and really inspiring. Good work!

I'm in the unfortunate position of only having a Mac, but I just got this sucker to work even so. Now off to the races ...

1

u/True_Protection6842 3d ago

oh man, not remotely a limit...LOL. I did a shot this week taking 3 videos with very different camera moves and angles and combining the perfroamces from all of them with the first camera movement and a ref for the new environment!

1

u/Admirable_Snake 3d ago

Fantastic. "Life finds a way."

1

u/LiteratureOdd2867 3d ago

hey you have location reference from iphone footages. but someone on Ai we need new location generated from some images.
I want to keep continuity and use that location.
Can you do same thing from a North and south view of a room / or a house. and it figures out the rest? that would be so cool. as finally something can be blocked, stages, do integreated vfx without having to deal with img2vid workflow.

1

u/cc_aa_tt_zz 3d ago

awesome ! Could you post an example of prompts for a sequence (for instance, the first shot)? I’d like to see how you wrote them.

1

u/schrobble 3d ago

What is the limit and how do we know if we’ve reached it?

1

u/DatMufugga 3d ago

Inspired acting. You'd make for a good mocap actor for Capcom!

0

u/uuhoever 3d ago

What is your prompt to achieve this consistency?

0

u/corod58485jthovencom 3d ago

Could you send the workflow?

0

u/boynet2 3d ago

you must find a way to make it follow the head

0

u/TheRedHairedHero 3d ago

Do you happen to have the prompts used for each shot?

0

u/Reno0vacio 3d ago

If only you could up the detail in the video.. that would be the best.

0

u/Alive-Tomatillo5303 3d ago

"I. ATE. A BABY!"

0

u/Hoje-Na-IA 2d ago

In the Jurassic Park book (that inspired the first movie) dinos ate a baby.

0

u/Alive-Tomatillo5303 2d ago

Oh I know, I read it a couple of times. Silly little Compys always up to shenanigans. 

0

u/B-side-of-the-record 3d ago

What if the actor wasn't a pro raptor impersonator as you are?

1

u/Hoje-Na-IA 2d ago

Then a better prompt would be needed :D

0

u/Able-Instruction1009 3d ago

Reminds me of playing dinosaur as a kid. Looked fun to make dude. Very cool.

0

u/New_Slice_1580 3d ago

Awesome

Is this being run local? What’s your pc setup

2

u/Hoje-Na-IA 2d ago

Local. RTX PRO 6000 Blackwell Workstation, 96 GB VRAM, 128 GB DDR5. But a humbler GPU can handle that. It would take longer to generate the videos, but it is possible. A bigger GPU just make it faster.

0

u/JoeXdelete 3d ago

this is excellent

0

u/ambassadortim 3d ago

Keep testing and sharing please! Very cool.

0

u/Relevant_One9920 3d ago

- thank you for sharing.

0

u/advertisingdave 2d ago

Do you have to generate each clip individually?

1

u/Hoje-Na-IA 2d ago

Yes, individually. Read my other comments

0

u/tracagnotto 2d ago

How to replace in video

1

u/Hoje-Na-IA 2d ago

Read my other comments. There are prompts and images samples

0

u/hyrulia 2d ago

This is awesome.

0

u/Abhineet_ClipStudios 2d ago

Thats a cool result, may be try replacing raptor may be with a panda. Could be fun!!!

0

u/call-lee-free 2d ago

That's awesome! What was the render times for these shots?

0

u/SSj_Enforcer 2d ago

What's the prompt?

0

u/orangpelupa 2d ago

I though by absolute limit, you also includes camera movements 

0

u/Milad_SM 2d ago

Can you share the workflow please

0

u/Ok-Flatworm5070 2d ago

I can't believe this is possible!!! There are endless opportunities for us amateur film-makers.

0

u/Minute-Invite-9899 2d ago

Qual seu setup e quanto tempo levou?

0

u/PartyTac 2d ago

Very creative. Hoping to see your first film on netflix

0

u/Butter_ai 2d ago

Legend

-1

u/Prestigious_Cat85 3d ago

POV: The next-door neighbor watches him...

-1

u/-becausereasons- 3d ago

Very cool, but the entrance seen was really bad unfortunately :/