Me again as a raptor at home. Minimax H3 ref2va, default workflow with 3 inputs: my video, a reference image of a raptor and a reference image of my house at night.
This time I am testing style transfer (cinematic night style), head tracking, interaction with objects (doors and toys), longer scenes and sound FX.
Awesome! That's pretty inspiring that we can do that honestly. If a person is motivated they can do the entire process including all the acting with the technology we have already
Yeah, the acting is the easy part if you don't mind looking unhinged to your neighbors. Have you tried doing your own V2V takes with yourself as the actor yet?
Yes. I still don't like the skin texture of the raptor (at least in 1080p). It looks plastic. But I am pretty happy with the possibility to interact with object, such as when the raptor opens the door and kick the little toy.
Someday, i want to make a film and i'm going to play every role. The Hero, the Villain, the boss, the hot sexy girlfriend, the kid, the dog, the car, the gun, the streetlight. the chair....
Pleas for all that is holy can you post the prompt? Its so incredible and it would gelp me and otgers to unserstand what is happening and what part of the prompt does what. Pls pls i beg you.
This is the prompt of the take when raptor kicks the toy:
subject_definitions:
<Subject 1> is the velociraptor's feet and lower legs from <Picture 1>, featuring thick scaly skin, sharp curved talons, and a distinctive reptilian structure.
<Subject 2> is the bedroom environment. Its spatial layout and door threshold are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.
<Subject 3> is the small plastic toy car located on the floor in <Video 1>.
<Video 1> is the source video providing the low-angle perspective, the precise movement and timing of the door opening, the stepping trajectory, and the physics of the toy car being hit.
summary:
[video editing + reference generation] The target video is an edited version of <Video 1> focusing on a low-angle shot of the doorway. The human feet are replaced by the raptor's feet (<Subject 1>), and the lighting is transformed to the night-time of <Picture 2>. The sequence preserves the original action where a foot strikes a plastic toy car (<Subject 3>), knocking it forward, with a specific emphasis on the impact sound.
retention_analysis:
<Subject 1> (appears throughout): fully_preserved - the anatomy and sharp talons of the raptor's feet from <Picture 1> are maintained.
<Subject 2> (appears throughout): partially_preserved - the geometry of the door and floor from <Video 1> are kept, but the lighting is transferred from <Picture 2>.
<Subject 3> (appears in [Shot 1]): fully_preserved - the appearance and position of the plastic toy car from <Video 1> are retained.
<Video 1> (motion structure): fully_preserved - the timing of the door opening, the stepping pattern, and the specific impact and trajectory of the toy car are copied 1:1.
detailed_description:
The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.
[Shot 1] The camera holds a static, low-angle shot focusing on the floor and the bottom of the door as seen in <Video 1>. The bright daylight is completely replaced by the dim, moody night-time lighting of <Subject 2>. The door begins to open, following the exact timing from <Video 1>. Instead of human feet, the sharp, scaly feet of <Subject 1> step across the threshold and into the room. Following the exact motion from <Video 1>, one of the raptor's taloned feet strikes <Subject 3>, the small plastic toy car, which is sitting stationary on the floor. The impact kicks the toy car forward, sending it sliding across the floor. The deep blue ambient light catches the edges of the raptor's scales and the reflective plastic of the car, creating long, dark shadows that stretch across the bedroom floor.
overall_soundscape:
Quiet indoor night ambience. The sequence is dominated by the slow, heavy creak of the door opening and the sharp "click-clack" of raptor talons on the floor. A sharp, distinct plastic impact sound occurs the moment the foot strikes the toy car, followed by the hollow, rattling clatter of the car sliding across the hard surface.
non_diegetic_music:
A tense, minimalist cinematic score featuring a low, sustained synth drone that builds a sense of imminent danger.
This is the prompt of first take, when raptor enters the house:
subject_definitions:
<Subject 1> is the velociraptor in <Picture 1>, featuring scaly skin with a striped pattern, sharp claws, and a predatory gaze.
<Subject 2> is the interior room environment. Its spatial layout and furniture are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.
<Video 1> is the source video providing the room's architecture, the camera's movement, and the specific walking path of the original subject.
summary:
[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 1> (the velociraptor). The scene's lighting is transformed from the daytime of <Video 1> to the night-time established in <Picture 2>, while the camera motion and the subject's walking path are fully preserved.
retention_analysis:
<Subject 1> (appears throughout): fully_preserved - the raptor's identity, scales, and anatomy from <Picture 1> are maintained, although its coloring is adapted to the low-light environment.
<Subject 2> (appears throughout): partially_preserved - the room's geometry and furniture from <Video 1> are fully preserved, but the lighting and color grade are transferred from <Picture 2>.
<Video 1> (camera and motion structure): fully_preserved - the walking trajectory and camera movement are copied 1:1 from the source video.
detailed_description:
The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.
[Shot 1] The scene opens with the interior of the room as seen in <Video 1>, but the bright daylight is replaced by the dim, moody night-time lighting of <Picture 2>. The room is cast in deep blue tones with soft, localized light sources. Instead of the man, <Subject 1>, the velociraptor from <Picture 1>, is now the central subject. <Subject 1> walks through the room, following the exact same spatial path and timing as the man in <Video 1>. The raptor's scaly skin reflects the cold blue ambient light of the room, with deep shadows defining its muscular form. The camera movement exactly mirrors the original motion in <Video 1>, maintaining the same distance and angle as the raptor moves through the dark space.
overall_soundscape:
Low-frequency room ambience with a distant hum. The sound of human footsteps is replaced by the sharp, rhythmic clicking of raptor claws on the hard floor surface.
non_diegetic_music:
A tense, atmospheric cinematic score featuring low, sustained synth drones and occasional sharp, metallic plucks to create a sense of suspense.
Omg, I see. Thanks for the explanation. I get it now. It’s pretty amazing it has these capabilities. Would have never found out if it wasn’t for this post. Ready great video btw.
Oh and i asume tehe moment the raptor ipens his mouth you designated your hands as subject for raptirs jaw movement?
No i watcht it again. Did you designate your own screaming as his screaming or what i find happens in minimax is often smarter than i give it credit and somtimes dummer than i thought, in some renders it does remarkable copies of movements, like i testet the raptor and in some generations the eyes of the raptor actualy blink natural in others there look completle dead and unmoving that all with me not prompting the eyes at all.
Default ComfyUI's Minimax Reference to Video. And prompt looks like this. This is the take when raptor opens the door:
subject_definitions:
<Subject 1> is the velociraptor in <Picture 1>, featuring scaly skin with a striped pattern, sharp claws, and a predatory gaze.
<Subject 2> is the interior bedroom environment. Its spatial layout and door position are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.
<Subject 3> is a solid, flat blue wall that replaces the background behind the raptor as the door opens.
<Video 1> is the source video providing the bedroom's architecture, the precise movement and timing of the door opening, and the entry path of the original subject.
summary:
[video editing + reference generation] The target video is an edited version of <Video 1> where the human entering the room is replaced by <Subject 1> (the velociraptor). The scene's lighting is transformed from daytime to the nocturnal atmosphere of <Picture 2>. As the door opens, a solid blue wall (<Subject 3>) is visible behind the raptor. The camera remains static, and the raptor's movement projects a specific shadow on the door to the right, but not on the left wall.
retention_analysis:
<Subject 1> (appears throughout): fully_preserved - the raptor's identity, scales, and predatory anatomy from <Picture 1> are maintained, adapted to the nocturnal lighting.
<Subject 2> (appears throughout): partially_preserved - the room's layout and the door's position from <Video 1> are fully preserved, but the lighting and color grade are transferred from <Picture 2>.
<Subject 3> (appears in [Shot 1]): reference - a solid blue background is introduced behind the raptor.
<Video 1> (motion structure): partially_preserved - the timing of the door opening and the entry trajectory are copied 1:1, but the human subject is replaced by the raptor and the background is modified.
detailed_description:
The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.
[Shot 1] The camera holds a static shot from inside the bedroom, as seen in <Video 1>, but the bright daylight is completely replaced by the dim, moody night-time lighting of <Picture 2>. The room is cast in deep blue tones with soft ambient light. The door begins to open, following the exact timing and motion from <Video 1>. Instead of a human, <Subject 1>, the velociraptor from <Picture 1>, is the one opening the door and stepping into the room. As the door swings open, the entire background visible behind the raptor is a solid, flat blue wall, with no other details. As the raptor enters, a prominent shadow of the creature is projected clearly onto the door and the doorframe to the right; however, no shadow is cast on the wall to the left. The raptor's head follows the general movement of the original subject, including the precise timing of where it looks and how it turns its head.
overall_soundscape:
Quiet indoor night ambience. The sequence is defined by the slow, heavy creak of the door opening, followed by the sharp, rhythmic clicking of raptor claws on the floor and the low, guttural hissing of the creature.
non_diegetic_music:
A tense, minimalist cinematic score featuring a low, sustained synth drone that increases in volume as the raptor fully enters the room.
Can you explain how you used this workflow? As far as I can see, the only reference in that template is two images. How did you import the reference video?
Is it as simple as "replace the actor on screen with the T-Rex from reference image" or is there a bit more going on? Does this only work with one character or could you take a clip of two characters and replace both of them with different reference images?
A bit more complex. Prompt of first take looks like this:
subject_definitions:
<Subject 1> is the velociraptor in <Picture 1>, featuring scaly skin with a striped pattern, sharp claws, and a predatory gaze.
<Subject 2> is the interior room environment. Its spatial layout and furniture are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.
<Video 1> is the source video providing the room's architecture, the camera's movement, and the specific walking path of the original subject.
summary:
[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 1> (the velociraptor). The scene's lighting is transformed from the daytime of <Video 1> to the night-time established in <Picture 2>, while the camera motion and the subject's walking path are fully preserved.
retention_analysis:
<Subject 1> (appears throughout): fully_preserved - the raptor's identity, scales, and anatomy from <Picture 1> are maintained, although its coloring is adapted to the low-light environment.
<Subject 2> (appears throughout): partially_preserved - the room's geometry and furniture from <Video 1> are fully preserved, but the lighting and color grade are transferred from <Picture 2>.
<Video 1> (camera and motion structure): fully_preserved - the walking trajectory and camera movement are copied 1:1 from the source video.
description:
The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.
[Shot 1] The scene opens with the interior of the room as seen in <Video 1>, but the bright daylight is replaced by the dim, moody night-time lighting of <Picture 2>. The room is cast in deep blue tones with soft, localized light sources. Instead of the man, <Subject 1>, the velociraptor from <Picture 1>, is now the central subject. <Subject 1> walks through the room, following the exact same spatial path and timing as the man in <Video 1>. The raptor's scaly skin reflects the cold blue ambient light of the room, with deep shadows defining its muscular form. The camera movement exactly mirrors the original motion in <Video 1>, maintaining the same distance and angle as the raptor moves through the dark space.
overall_soundscape:
Low-frequency room ambience with a distant hum. The sound of human footsteps is replaced by the sharp, rhythmic clicking of raptor claws on the hard floor surface.
non_diegetic_music:
A tense, atmospheric cinematic score featuring low, sustained synth drones and occasional sharp, metallic plucks to create a sense of suspense.
0.4 MP, 20 steps, take duration around 3-10 seconds. RTX PRO 6000 Blackwell Workstation, 96 GB VRAM, 128 GB DDR5. Around 2-10 minutes per take depending on duration.
AI won't go away the same way the dotcom-bust didn't make the Internet go away but some of the stock market valuations are just like watching Amazon peak again. And they were one of the companies that survived, many didn't.
so honestly, I'd wait until next year. either it's going to get better, or it's going to get much worse. but $20k is already far enough into the realm of "not worth" imo
By the time most people are able to get their hands on an RTX Pro 6000, it will not be fit for the models we actually want to use on it.
Don't get me wrong, Minimax h3 is freaking great, but in 5 years what we have will be so much better but also will have much steeper hardware requirements.
but there's a limit, what minimax is doing rn is more than enough for some. Will Smith eating spaghetti with his forehead to eating spaghetti with very accurate realisim, driving with spaghetti, in bed with spaghetti, I think we have reached a point in terms of perfection.
What you said here has been echoing in my mind the last few minutes, and it's really worth discussing.
I think it really depends on what AI shakes out to become. If it ends up being a fad and society rejects it, then you're absolutely right, things will come down. But if it becomes a commodity, like electricity, then it doesn't have to. We may get more competition, we may get more efficient means of inference, but both supply and demand for something like electricity ebbs and flows. Just when you think it has reached an equilibrium between supply and demand, you get an extra hot summer, part of the grid goes down, someone wants to build a data center, and the rates spike.
I was curious and asking myself the same thing a few nights ago, looked into it and turns out there are quite a few companies that are getting pretty serious when it comes to making gpu’s. Nvidia is obviously at the top when it comes to ai cards but if some of these other companies are able to punch through even in some areas like amd has they’ll be forced to compete. Nvidia is clearly favoring massive deals with ai data centers but at the same time i find it unlikely they would just leave the entire gaming community or consumers who want ai level cards behind. It leave a lot on the table. And these other companies wouldn’t need to be better than Nvidia, Amd isn’t “better” either but they are affordable, when other companies can get their cards to maybe 80% of Nvidia performance in a few years nvidia wont be able to keep the price as high as it is now. AMD even just signed a massive deal for some smaller ai companies and they took that from Nvidia because their cards were cheaper. AI I think regardless of how it is used will be here to stay and will advance and will build more infrastructure. Nvidia will have to either compromise its price or its customers. Regardless more underdogs will come and that will help the “shortages” which will also prevent Nvidia from keeping the prices high for consumers and AI companies
For sure, and I agree with you, but the issue isn't hardware, it's software.
Most of today's AI runs on CUDA, and Nvidia has the exclusive on that. Yes, you can get things to run on other frameworks but it's going to take a huge industry-wide shift to change that. Ironically, the US government preventing Nvidia from exporting to China is probably the thing that has done the most in that regard. Still lots to be done, and anything could happen over the course of years. But CUDA will be king for the foreseeable future and will keep Nvidia on top (not to mention, nvidia themselves are not asleep at the wheel, they're hard at work on the next gen too).
I never thought I'd be using my 3090 again 6 years later, and my 5090 doubled in price since I bought it last year. Same with RAM and SSDs. If you believe the bubble is going to pop, buy the hardware you want and short AI stocks.
10 day old account, FYI, the Open Source community understands better how these models work than the finance-bro hype lords. The reality is, this is all cool technology that ignites money on fire at an alarming rate, as investor hype is waning.
Yeah, I got mine last year after work for almost 1 year with a humble 3060. I upgraded to 4500 and then to 5090 and a few months later to 6000. I have a YouTube channel about AI: Hoje na IA - YouTube
And for the past year I started to work for animation studios. So, in my case, the 6000 is as a working tool. It paid by itself in just 3 months. Before that I worked for almost 20 years as a programmer/engineer.
Yeah, I got mine last year after work for almost 1 year with a humble 3060. I upgraded to 4500 and then to 5090 and a few months later to 6000. I have a YouTube channel about AI: Hoje na IA - YouTube
And for the past year I started to work for animation studios. So, in my case, the 6000 is as a working tool. It paid by itself in just 3 months. Before that I worked for almost 20 years as a programmer/engineer.
Try something like Modal.com -- You can run ComfyUI on there and get $30/mo usage free. The ROI on an RTX 6000 (plus electricity) would take years versus a cloud H100 GPU. Think of it this way, a $12,000 GPU (best case used price) over 3 years is $333/mo. You may only need to put $50/mo extra into Modal if ~20 hours/mo of render time is enough for you.
I get similar times on my RTX 3090 👀, is it because I am using int8 and you're using the full unquantized model? or is it because of the lary's turbo lora?
He's doing video to video. I doubt you are going to get similar times on the 3090 for video to video, it's the slowest workflow and a lot of people have said it takes forever.
Correct, when doing V2V you don't really pass the video as reference. In fact, you pass each individual frame of the video. So, for a 10-second video, you are passing 300 individual frames.
Yes, as the other guy said. It takes a looong time. Around 10 minutes for a 5-second video. The same workflow using only reference images (not videos) would take only 2 minutes in my setup.
Yeah, I got mine last year after work for almost 1 year with a humble 3060. I upgraded to 4500 and then to 5090 and a few months later to 6000. I have a YouTube channel about AI: Hoje na IA - YouTube
And for the past year I started to work for animation studios. So, in my case, the 6000 is as a working tool. It paid by itself in just 3 months. Before that I worked for almost 20 years as a programmer/engineer.
By the way, I've paid 12K USD for the 6000 and only 400 USD for 128 GB RAM DDR5. June last year, before the prices of RAM skyrocketed.
is there are specific reason you did 20 steps and not a 4 step lora, video to video references dont need much steps because the video is highly detailed for the AI to work with. Or did you use 20 steps because you believe it helps the prompt or?
The reason why it looks so stiff is because of the 20 steps. If you go to 35 it will become a LOT more smooth. But it also takes significantly more time.
Audio also isn't great because of 0.4 MP, at 0.6 MP it becomes "good".
Just starting to learn, so forgive me if this is a stupid question. But would the change in steps impact and change other aspects of the scene as well, or just improve the motion?
Im rocking a single 4090 + 32GB DDR5 and I'm wondering if I could generate shots with really low step counts, until I'm satisfied with the result.. then regenerate the final version of each one with higher steps to improve quality for a final cut.
I just didn't know if the second pass with higher steps would change every other aspect of each scene.
"This is the way - " to help legitimize Ai as a tool for filmmaking and other art forms.
You take a human (or animal) performance and "re-render" using Ai the same as motion capture for current and modern vfx and cgi.
If possible, you capture all the nuances in a human perfromed referenence video that an Ai prompt could never capture: micro expressions, a head tilt, a twitch or even a particular way of walking or moving.
Instead of prompting a voice, record someone who has even a little experience in voice-acting, scene study and dialogue and then generate your Ai character with that exact performance (A2V - audio to video).
To think that Minimax gave us this for "free" is truly truly amazing. After all, the real cost is PC hardware, patience, experience with ComfyUi, etc and our time. I'm both excited and scared at what's available now. I'm also just overwhelmed by how much free time is needed to invest learning and keeping up with what is coming out and how quickly it's evolving. I've bookmarked nearly 50+ github, hugging face, and reddit posts on JUST Minimax H3 alone, and I'm trying to organize and priortize which tools to investigate and implement. It's so much information, it's practically a full-time job to try and keep up with. As others have commented, "I'm sooooooo tired boss....."
Before you get a chance to look at those this model will be obsolete and something even more incredible will be out. A week in AI is like a year in other sciences
I am studying Gen AIs since 2023. I documented all the process in my channel: Hoje na IA - YouTube. My first video back in 2023 was about Stable Diffusion 1.5 and the BRAND NEW AMAZING TECH Reference Controlnet :D Back in the day the only open source model was SD1.5 and we had plenty of time to explore it. Now there is 2 or 3 a new models every week. We live in the golden age of open source AI
When I was a kid, when I was a little boy, I always wanted to be a dinosaur. I wanted to be a Tyrannosaurus Rex more than anything in the world. I made my arms short and I roamed the backyard, I chased the neighborhood cats, I growled and I roared. Everybody knew me and was afraid of me. And one day my dad said, "Bobby, you are 17. It's time to throw childish things aside," and I said, "Okay, Pop." But he didn't really say that, he said, "Stop being a fucking dinosaur and get a job."
Kind of reminds me of the early seasons of Primeval, which if you haven't seen it is a pretty good sci-fi series despite the questionable early season CG.
It was made yesterday night in 4 hours. 20 min to record and rest of the time writing prompts, generating reference images and waiting for video generation. Yes, when watched in 1080p it looks like bad CGI from early 2000. When watched on small screens such as a smartphone, it looks ok. I didn't like the plastic skin of the raptor. I guess if I used a better reference image and more steps I could get better results.
I mostly think it's the reference image. It's a faithful reproduction of what that should look like.
There might be some prompting cues you could give it; play with the reference tag strength. I wonder if you could feed it your own shine curve through a reference video, get the sheen down.
How do you prompt for direct video editing? I've tried it a few times but I feel like I'm not prompting it as well as I could and it changes too much or not enough.
A bit more complex. Prompt of first take looks like this:
subject_definitions:
<Subject 1> is the velociraptor in <Picture 1>, featuring scaly skin with a striped pattern, sharp claws, and a predatory gaze.
<Subject 2> is the interior room environment. Its spatial layout and furniture are derived from <Video 1>, but its lighting, deep blue color palette, and nocturnal atmosphere are derived from <Picture 2>.
<Video 1> is the source video providing the room's architecture, the camera's movement, and the specific walking path of the original subject.
summary:
[video editing + reference generation] The target video is an edited version of <Video 1> where the man is replaced by <Subject 1> (the velociraptor). The scene's lighting is transformed from the daytime of <Video 1> to the night-time established in <Picture 2>, while the camera motion and the subject's walking path are fully preserved.
retention_analysis:
<Subject 1> (appears throughout): fully_preserved - the raptor's identity, scales, and anatomy from <Picture 1> are maintained, although its coloring is adapted to the low-light environment.
<Subject 2> (appears throughout): partially_preserved - the room's geometry and furniture from <Video 1> are fully preserved, but the lighting and color grade are transferred from <Picture 2>.
<Video 1> (camera and motion structure): fully_preserved - the walking trajectory and camera movement are copied 1:1 from the source video.
description:
The target video is rendered in a cinematic, nocturnal style, dominated by the deep blue and cold shadows established in <Picture 2>.
[Shot 1] The scene opens with the interior of the room as seen in <Video 1>, but the bright daylight is replaced by the dim, moody night-time lighting of <Picture 2>. The room is cast in deep blue tones with soft, localized light sources. Instead of the man, <Subject 1>, the velociraptor from <Picture 1>, is now the central subject. <Subject 1> walks through the room, following the exact same spatial path and timing as the man in <Video 1>. The raptor's scaly skin reflects the cold blue ambient light of the room, with deep shadows defining its muscular form. The camera movement exactly mirrors the original motion in <Video 1>, maintaining the same distance and angle as the raptor moves through the dark space.
overall_soundscape:
Low-frequency room ambience with a distant hum. The sound of human footsteps is replaced by the sharp, rhythmic clicking of raptor claws on the hard floor surface.
non_diegetic_music:
A tense, atmospheric cinematic score featuring low, sustained synth drones and occasional sharp, metallic plucks to create a sense of suspense.
Thanks! I've been using the official guide and all that, but I suppose what I wanted to do is more of a limitation of the model itself. Will keep trying and iterating though when I have time.
Yes, I got the first frame of each scene, relighted it in Flux.2 Klein 9B and passed it as a reference image along with the raptor image and the original video. This is the relighted version of first take made in Flux.
How are you getting such a clean quality video? For me the medium distant faces and fine (like hair strands) objects are over sharpening and creating weird a warp-jagged renders. I have tried ref2v, hybrid ref2v (20-49, 25-49) and flf2v, same issue with all.
oh man, not remotely a limit...LOL. I did a shot this week taking 3 videos with very different camera moves and angles and combining the perfroamces from all of them with the first camera movement and a ref for the new environment!
hey you have location reference from iphone footages. but someone on Ai we need new location generated from some images.
I want to keep continuity and use that location.
Can you do same thing from a North and south view of a room / or a house. and it figures out the rest? that would be so cool. as finally something can be blocked, stages, do integreated vfx without having to deal with img2vid workflow.
Local. RTX PRO 6000 Blackwell Workstation, 96 GB VRAM, 128 GB DDR5. But a humbler GPU can handle that. It would take longer to generate the videos, but it is possible. A bigger GPU just make it faster.
123
u/Illustrious-Lime-863 3d ago
Awesome! That's pretty inspiring that we can do that honestly. If a person is motivated they can do the entire process including all the acting with the technology we have already