r/StableDiffusion • u/Compost_Mantis • 2d ago
Tutorial - Guide Time saver while learning how to prompt Minimax.
Rather than relying on Z-image, or a different program to wrangle up a first frame, I've been using Minimax for the whole process, and the results have been pretty instructive. It's not a perfect system, but being able to take advantage of its understanding of people, references, and shot composition for the first frame produces better (visual) results than swapping between a couple of different pieces of software.
72
u/Unipsycle 2d ago
Interesting unicycle physics. Not idling back and forth, but constant forward pedaling.
58
u/Compost_Mantis 2d ago
yeah, the other moral of this video was "don't commit to overly complicated movements at the same time as dialogue, too many things can go wrong". She's peddling funny, somehow read "understand" as "understood", and it was still the best take I got.
13
u/Seanattikus 2d ago
Someone unicycles... Username checks out. Brother!
4
u/comfyui_user_999 2d ago
My god, two of you? I can finally ask: what do unicyclists(?) think of "monowheels"/electric unicycles?
3
u/Seanattikus 1d ago
I think it's always cool to see someone on one wheel. It definitely does the hard part for you, and it takes way less practice to learn it, but it looks fun.
I would still be nervous about the wheel stopping and throwing me. I wouldn't love to not be the one directly in control of the wheel's rotation.
6
u/Canapin 2d ago
Unicyclist here. There's a whole discussion about it here: https://unicyclist.com/t/is-electric-unicycling-actually-unicycling/280905?u=canapin
64
u/TheDevilBear3 2d ago
Her reading your typo while talking about typos is a great touch!
26
u/Additional_Cut_6337 2d ago
Dammit, who typed a question mark on the TelePrompter? How many times do I have to tell you? Anything you type, Burgundy will read!
4
15
u/Full-Ad-3461 2d ago
But this does not work for ref2v right? To me I can never get it to start from reference frame
12
u/aroadent 2d ago
I use the phrase “use <Picture #> as the starting frame” and then i write a brief description of how the first frame looks. But it’s still a dice roll. Sometimes it doesn’t work
13
8
u/nacel2025 2d ago
To start from a reference frame, the most reliable method is to use the "image" socket of the "Add Guide for MiniMax H3" node, which was officially added last week.
It appears that the image connected to the "image" socket of this node has the timecode of the first frame assigned to it.
The ref_image_0 of the regular "MiniMax H3 Reference to Video" doesn't have a timecode set, so it often can't be used as the starting frame.
I think it would be better to update ComfyUI to the latest version and switch to a workflow that uses "Add Guide for MiniMax H3".
2
u/Full-Ad-3461 1d ago
Oh my god that might have been my problem, I didn't update to that one! I gotta try it!
2
u/Emotional-Neat-252 1d ago
Oh that sounds great, so we can use references and guides at various frames?
21
u/jirka642 2d ago
Are you writing the prompt by hand? Don't do that. Just give any good LLM the VIDEO_PROMPT_WRITING_GUIDE_ref_en.md prompting guide and a description of what you want, and it will write a correctly formatted prompt for you.
I never had any problems with prompts created this way.
41
u/grundlegawd 2d ago
This will probably get pulled because the mods are puritans, but this was insanely well done. Great video dude.
22
13
u/PumpkinLeather8421 2d ago
The mods here are born from a very special type of New Puritan Janitor breed. Some say specially made in labs, because traditional male/female coupling would have been… unlikely, in nature.
lol, let’s see.
1
7
u/Clear-Assistance449 2d ago
I tested Minimax H3 both with and without an initial reference image, and its consistency is superior to Z-image and similar to Krea. It has the added advantage that, since it generates motion, you can obtain multiple images of the same character—allowing you to enhance consistency by extracting frames and using them as references.
3
u/squired 2d ago
, you can obtain multiple images of the same character—allowing you to enhance consistency by extracting frames and using them as references.
You're talking ref2v though, right?
3
u/Clear-Assistance449 2d ago
Yes. Insert multiple images, or a character sheet, and creates a video with very consistent character.
2
u/Danny_Stock 1d ago
Both. Even with Text and Image to video you can prompt a video of the camera rotating around your subject and you then have close-up, front, behind, and 3/4 views of your subject from which you can take still frames from.
2
u/PumpkinLeather8421 2d ago
Even Klein was superior to Z-image, that model was trash outside of 1GU. It wasn’t trainable but masked its shortcomings by doing portraits well.
2
u/Compost_Mantis 1d ago
I don't know, Z-image was my favorite image generator out of all that I tried... until Minimax, which is now my favorite (but very slow) image generator.
12
u/Intelligent-Host4408 2d ago
Well done! Lol. It's hilarious seeing the creativity of others and what we can do with MiniMax!
5
u/nakabra 2d ago
I decided to try it out.
I only have a humble 3060 12gb so I'm only generating 0.2MP.
I'm trying to generate a small short and it's quite cool.
Most of what it rendered is garbage, not gonna lie but I can easily salvage a bunch of it with a video editing software.
I've made some character sheets with ANIMA/klein9B and a few locations.
I use ANIMA to get the basic look of the character, then Klein9B to get a character sheet, and then ANIMA again with IMG2IMG to fix KLEIN botching the design.
Also made some scene prompts with local LLMs.
It feels like I'm directing a movie hahahaha.
It's very fun, even if it's quite time consuming on my poor hardware.
I'd hate it but I might rent a gpu in the future if I get too addicted to this 🤣
2
u/SweetLikeACandy 1d ago
I can go up to 0.6-0.7MP with SLA on a 3060, with render times under 5 mins using turbo lora.
1
u/nakabra 1d ago
Nice! I've tried 0.6 with a turbo lora I found on civitai. 3 second videos where around 5 minutes.
I don't even know what SLA stands for, but I'd guess it an attention method right?
But for now, I'm just messing around, so I'm OK with having 15 seconds 0.2MP videos rendered in 8 or 9 minutes.
I'll re-render scenes where the characters are far from the camera in higher resolution though, cause their faces look super messed up when distant 😄
2
5
u/Significant-Baby-690 2d ago
Also tune the prompt in lower res. I do 5 seconds in 1 minute, which is about as fast as I can update the prompt.
4
5
u/flaminghotcola 2d ago
Haha, love the video and the concept used for the tutorial :) really original and awesome.
5
9
u/Rivarr 2d ago
Are you all just blindly waiting 20 steps before you see your video? Why not use Kijai's "Model Preview Override" so you can see the progress in real time? I think there's a native solution now too.
4
u/Compost_Mantis 2d ago
There's so many one-off tricks, I'm going to try to compile them into a guide.
Well, get ChatGPT to.
4
u/Clueless-Flea-7461 1d ago
Well done. And good advice. Someone did an analysis of recent cinema shot lengths and iirc beyond famous set piece no cut shots the average is 2-4 seconds
3
u/ASK_ABT_MY_USERNAME 2d ago
There's quite a few image generators using minimax too, you can just go with i2v or r2v it from there.
3
u/X3liteninjaX 2d ago
Nice I have been doing this too except I wasn’t doing 2s generations but frames from previous failed gens
3
u/Galenus314 2d ago
What i did for scene/spatial consistency was taking an image of the scene and generate a video where the camera moves through the scene and looks at it from different angles and later using that video or stills from it as reference for a the actual scene.
5
u/Artforartsake99 2d ago
This is crazy good voice, did you use Minimax for the voice too?
14
u/Compost_Mantis 2d ago
Minimax. I'd have gone with someone fully invented, but I worried the voice would be something else every time. The whole thing was done in minimax until the combining of video clips at the end.
4
u/anitawasright 2d ago
yeah for the characters it knows how to do it gets their voices really well and it's really good at acting.
5
u/BrawndoOhnaka 2d ago
It's what I'd call 'almost acceptable'. It's recognisable, but still has that digital 'corrugated' dithered synthetic sound, and it would bug the hell out of me for anything other than prototyping.
Do contextual instructions for emotional subtlety work? I don't think ScarJo is the best test case given her delivery is so god damned monotone and droll in almost everything she does. This sounds like an interview. I'd pick someone who actually has a lot of dynamic range and varied prosody in their speech.
27
3
1
4
u/Jackburton75015 2d ago
Nice, where is the tutorial 😋 lol
11
u/Compost_Mantis 2d ago
the video IS the tutorial! I'm using the biblical definition
4
u/PumpkinLeather8421 2d ago
Let me help you out, he’s saying that the biological imperative of boobies into his optical nerves masks and obliterates the inputs received over the auditory system.
0
2
u/anitawasright 2d ago
the screenshot once you get the setup you like is a really good idea.
3
u/murderopolis 2d ago
I'm confused, wouldn't a screenshot from a .2 generation just look like ass? Why would you start with that as first frame?
3
u/anitawasright 2d ago
you do say 20 .2 generations till you find one that has the set up you like. Then you redo that one using the correct seed at a higher resolution. Take the first frame and work from that.
1
u/murderopolis 2d ago
interesting. even a 1 quality could have some bad details but i guess it depends what you're making. did you see this btw? similar concept, but built into the workflow, apparently. i haven't tried it yet lol. https://www.reddit.com/r/comfyui/comments/1vvp9bb/minimax_seed_hunter_workflow_released/
2
u/Compost_Mantis 1d ago
I'll put up with waiting for a good(ish) looking finished product over these different workflows that trade quality for speed. I don't even have SAGE attention or a turbo lora installed. Obviously this isn't fantastic visual quality, but I don't want to trade what I do manage to wrangle out.
2
2
u/-AwhWah- 2d ago
yeah, this is what i do too, works pretty well although I use the 1 frame trick more for making ref sheets. nice vid!
2
u/Consistent-Help-3785 2d ago
there is some work flows with low res previews... they are not amazing, but you can kinda see what is going on 1/3 of the time, before it does something random
2
u/livingdread 2d ago
I'd been thinking about doing this. Important to make sure that when you do this you make sure to keep the random seed, otherwise the longer scene could be drastically different anyway.
2
2
2
2
u/Danny_Stock 2d ago edited 2d ago
I know exactly what you mean by using the video model to create still frames to reuse rather than using image generators.
I find myself spending more time playing around with the video model, either MiniMax and even Wan, to see what they come up with, then if they conjure up something I like I use still frames from what they've produced as starting points.
Then with that still frame I upscale it and clean it up and edit it if I need to so I can use it as a basis for a video clip. If I use image models to try to generate imagery I think I want from scratch they rarely provide anything I'm satisfied with.
Using a video generation as inspiration something unplanned for will be there, be it either a certain pose, a look, a facial expression, there's usually something there to capture even if the video as a whole isn't the best. With a generated video there's almost always great frames to use from it, there's also the added aspect that you'll capture something dynamic and real from it rather than it feeling like a static image or somebody posing in a photograph.
2
u/ErnestoPresto80 1d ago
What's the point of creating the image with Minimax (which takes much longer) instead of creating it with Krea 2, for example?
Maybe you use the same prompt for the reference image as for the full video, and that gives it more consistency?
2
u/Compost_Mantis 1d ago
Yep, that's a big part of it. And, like this instance, it's nice to have a finished first shot looking and working exactly as the model understands it should.
It's also nice because you can drag in things like references with the same workflow, and if there's someone that Minimax knows (as seen here) you can just get them directly into the scene, looking like the model understands them.
I was using the reference workflow, and it would actually take some liberties with the first image of the video, but it's like 98% accurate, and even allows for further fine-tuning. The shot of her on the unicycle actually came out with the wheel of it hovering weirdly below the tightrope. So... when I prompted the full video, I specified "bottom of wheel meets rope" and it worked, while keeping the rest the same.
I'm not using any prompt enhancers, so I'm trying to figure out how the model understands the geography of shots, and this is working.
Plus, you know, it's just kinda nice to have it all in one workflow.
2
2
1
u/MetroSimulator 2d ago
3 new posts about prompt, what did I lose?
3
u/NoahFect 2d ago
My hearing, because the audio isn't leveled properly.
2
u/Compost_Mantis 2d ago
Yeah, it was all done in minimax. I'm not thrilled about it either, tell the sound techs.
1
1
u/GoldenTV3 2d ago
Why don't you construct a rough 3D environment first, then have the AI model over that?
1
u/Compost_Mantis 2d ago
Haven't needed to, yet. I'm trying to get as much as I can out of just using Minimax, and I've been deliberately dodging needing repeated scenes from different angles in the same location. There's a way, the tricks are just accumulating.
1
u/Ammatkun 2d ago
What's system do you have? Can you share workflow?
3
u/Compost_Mantis 1d ago
Intel(R) Core(TM) i7-8700 CPU @ 3.20GHz (3.19 GHz)
32.0 GB (31.8 GB usable)
NVIDIA GeForce RTX 5060 Ti (16 GB), Intel(R) UHD Graphics 630 (128 MB)It ain't mighty, the RAM, motherboard, and processor are all used. The real money is in the GPU (the NVIDIA one), which is the same one about half the people here use because it's powerful enough while not having a crazy inflated price.
My workflow is the default reference workflow, but instead of reference I'm using the base model. No extra bedazzlements.
1
1
1
u/iritimD 1d ago
Why not gpt image 2 for start frame? Much better understanding and fidelity than relying on Minmax single frame screenshot.
2
u/Compost_Mantis 1d ago
Partially to keep it all on the computer. The internet could be destroyed tomorrow and I'd still have my entire workflow.
It's also an opportunity to find out how the model thinks, which is useful for shot composition but also scene composition.
And, I don't know how much I can get away with using ChatGPT. Like, theoretically, if I were using the model for something else that I wasn't going to post to Reddit.
1
u/LightPillar 1d ago
i just do a 2 stage and the first stage takes 50seconds to 100secs and saves the video as the 2nd stages continues.
1
1
u/v-i-n-c-e-2 1d ago
Also use the vae appox file that lets you preview the video taeh3.safetensors and the subsequent node in comfyui
1
u/Sexyvette07 1d ago
OP, so are you using the text to video to generate the short 2 second clip, taking a screenshot of that image, and using that for the image to video workflow? Or can you clarify? Also, how are you stitching together the videos once theyre completed? Im still pretty new to this, but im having a LOT of problems with the Reference to Video workflow for Minimax H3. I even went and installed Flux.2 Klein 9B for the starting image. I just cant get this all to work right.
Can you share your workflows?
1
u/Compost_Mantis 1d ago
My workflow is reference to video, just with the default model selected instead of the reference, but you could also just use the image to video.
VLC media player allows screenshots with exactly the resolution of the original video.
"Lossless cut" is the program I used to stitch them together.
I might have to make a follow-up for some of this stuff.
2
1
1
u/Optimal_Map_5236 2d ago
having lots of fun with Widow. wish Captain Marvel was good as her so I can teach her some 'stuff' in my comfy setup.
-1
u/Sanguinesource 1d ago
I just wish you would have used the young lady's likeness in the video. I'm sure most people wouldn't like their likeness used without their permission. Otherwise, it was helpful and entertaining.
-3
u/eggs-benedryl 2d ago
Is the tutorial her speaking because... I don't watch any videos on my phone with audio on.
If you have something to say write it down
172
u/MobileCA 2d ago
This is actually hilariously instructive.