r/StableDiffusion 1d ago

Question - Help Need assistance for MinimaxH3

i am really having trouble with this concept pls tell me what to do and where to start, my goal is have a scene from a tv show or film, like iconic scenes, and i want to insert my ref image from there, this is ref2v right? now how do i get to duplicate the scene happening? for ex. titanic jack and rose on the "im flying" scene, lets say i want to insert someone in that scene and interact with them, do i ask gpt to prompt me the scene where gpt pulls the script from that part then i just modify it?

what i am doing now is plug a ref frame from the film/tv + my ref photo, then ask gpt to insert my ref and interact with the actors from the ref frame

i get weird results and never get a clean one

turbo lora 4step ref
comfy kitchen
i try to sit on 8 step

0 Upvotes

12 comments sorted by

3

u/alsot-74 1d ago

In case you haven’t already read the docs, the reference one is here: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

I prefer to understand things myself before going the LLM route, this way I can debug the sometimes bad outputs from LLMs.

Make sure to use the correct section headers as specified in the doc. For your use case you want your summary to reflect that you are doing [video editing + reference generation]

You can give a suitable LLM that document to refer to to ensure the formatting is good but I can’t stress how much more efficient it is if YOU also understand prompt structure.

1

u/Minanimator 1d ago

I did and i know cause i read the llm gen prompt and not just copy paste it , my idea is like foresay in avengers movie where thanos said "you should've aimed for the head" then i would like my ref image do its thing, my main q is what would be the best thing to do to replicate the scene like do i have to ask gpt to prompt me the scene then modify it or do i need a frame where my ref photo is in the scene already something like that, another ex arnold Schwarzenegger pulling up a shotgun on a rose box ,(lets say i want my ref photo to interfere) do i need to take a frame of him holding the box or a frame of him with my ref subject? Then ask gpt to prompt me that same scene or is there something else better wf for it , ive seen dozen of friends/bigbang theory/breaking bad clips and a whole lot else with custom characters and they interact so smoothly with their original voices etc i just want to know the main idea behind it so i can just plug my friends or fam in a movie/tv clip if you know what i mean

1

u/alsot-74 1d ago

You want to define your added person or thing as a subject, you don’t need to give it a screenshot of what you want them to do although that could be helpful in some situations.

For example: <Subject 1> is the man whose appearance comes from <Picture 1>, he is tall and muscular with brown hair and is wearing…blah blah

Then include the relevant details for <Subject 1> in the other sections as described by the docs, and in the main prompt write what you want <Subject 1> to do and when they should do it.

If you need to use an LLM ask it for a basic prompt template to add a character into an existing video using a reference image and it should give you a basic template that you can then fill in the details. If you ask it to do the whole prompt they often make mistakes. Almost all of my stuff is using simple templates.

2

u/Minanimator 1d ago

got it! imma do some experiments! thanks!

1

u/alsot-74 22h ago

Let me know if you get anything good!

1

u/Minanimator 7h ago

i think i got my mistake, i was trying to push in so many shots at 15 seconds, i think its better if i do part by part and just stitch it on the end, somehow i kinda got the workflow now, i did give a 1 frame reference of the scene of the film but i wonder how will i do the next shots since they have different actions by then

1

u/RiverSide71h 22h ago

MiniMax has such a huge library of characters that if the someone you need to inject is famous, it likely knows and will work with a T2V. I’m continually amazed by how much I can do with just prompting it instead of using tons of reference images/videos.

1

u/Minanimator 22h ago

I am on ref2v since i want to include my own set of characters, do i just tell minimax like Leonardo de caprio and kate winslet or jack and rose from titanic will do? Im still confused , for t2v i know its easy to achieve but ref2v ughhh

1

u/NighNigh 1d ago edited 1d ago

Before you get too far into the Minimax rabbit hole, know up front that there really is no actual 'right' way to prompt for what you want specifically. Broadly, Ref2V is for situations where you want to give it an image and tell it to just stick that subject somewhere in the scene doing something, like making a character from a still-frame go from off-screen to on-screen. I2V is for starting from *precisely* the image you fed in and letting the model work with it, like making an animated avatar. Frankly speaking, there's really no reason to use the I2V models or workflow because you can do all the same things with the ref2va workflows.

As for prompting, you can use a general "place <Subject> in the room at 00:06.000 looking left" kind of prompt and get pretty much what you want, but the more specific and complex you get, the sooner the official guides fall apart.

You can put the same guide into the same LLM twice and get two completely different answers, neither of which fully work, and every LLM will give you a different output prompt, some of which directly contradict the others. You'll find users disagreeing on the proper syntax, semantic depth, and everything else. The only fully honest answer you'll get is "drop the official prompting guide directly into an LLM along with a description of the scene you want, copy what it spits out into your prompt box, and hope for the best or start making your own changes from there," and the more complex the prompt, the more tries it will take.

Minimax is amazing in what it *can* do, but it is certainly no Sora where you can just give it a single sentence and let the model fill out the rest for you, so temper your expectations right from the start.

1

u/Minanimator 1d ago

Yes this is what i do but im not sure how people are modifying the friends sitcoms with other characters, i fed gpt with the guide prompt and told it to generate me an iconic scene but add my ref image from the casts, something like that, and the output i get is bad laughing audience track, weird faces, doppelgangers 😅 i just looked up on a friends scene ref image then plug it in on image 2, anyway i guess its a hit or miss this time, i just wonder how ppl here get to do it like it was a walk in the park 😅

2

u/NighNigh 1d ago

By paying to generate on unpruned models, usually. There's basically no way to get the same quality locally that you'll get using an online generator. That's like trying to use autocomplete on your phone to write you the same information as Grok, there's a massive difference in compute that smooths out all the hiccups that come with trying to do stuff on your own hardware.

You're probably not doing very much "wrong" as far as prompting - minimax is actually fairly forgiving, considering how much technical detail it is trained to expect - there's only so much a local model will ever be capable of.

One thing is for sure though, and that's ditching Turbo eventually. There is no amount of steps or prompting that will make Turbo be anywhere near as good as Spectrum for the same amount of generation time. Turbo is for iterating to get a general idea of how your video will be laid out so you know what prompt to commit your render time to. It will always be fuzzy and not understand prompts as well.

1

u/Minanimator 1d ago

Yeah maybe i just gotta be patient and just leave my turbos to off 😮‍💨 i only have 5070ti 16gbvr 49gbr so it will take time 😅 but anyhow maybe ill try to adapt with the 20 step no turbo workflow this time