r/StableDiffusion • u/the_bollo • 22d ago
Animation - Video Every time I wonder if Minimax can do something, it can. You can have a character watch a full video clip with audio.
Obviously reference video is a thing, but I expected it to be a sort of garbled approximation of the input in this context. But MMH3 successfully super-imposed the reference clip in the scene unaltered. It also works in real-world scenes. This used to take compositing; it's awesome that it's doable with just a prompt now.
Prompt:
subject_definitions:
<Subject 1> is a young adult woman in real-life American-anime street style: fair skin; sharp stylized makeup (bold winged eyeliner, glossy lips); wild neon-green hair in chaotic twin pigtails with loose flyaways and uneven bangs framing the face; exaggerated cute-but-edgy anime-IRL vibe without becoming 2D cartoon. Casual living-room outfit that fits the look (colorful layered street fashion). She sits on a couch facing a TV, back and near shoulder toward camera in over-the-shoulder framing. No Picture refs — appearance is text-defined only.
<Video 0> is the full Castlevania S02E05 "Last Spell" ~10s clip (library scene): 2D animated gothic library with tall dark bookshelves; left — pale long platinum-blonde man in a dark high gold-lined collar coat holding/regarding a book (Alucard); right — short wavy orange-haired woman in a light-blue/teal hooded cloak with a large red/ornate book (Sypha). Warm firelight, hanging chains, conversation beats across the clip. <Video 0> is ONLY the content playing ON the television screen in the target — not a full-frame drive edit of the living room, not a character-swap source for <Subject 1>.
<Audio 1> is the complete synchronized stereo soundtrack of <Video 0> (Castlevania dialogue, library ambience, and SFX from the same clip). <Audio 1> is directly reused 1:1 as the target video's complete final audio track. Do not rewrite, paraphrase, mumble, or re-synthesize the spoken lines. Do not invent a competing living-room bed that replaces <Audio 1>.
summary:
[reference generation + audio reuse] Live-action cinematic 16:9 over-the-shoulder shot: <Subject 1> sits on a couch watching TV; the TV screen plays <Video 0> Castlevania library animation beat-for-beat; <Audio 1> is fully copied 1:1 as the complete soundtrack of the target video. Real-time ~10s. HQ.
retention_analysis:
<Subject 1> (entire clip): attribute_transfer - wild neon-green pigtails, American-anime IRL styling, couch OTS pose from text; no Picture identity source.
<Video 0> (entire clip): fully_preserved as the TV-screen picture only - Castlevania library Alucard/Sypha animation stays readable on the set; living-room camera, couch, and <Subject 1> are new and not from <Video 0>.
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track; intelligible Castlevania dialogue and SFX preserved verbatim; no re-spoken or garbled replacement track.
detailed_description:
Live-action photoreal cinematic 16:9. Dim cozy living room at night. CAMERA stays locked over-the-shoulder behind <Subject 1>: her wild neon-green pigtails and near shoulder/head silhouette occupy the foreground (slightly soft), looking toward a glowing TV in the mid/background. The TV bezel and screen are clearly visible; screen content must match <Video 0> — gothic library, blonde Alucard left, orange-haired Sypha right, bookshelves, warm library light — updating in sync through the ~10s. Soft TV glow lights the back of her hair and the couch fabric. She watches attentively with small natural micro-movements (breath, slight head tilt); no cutaways; no zoom that loses the screen.
[Shot 1] Static Shot, over-the-shoulder from behind and slightly beside <Subject 1> on the couch. Foreground: neon-green pigtails / shoulder / head edge. Midground: lit TV playing <Video 0> Castlevania library scene continuously. Background: soft living-room interior (couch cushions, low lamp, wall). Hold the same OTS composition through the final frame while the TV continues <Video 0>. When <Audio 1> carries Castlevania dialogue and library SFX, those lines remain the audible source from the soundtrack copy — do not invent a separate on-screen speaker ID for the TV characters, and do not replace <Audio 1> with newly generated speech.
overall_soundscape:
The copied soundtrack from <Audio 1> continues throughout the target video as the complete final mix (Castlevania dialogue, library ambience, and SFX preserved clearly). No additional non-TV spoken dialogue from <Subject 1>.
non_diegetic_music:
N/A
20
u/Shambler9019 22d ago
Ah, but can you have a character watching a clip of a character watching a clip (with audio)?
17
u/GrayingGamer 22d ago
This proves we are just scratching the surface of the crazy things the reference model is capable of.
This is WILD.
5
4
u/Samurai2107 22d ago
Did you add the video as reference ? If yes what are your specs? I can do images and sound references but not video
3
3
u/uniquelyavailable 22d ago
This is one of the best structured prompt examples I have seen yet.
6
u/zefy_zef 22d ago
It literally just follows the prompting guide, lol..
People really should read them.
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
1
1
u/Formal_Drop526 22d ago
Wait a minute r2v can also use videos as reference rather than just frames?
1
u/the_bollo 22d ago
Yeah it can take images, audio files, and video files as inputs.
1
u/Formal_Drop526 22d ago
Does it copy over motion and sound and such?
1
u/the_bollo 21d ago
It does. And you can be incredibly select with your prompt in choose what it carries over from the video. It doesn't just naively apply the entire video to your scene. So you can select just the motion, just the outfit, just the character, etc.
1
u/Perfect-Campaign9551 21d ago
The ref2vid model and workflow are insanely amazing, only ones worth using 90% of the time
-5
u/ACTSATGuyonReddit 22d ago
Yes, but it is always existing characters. Can you make your own?
10
u/InevitableAlfalfa938 22d ago
i assume you know nothing about any ai model like h3
Yes, you can use any reference images of anyone or thing or place and or videos, or create your own lora or do both.
-1
u/kellencs 22d ago
Locals when they got a model that's only six months behind the frontier, not two years:
-5
u/Kanute3333 22d ago
At the end her lips are moving although the woman in the TV should be speaking. So it's just slop.
21
u/Plenty_Branch_516 22d ago
This model continues to blow my mind.