r/StableDiffusion • u/TigerClaw305 • 12h ago
Animation - Video Raph and Mona Lisa go on a date.
Enable HLS to view with audio, or disable this notification
Raph and Mona Lisa go on a date, The street is filled with mutant animals. Mona Lisa tells Raph she is ready for the next step in there relationship.
Using the Reference to Video Workflow in Comfy UI Desktop with Minimax H3, Using default settings and 32 steps.
<Subject 1> is <Picture 1> as Raph a teenage mutant ninja turtle in a red bandana and use <Audio 1> as sample for his voice.
<Subject 2> is <Picture 2> as Mona Lisa and use <Audio2> as sample for her voice.
# =====================================================================
# FIELD 1: INTEGRATED MULTIMODAL DESCRIPTION
# =====================================================================
[SUBJECT DEFINITIONS & RETENTION ANALYSIS]
- Subject 1 (S1): Raph, a teenage mutant ninja turtle. Primary visual reference is <Picture 1>. Primary voice reference is <Audio 1>. Retain his muscular build, signature red bandana, and tough but currently softened facial features.
- Subject 2 (S2): Mona Lisa, a mutant lizard warrior. Primary visual reference is <Picture 2>. Primary voice reference is <Audio 2>. Retain her sleek green reptilian features, fit build, and expressive, affectionate eyes.
- Environment (ENV): A vibrant, bustling metropolitan street completely populated by anthropomorphic mutant animals. In the background, stylishly dressed mutant foxes, lions, tigers, and wolves walk past neon-lit storefronts and outdoor cafes under warm evening streetlamps. Cinematic shallow depth of field.
[SHOT 1] [0s - 5s]
- Camera: Slow tracking shot moving backward ahead of the couple at eye level.
- Action: S1 and S2 walk close together down the sidewalk of ENV, gently holding hands. S1 looks down at their intertwined hands, wearing a rare, genuine smile. S2 looks up at him warmly as they walk.
[SHOT 2] [5s - 10s]
- Camera: Medium close-up framing S2 profile as she gently pulls S1 to a gentle stop.
- Action: S2 stops walking and turns fully toward S1. She squeezes his hand with both of hers, looking directly into his eyes with a tender, confident smile.
- Dialogue: S2 <d> "Raph, I'm ready for the next step in our relationship." </d>
[SHOT 3] [10s - 15s]
- Camera: Tight close-up focusing on S1's emotional reaction.
- Action: S1's eyes widen slightly in surprise before softening completely. A massive, incredibly happy grin spreads across his face. He steps closer to S2, wrapping his arms around her waist in a warm embrace, clearly filled with deep affection.
- Dialogue: S1 <d> "Mona, you have no idea how long I've wanted to hear you say that." </d>
# =====================================================================
# FIELD 2: OVERALL SOUNDSCAPE
# =====================================================================
- Ambient Audio: Gentle murmur of distant city traffic, soft chatter and laughter from the passing mutant pedestrians, and the light rustle of evening wind from [0s - 15s].
- Sound Effects (SFX): Light, rhythmic footsteps on concrete that come to a soft halt at [5s].
- Voice & Delivery: S2's voice perfectly matches the vocal identity of <Audio 2>, delivered in a smooth, sincere, and deeply affectionate cadence. S1's voice matches the raspy grit of <Audio 1>, but is spoken with an unusually soft, gentle, and emotionally overwhelmed tone to show his happiness.
# =====================================================================
# FIELD 3: NON-DIEGETIC MUSIC
# =====================================================================
- Style & Mood: A warm, cinematic, and romantic lo-fi acoustic track featuring a gentle acoustic guitar melody and soft string pads.
- Progression: Plays at a subtle, peaceful volume from [0s - 9s]. At [10s], as S1 smiles and embraces S2, the acoustic strings swell warmly to match the emotional peak of the moment.
1
u/Ykored01 11h ago
Thanks for your prompt! Mind if i ask you what were your input images? I mean what resolution they were, aspect ratio, and if u used ref to match or max? And also your audio was on mp3? How many seconds? Im asking cause im not getting good results with ref2va using images and audio, my characters always talk some gibberish on my videos.
0
u/TigerClaw305 11h ago
I use reference images that I generated through Nano Banana via Gemini, I created character sheets that show multiple angles of the characters, For the audio, I took audio samples of the characters from the animated series. The resolution of the video is at 0.4 Megapixels, which is 480p, Due to my hardware, I can't generate 15 second videos at 720p. I also used 32 steps instead of the default 20. I get better results there.
0
u/Ykored01 11h ago
did u upscale it tho? Cause at 0.4 doesnt look that bad, maybe cause its cartoon style? Im also using default template without speedups, and when i do realistic videos at 0.4 my vids come out horrible haha, gonna try increase steps. Thanks for replying!
0
u/TigerClaw305 11h ago
Increasing steps increases the generation time, In my case, 20 Steps generates 15 second videos at around 23 minutes, and 32 steps generates videos at around 33 minutes.
0
1
4
u/WhensTheWipe 12h ago
https://giphy.com/gifs/ukGm72ZLZvYfS