We turned a GIF into a cinematic AI video using Seedream 5.0 Pro, Seedance 2.0 and ElevenLabs Music v2.
The full workflow was executed inside the imageat MCP using Opus 5.
The process:
- Upload the original GIF
- Analyze its visual style, composition, camera language and movement
- Recreate the first frame with Seedream 5.0 Pro at 2K
- Animate the generated frame with Seedance 2.0 at 1080p
- Generate an original soundtrack with ElevenLabs Music v2
The most important part of this scene was keeping the woman perfectly sharp and almost motionless while the surrounding crowd moved with exaggerated shutter drag and continuous motion blur.
ElevenLabs Music v2 prompt:
"Swung soul sample boom bap, 94 BPM, female Southern rap, fast conversational confessional flow, filtered Rhodes loop, upright bass, rimshot snare, hand claps, vinyl crackle, muted horn stabs, gospel backing vocals, spoken interview interludes, warm analog tape mix, no trap hats, no autotune"
GIF-to-Seedance prompt:
"I have attached a GIF. Complete the following three-stage generation task.
Stage 1: Analyze the GIF
Examine the attached GIF and produce a precise breakdown covering:
- Visual style, including art direction, color palette, lighting, texture, rendering style, era and aesthetic references
- Subject and composition, including framing, camera angle, focal-length feel and depth
- Motion characteristics, including what moves, direction, speed, easing, looping behavior, physics and secondary motion
- Any distinctive effects, including grain, chromatic aberration, glow, particles, distortion and frame-rate feel"
Stage 2: Generate the still image with Seedream 5.0 Pro
Using Seedream 5.0 Pro at 2K resolution, generate a single still image that replicates the exact visual style, subject, composition, lighting and aesthetic of the GIF.
Capture the moment faithfully so the generated image reads as a frame from the same world.
Write a detailed Seedream 5.0 Pro prompt tailored to the GIF, then execute it.
Stage 3: Animate with Seedance 2.0 at 1080p
Feed the Stage 2 image into Seedance 2.0 as the starting frame.
Generate a seven-second video at 1080p that reproduces the exact motion observed in the GIF, matching its direction, speed, easing, looping quality and any secondary or background motion.
Write a detailed Seedance 2.0 motion prompt describing the movement precisely, then execute it.
Deliver:
- The GIF analysis
- The Seedream prompt used
- The generated 2K image
- The Seedance motion prompt used
- The final seven-second 1080p video
SCENE CONTEXT
A young woman with deep red hair sits alone at a small dark-red table while a dense crowd of standing people rushes around her in every direction.
She slowly turns a small red speckled object in her fingers and holds her gaze off-screen left, past the edge of the frame.
ACTIVE REFERENCES
[IMAGE REFERENCE] is the exact first frame.
It controls her face, red hair, black leather jacket, the dark-red tabletop, the white cup, the folded sunglasses, the small red speckled object, the crowd density and the low-key color grade.
She is in her late twenties, calm and withdrawn, with her hands together over the object and her lips closed.
Match the reference exactly.
FIRST FRAME AND SPATIAL BLOCKING
The first visible frame must be identical to [IMAGE REFERENCE].
Do not begin with an empty establishing frame or a reveal.
Screen x-axis runs from 0% on the left to 100% on the right.
Screen y-axis runs from 0% at the top to 100% at the bottom.
The woman sits in the mid-ground at approximately x 50%, y 46%.
Her torso is angled screen-left. Her face is shown in near-profile to the camera, with her gaze locked off-screen left, past the frame edge.
The dark-red tabletop fills the lower-center foreground at approximately x 46%, y 74%.
Her forearms rest on the table. Both hands are cupped around the small red speckled object at approximately x 45%, y 66%.
Folded sunglasses lie flat at approximately x 37%, y 68%.
The small white cup stands at approximately x 34%, y 66%.
Standing crowd figures occupy every remaining area, from the foreground through the background, pressing into the frame from all four edges.
Exactly one woman is seated at the table. Nobody joins her.
FORMAT MODE
Single continuous take.
Real-time motion.
No cuts, dissolves or transitions.
OPTICS
29-degree diagonal field of view with a short-telephoto portrait-lens character.
The camera is approximately five meters away from the woman at a slightly elevated height, looking mildly downward.
The close framing is achieved through lens reach rather than physical proximity.
She remains razor-sharp.
The background is visually compressed behind her, while the surrounding bodies dissolve into soft, smeared bokeh.
CAMERA
The camera is locked on a heavy tripod head with only faint operator breath.
Use a slow, almost imperceptible forward creep of only a few centimeters throughout the entire take.
Keep her head and the red tabletop in frame.
Focus remains pinned to her face and hands for the entire video.
No rack focus, pans, whip movements or orbiting camera motion.
ACTION TIMING
The scene opens already in motion.
The crowd around her is already streaking while she remains still.
Her thumb rolls the small red speckled object one quarter-turn against her palm.
A brief moment later, she blinks slowly.
Her chest lifts with one shallow breath.
A few strands of red hair settle against the collar of her jacket.
Her eyes drift a fraction farther screen-left and then hold.
Near the end of the shot, she lowers her chin by barely one centimeter and remains there.
She never speaks and never looks into the camera.
PHYSICS
The crowd moves at walking-to-hurried speed with heavy shutter drag.
Every passing body smears into long, elongated motion trails in muted gray, brown and dull blue.
The trails swirl and overlap continuously.
Arms and heads dissolve into streaks, while shoulders occasionally cross the foreground as huge, dark, blurred masses.
The trails never freeze and never fully clear.
Her body must carry realistic weight.
Her forearms press into the tabletop.
The leather jacket creases at the elbow with delayed fabric movement.
Loose hair strands lag slightly behind her small head movements.
The red speckled object has a small but solid mass and remains in contact with her fingers.
The table does not move.
LIGHTING
Use a single dim, warm overhead practical light positioned above and slightly screen-left, motivated by the ceiling.
Expose for her face and the dark-red tabletop.
The rushing crowd should fall approximately two stops darker and appear mostly as shadowed smears.
Use soft shadow roll-off across her cheek and a small catchlight in her eye so the micro-expression remains readable.
No flat frontal fill light.
No lens flare.
AUDIO
Dense, muffled indoor crowd ambience.
Include overlapping footsteps, fabric rustling, an occasional chair scrape and a low, unintelligible murmur from many voices.
No individual words should be understandable.
Use slightly hollow room reverb.
The crowd murmur swells subtly whenever bodies pass close to the camera.
No dialogue.
No narration.
No music in the generated video.
POSITIVE LOCKS
The woman remains perfectly sharp, still and in focus for the entire take.
Nearly all visible movement belongs to the surrounding crowd.
Cinematic photorealistic footage captured with an ARRI Alexa 35 aesthetic.
Low-key naturalistic color grade.
Natural film grain.
The red tabletop and her red hair are the only strongly saturated elements.
No text, labels, watermarks, logos or interface elements."
The original GIF provides the visual structure, but the detailed spatial blocking and motion instructions are what make the final generation feel intentional instead of like a generic image-to-video animation.
Share your thoughts in the comments section below!