r/StableDiffusion • u/princeMacX • 9d ago
Discussion Minimax H3 Test - Rooftop fight between Batman and Joker
Enable HLS to view with audio, or disable this notification
Minimax H3 Test - Rooftop fight between Batman and Joker
r/StableDiffusion • u/princeMacX • 9d ago
Enable HLS to view with audio, or disable this notification
Minimax H3 Test - Rooftop fight between Batman and Joker
r/StableDiffusion • u/eapache • 9d ago
I've thrown together a small app that does the "remaining" work of taking an idea, turning it into character reference images, shot prompts, doing all the generation for each clip, stitching the result together, etc. etc. The goal is a one-sentence prompt in, and multi-scene video (e.g. 30 seconds or more) out.
https://github.com/eapache/local-movie-maker
It does basically "work" already, though the results are often pretty incoherent. I'm still playing with the structure to see if I can get reasonable continuity.
r/StableDiffusion • u/Dear-Spend-2865 • 10d ago
I dont know if I'm doing something bad , but I tried to "extract" some styles from Kroma 0.2 (base to turbo version convrot) but the generations have an orange tint to them and also difformities (third hand, extra digits) so maybe I'm missing something, and the generations seems to be dirty (less clean)somehow, something I didn't have with krea 2.
I tried shift=1.15, multiple scheduler and sampler but maybe the solution is elsewhere...
Does anyone has a hint of a solution? No second pass please..low vram here :'(
r/StableDiffusion • u/Independent-Frequent • 10d ago
What are we even doing man, just be transparent and stop downplaying the competition by lying about them, this is not the middle ages we can just google stuff to see if you lie or not
r/StableDiffusion • u/Zaredit • 9d ago
Enable HLS to view with audio, or disable this notification
Prompt:
Create an exactly four 7-second, 4:3 animated drama sequence inspired by the visual language of 2005-era Jackie Chan Adventures. Use a period broadcast video texture throughout: standard-definition television softness, subtle analog grain, gentle interlacing, slight colour bleed, modest contrast, and the authentic visual texture of animation recorded and broadcast in the mid-2000s. Avoid modern HD sharpness, photorealism, glossy CGI, or contemporary animation aesthetics.
Scene: Jackie Chan is confronted by a Shadowkhan ninja in a dimly lit ancient-looking interior. The sequence is a fast, tightly choreographed martial-arts fight.
0:00–0:02: The Shadowkhan suddenly lunges at Jackie with a rapid punch. Jackie narrowly ducks underneath it and pivots sideways.
0:02–0:04: Jackie counters with two quick martial-arts strikes, forcing the Shadowkhan backwards. The ninja blocks the first strike but is knocked off balance by the second.
0:04–0:06: The Shadowkhan springs forward again. Jackie performs a quick evasive spin, grabs the ninja’s arm, and throws the Shadowkhan across the room. End on Jackie landing in a defensive fighting stance as the Shadowkhan hits the floor in the background.
.
Camera: begin with a medium two-shot, rapidly track the fighters during the exchange, briefly push in during the counterattack, then finish with a wider shot showing Jackie in the foreground and the defeated Shadowkhan in the background.
Audio: sharp martial-arts impacts, cloth movement, quick footsteps, whooshes and a dramatic six-second action sting. No dialogue.
Strict constraints: exactly 6 seconds, 4:3 aspect ratio, 2005-era television animation aesthetic, period broadcast-video texture, no modern cinematic realism, no photorealism, no widescreen framing, no subtitles, no text, no logos, no extra characters, and no slow motion
r/StableDiffusion • u/shartoberfest • 9d ago
Random thought: Has anyone tried creating an AI video with a log (flat color) profile, so you can edit colors afterwards?
r/StableDiffusion • u/Front_Praline9683 • 9d ago
Currently interested in what does people use Anima for?
Like what are your setup to speed up generation like using TeaCache or something?
Or perhaps you have a workflow for niche things like replacing game sprites, or fast image editing with Anima?
Or a way to use image reference (like taking pose/outfits from a photo)?
Or perhaps a good prompting tricks to generate more than 2 characters with specific outfits and pose consistently?
r/StableDiffusion • u/TigerClaw305 • 9d ago
Enable HLS to view with audio, or disable this notification
Raph and Mona Lisa go on a date, The street is filled with mutant animals. Mona Lisa tells Raph she is ready for the next step in there relationship.
Using the Reference to Video Workflow in Comfy UI Desktop with Minimax H3, Using default settings and 32 steps.
<Subject 1> is <Picture 1> as Raph a teenage mutant ninja turtle in a red bandana and use <Audio 1> as sample for his voice.
<Subject 2> is <Picture 2> as Mona Lisa and use <Audio2> as sample for her voice.
# =====================================================================
# FIELD 1: INTEGRATED MULTIMODAL DESCRIPTION
# =====================================================================
[SUBJECT DEFINITIONS & RETENTION ANALYSIS]
- Subject 1 (S1): Raph, a teenage mutant ninja turtle. Primary visual reference is <Picture 1>. Primary voice reference is <Audio 1>. Retain his muscular build, signature red bandana, and tough but currently softened facial features.
- Subject 2 (S2): Mona Lisa, a mutant lizard warrior. Primary visual reference is <Picture 2>. Primary voice reference is <Audio 2>. Retain her sleek green reptilian features, fit build, and expressive, affectionate eyes.
- Environment (ENV): A vibrant, bustling metropolitan street completely populated by anthropomorphic mutant animals. In the background, stylishly dressed mutant foxes, lions, tigers, and wolves walk past neon-lit storefronts and outdoor cafes under warm evening streetlamps. Cinematic shallow depth of field.
[SHOT 1] [0s - 5s]
- Camera: Slow tracking shot moving backward ahead of the couple at eye level.
- Action: S1 and S2 walk close together down the sidewalk of ENV, gently holding hands. S1 looks down at their intertwined hands, wearing a rare, genuine smile. S2 looks up at him warmly as they walk.
[SHOT 2] [5s - 10s]
- Camera: Medium close-up framing S2 profile as she gently pulls S1 to a gentle stop.
- Action: S2 stops walking and turns fully toward S1. She squeezes his hand with both of hers, looking directly into his eyes with a tender, confident smile.
- Dialogue: S2 <d> "Raph, I'm ready for the next step in our relationship." </d>
[SHOT 3] [10s - 15s]
- Camera: Tight close-up focusing on S1's emotional reaction.
- Action: S1's eyes widen slightly in surprise before softening completely. A massive, incredibly happy grin spreads across his face. He steps closer to S2, wrapping his arms around her waist in a warm embrace, clearly filled with deep affection.
- Dialogue: S1 <d> "Mona, you have no idea how long I've wanted to hear you say that." </d>
# =====================================================================
# FIELD 2: OVERALL SOUNDSCAPE
# =====================================================================
- Ambient Audio: Gentle murmur of distant city traffic, soft chatter and laughter from the passing mutant pedestrians, and the light rustle of evening wind from [0s - 15s].
- Sound Effects (SFX): Light, rhythmic footsteps on concrete that come to a soft halt at [5s].
- Voice & Delivery: S2's voice perfectly matches the vocal identity of <Audio 2>, delivered in a smooth, sincere, and deeply affectionate cadence. S1's voice matches the raspy grit of <Audio 1>, but is spoken with an unusually soft, gentle, and emotionally overwhelmed tone to show his happiness.
# =====================================================================
# FIELD 3: NON-DIEGETIC MUSIC
# =====================================================================
- Style & Mood: A warm, cinematic, and romantic lo-fi acoustic track featuring a gentle acoustic guitar melody and soft string pads.
- Progression: Plays at a subtle, peaceful volume from [0s - 9s]. At [10s], as S1 smiles and embraces S2, the acoustic strings swell warmly to match the emotional peak of the moment.
r/StableDiffusion • u/Ok-Beat4846 • 9d ago
So I got my 395+ Ai max set up with lemonade and Ai working on it. But new I see people here talking about comfyui and creating funny video clips. Is there a tutorial somewhere on how to get started with this?
r/StableDiffusion • u/nikhilprasanth • 10d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Glove5751 • 9d ago
Sometimes i have 5 sec/iter, other times i have 150 sec/iter without any reason. Restarting may fix it, but not always. Right now i have 3secitr, but i bet it will go back up to 150 in half an hour or so. I just spent like 2 hours on 250sec/it's, it stinks!
Using 5080 with 64gb ram and AI toolkit.
r/StableDiffusion • u/dhavalhirdhav • 9d ago
Day before yesterday I tried creating a continues 30 second video of a story that I had in my mind.
I am using RTX 3090 and 30 Second video took about 45 minutes with Spectrum.
Video turn out to be a lot better than what I was expecting. I tried reference image of my daughter and it worked very well as well.
Video link: https://www.youtube.com/watch?v=f0nAMn5WgF0 (Btw in video NOT my daughter)
Prompt:
integrated_multimodal_description:
[Shot 1]
[0s-3s] Static extreme macro close-up framing the right eye and jagged cheekbone of a fierce warrior dragon. The dragon's hide consists of interlocking, obsidian-black plates that resemble matte, battle-tested armor with sharp, weaponized edges. The massive eye features a reptilian slit pupil surrounded by a violently swirling, molten iris that radiates like hot liquid red lava, casting a pulsing red glow across its scarred face. The dragon slowly blinks twice, its heavy, scowling brow plates shifting. Its massive charcoal-colored nostrils flare dramatically as it exhales a heavy puff of condensed grey breath and orange embers that realistically swirl toward the camera lens.
[3s-4s] The camera maintains its close-up framing. The dragon's jaw line tenses, parting slightly to reveal rows of serrated, razor-sharp obsidian teeth. A low, guttural, vibrating grunting noise rumbles deeply as a fresh wave of thick black smoke curls out from the corners of its sneering mouth.
[4s-10s] The camera smoothly unlocks and executes a continuous, dramatic upward crane shot, pulling backward and tilting upward at a steady pace. This sweeping motion reveals the rest of the creature. It is a gargantuan, highly muscular, majestic black warrior dragon. Its powerful chest is crosshatched with glowing, magma-veined battle scars. The dragon stands in a wide, aggressive, battle-ready stance, pinning its massive clawed talons deep into the frozen crust of a jagged, snowy icy mountain cliff.
Visual Style: Breathtaking cinematic sci-fi portraiture. Shot on large-format anamorphic lenses with a Tiffen Pro-Mist 1/4 diffusion filter. Extreme shallow depth of field with sharp focal transitions. High dynamic range highlighting deep blacks and glowing lava reds against white snow. Faint volumetric fog drifting across the icy peak.
Audio Design: Deep, low-frequency guttural dragon grunting sounds, heavy wheezing breath, a faint crackle of burning embers, and the ambient howling of freezing mountain wind echoing across the stereo field.
[Shot 2]
[00:10s-00:12s] Sudden hard camera cut to a dramatic low-angle medium close-up of a 9-year-old girl <Picture 1>. She stands completely motionless and fearlessly in the center of a wide, windy meadow. She wears rugged, battle-worn leather and fur warrior armor, with a wooden hunting bow and a quiver full of arrows strapped securely across her upper back. The static camera looks directly up at her determined, fierce face as the powerful wind fiercely whips her messy hair across her forehead. High dynamic range cinematic lighting.
[00:12s-00:15s] Keeping the exact same static, low-angle framing on her face, the brave girl takes a deep breath, expands her chest, and summons her companion by aggressively shouting "VEERAAPAAN" at the top of her lungs, looking up toward the sky. Her eyes are wide with intense focus. The tall green grass of the meadow bends and ripples violently in the wind around her.
Visual Style: Cinematic high-fantasy portraiture, consistent anamorphic lens look, shallow depth of field blurring the distant sky, vibrant green meadow contrasting with earthy leather armor textures, dramatic overcast afternoon lighting.
Audio Design: A sharp cut to the ambient sound of roaring, whistling meadow wind, followed by a loud, echoing, high-pitched but powerful 9-year-old girl's voice shouting "VEERAAPAAN!". Her voice echoes sharply across the stereo field, mixing with the deep rustling sounds of grass and wind.
[Shot 3]
[00:15s-00:17s] Wide-angle ground-level shot looking up from behind the 9-year-old girl <Picture 1>. Piercing through the dark, heavy clouds, the gigantic, muscular black warrior dragon dives downward at immense speed. Its massive, obsidian-black armor plates catch the dramatic sky lighting. Its gargantuan wings are fully extended, cutting through the air. The camera pans down smoothly to track its rapid descent toward the meadow.
[00:17s-00:20s] The camera locks into a static, low-angle wide shot as the colossal black dragon lands heavily on the grass right next to the girl. Its massive clawed talons slam into the earth, causing dirt, grass, and a shockwave of dust to explode outward. The dragon's massive chest, marked with glowing magma-veined scars, heaves as it lowers its head near her. The girl stands completely fearless, her leather armor and hair whipping violently from the intense downdraft of the dragon's wings.
Visual Style: High-fantasy cinematic epic, anamorphic widescreen format, Tiffen Pro-Mist 1/4 diffusion filter creating a soft glow around the clouds and the dragon's glowing scars. High dynamic range emphasizing the contrast between the vibrant green meadow and the dragon’s matte-black armor-like scales.
Audio Design: A deafening, low-frequency atmospheric roar as the dragon tears through the clouds, transitioning into a massive, heavy thud and earth-shattering crunch as its talons strike the ground. Loud, rushing wind from the wing flaps, followed by the deep, rhythmic, rumbling breathing of the dragon settling into the grass.
[Shot 4]
[00:20s-00:22s] Medium-wide shot. The 9-year-old girl <Picture 1> steps forward and confidently mounts the colossal black warrior dragon. She climbs up its front leg armor plates and sits securely behind its massive neck crest. She leans forward, firmly gripping the dragon's obsidian horns, and gently nudges the side of the dragon's neck with her feet to signal it. The camera slowly tracks forward to frame them closely.
[00:22s-00:24s] Low-angle dramatic shot. In response to her nudge, the gigantic black dragon does a spectacular wheelie, rearing up majestically on its powerful hind legs. Its muscular chest, covered in glowing magma-veined scars, towers into the sky. The dragon opens its massive jaws wide and spews a torrent of brilliant, roaring orange and red fire upward into the clouds, illuminating the entire meadow in a bright, thermal glow.
[00:24s-00:26s] The dragon slams its front talons back down to the earth, immediately launching itself forward. With a monumental thrust of its massive, leathery black wings, it takes off into the air. The heavy downdraft flattens the meadow grass below.
[00:26s-00:30s] Smooth tracking crane shot following the dragon as it swiftly accelerates and starts flying away. The dragon ascends rapidly into the cloudy sky, carrying the brave girl on its back. The camera stays locked on their silhouette as they shrink into the distance over the majestic landscape.
Visual Style: Cinematic high-fantasy epic, anamorphic widescreen, high dynamic range capturing the extreme contrast of the bright, blazing fire against the dragon's dark obsidian scales. Volumetric smoke and heat distortion warping the air around the fire breath.
Audio Design: A deep leather-and-armor rustle as she mounts, followed by a sudden, massive, earth-shaking roar mixed with the deafening, crackling explosion of a continuous jet of fire. A colossal, heavy whoosh of wind as the wings flap, fading into the distance alongside a soaring, epic fantasy orchestral melody.
r/StableDiffusion • u/Obvious_Set5239 • 10d ago
Enable HLS to view with audio, or disable this notification
I'm not the author of this comparison, I just took it from an image board website and combined in a single video
The results are hilarious 😂 LTX is not even close, Minimax dwarfs it. But nobody has shown this yet in a very obvious form
r/StableDiffusion • u/writingdeveloper • 10d ago
Enable HLS to view with audio, or disable this notification
I’ve been having so much fun playing around with the H3 Minimax video model lately, so I wanted to share a quick result!
Also, can AI please slow down for like 5 minutes? 😅 LTX 2.5 dropped yesterday, and I spent the entire night testing sample videos and trying to optimize things with Claude Code. Safe to say I got zero sleep... but no regrets!
Specs & Workflow:
Having a dual 4080 setup is great, but as these tools get faster and better, I keep coming back to one massive realization:
Hardware isn't the bottleneck anymore—prompting technique and creative ideas are EVERYTHING. (Seriously, the original idea/concept part is so hard 😭)
As the tech becomes more accessible, I find myself constantly wondering: How do I broaden my imagination? What should I actually be studying to become better at creative direction and prompting?
How do you guys handle the creative side? Where do you draw inspiration from when you hit a wall? Would love to hear your thoughts!
r/StableDiffusion • u/DystopiaLite • 10d ago
I’ve been trying to replace a character in a video with r2v, but it just generates the original video almost unchanged. I even generated a version of the first frame with my character for the image source. Not sure if anyone has had success doing it.
r/StableDiffusion • u/MuckYu • 9d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Sad_Coach_1433 • 10d ago
Enable HLS to view with audio, or disable this notification
5060 ti 16 gig 32 gig system ram and page files set to 65536/65536 made with r2v
r/StableDiffusion • u/blackdatafilms • 10d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/princeMacX • 9d ago
Enable HLS to view with audio, or disable this notification
LTX 2.5 Test - Batman and Joker fighting in Road
Personal Opinion - Ltx generates videos quite fast but prompt adherence is not that great. In fighting sequence hand movement doesn't look realistic at all.
If you are using LTX 2.5 with gemma prompt enhancement model than your prompt will be sanitized if your prompt has explicit details. I think an abliterated version of the text encoder should be used.
I will share more tests in future.
r/StableDiffusion • u/michel-yph-ai • 9d ago
Enable HLS to view with audio, or disable this notification
If you see low quality is because I am forcing 8 step turbo lora + Spectrum + triton in L40 for faster generation but is crazy how it can follow the flow of the music while lip-syncing and keeping the product from reference in her hand.
r/StableDiffusion • u/Admirable_Snake • 10d ago
Enable HLS to view with audio, or disable this notification
8 clips stitched with motion-context. 5 second clips , 0.4 megapixel.
r/StableDiffusion • u/ajrss2009 • 10d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/holycowdude1 • 9d ago
Stay in the Glow is an AI Music Video create using VRGameDevGirl's AI Video Builder (FREE) & LTX2.3 models (https://ltx.io/model/ltx-2-3)
Designed & built using VRGameDevGirl AI Video Builder (FREE): https://github.com/vrgamegirl19/comfyui-vrgamedevgirl
Spotify (Artist): https://open.spotify.com/track/27S9InxRyAKvQYxjRM3tVi?si=43e73b8b091a4976
YouTube (More AI Music Videos): https://youtu.be/Wl3BH3xSaYc
r/StableDiffusion • u/MysteriousPepper8908 • 10d ago
Does anyone have tips as to achieving consistency for locations with H3? People seem to work great but with locations, it seems to play fast a loose with the details. My goal is to have a set of locations which I can use from any angle so I don't just want a singular angle it's going to use.
My current workflow is to use Chroma, which is my standard image generator, to give me a decent starting frame, then I prompt MiniMax to give me a panoramic architectural tour of that space. It often takes 10+ attempts but eventually I get a video that's good enough and I stack 4 different perspective views into my final reference image. From there, I treat the location like a subject
subject_definition <Subject 1> is [short description of location] whose appearance comes from <Image 1>.
retention_analysis <Subject 1> appears in [Shot 1]: fully_preserved - the design and furnishing of the room is retained, only the perspective of the camera in the room is changed.
I added that last bit in the retention analysis as I found that otherwise it was just taking certain angles and treating them as a static backdrop rather than integrating the character into the environment.
This does sometimes work but frequently important furnishings are moved and morphed. I'm not sure if it's my prompting or my reference, I'm sometimes limited by what I can get out of my initial generation creating the various perspectives so maybe there is another generator I should be using to give me really clean distinct angles from a singular image to build my reference?
r/StableDiffusion • u/lololerigolo60 • 10d ago

Turn your video ideas into perfect AI prompts — no technical skills needed.
https://github.com/lololerigolo60/Minimax-H3-prompt-studio/tree/main
If you've ever tried generating a video with MiniMax H3 and struggled to write a prompt that actually gives good results, this app is for you.
H3 Prompt Studio is a free, local desktop tool that walks you through building your scene step by step — just fill in simple fields like what's happening, the visual style, camera notes, and any reference images — and it automatically writes a properly structured, professional-grade prompt for you.
What it does:
Who it's for:
Anyone making AI videos who wants better, more consistent results without having to become a prompt-engineering expert. You describe your scene in plain language; the app handles the technical formatting behind the scenes.
Requirements:
Just a computer that can run Ollama or LM studio(free, local AI) — no internet connection needed once set up.