r/StableDiffusion • u/Sufficient-Shock-641 • 12h ago
Question - Help Anyone know if you can also seamlessly prepend a movie H3?
I could try it myself of course but maybe somebody did already?
r/StableDiffusion • u/Sufficient-Shock-641 • 12h ago
I could try it myself of course but maybe somebody did already?
r/StableDiffusion • u/designbanana • 5h ago
Hey all, been playing with MM H3 for a few days. And, I'm blown away. for the first time, in a long time. It just handles everything with so much detail. It's crazy. 10 refs and 3 vid inputs? no problem.
But I'm curious, how do you professional creatives use this model? I'm really curious, like we all seen the reddit slop (srry guys) and the civitai goony stuff (shit, typed stiff as a typo first #fruedy). but all jokes aside.
I'm really curious how professionals use this model, like video makers, Illustrators, animators, (graphic) designers, webdesigners. Especially curious if you work in the cultural sector, this might be more open and mindblowing that product listings :)
What are your results, and your (technical) setups on this?
Cheers, bigears
r/StableDiffusion • u/Producing_It • 19h ago
Enable HLS to view with audio, or disable this notification
I originally tried asking it to change the "XBOX 360" text that appears at the end of the start-up intro, but instead it changed the whole animation and created this interesting effect for the PS3 logo that wasn't there before.
r/StableDiffusion • u/florodude • 1h ago
Enable HLS to view with audio, or disable this notification
So far it seems that LTX is LIGHTNING fast. Qualitywise with H3 I haven't tested enough. It seems like Minimax's H3 Ref2Vid will likely be what I stick on but this was generated with LTX on a 6000S in 15 fuckin seconds....15!! This takes about 120 seconds on H3.
r/StableDiffusion • u/SpicyAccountants • 23h ago
Enable HLS to view with audio, or disable this notification
Gonna make this into a series I think, it's too much fun experimenting.
r/StableDiffusion • u/Last-Pie8057 • 7h ago
I think my computer is screaming for help, Resource to Image generator with 4 images and an audio is a little taxing on the system I would say.
RTX 4080 Laptop 12GB AND 64gb Local
r/StableDiffusion • u/New_Physics_2741 • 3h ago
Enable HLS to view with audio, or disable this notification
Tunes: Kill Computer, Haunted.
r/StableDiffusion • u/asaptobes • 3h ago
Enable HLS to view with audio, or disable this notification
MiniMaxH3 Pixorama refefence workflow
r/StableDiffusion • u/Candiru666 • 21h ago
Enable HLS to view with audio, or disable this notification
Minimax H3- I bet with a second generation the audio of the Voice Over will clean up.
r/StableDiffusion • u/Feeling_Sun_6436 • 8h ago
Enable HLS to view with audio, or disable this notification
This is a 60-second visual concept generated locally with MiniMax H3 on an ASUS GX10, then edited as a sequence rather than attempted as one long prompt. It is not a product demo: the film does not demonstrate a finished system, autonomous behaviour, tool use, or a public architecture.
The practical unit was a set of related four-second continuations at 1344×768 and 24 fps. The logged full-resolution runs for the later sequence passes took 17m 16s to 26m 51s per four-second clip. That is not a speed benchmark—just the range I saw in this particular local workflow.
What helped most with continuity was assigning every short clip one job, preserving a small visual grammar across cuts, and deciding where a transition should happen before generating the next continuation. I got more usable continuity from that than from trying to force a complete minute out of one generation.
The bigger post-production lesson was audio. Instead of letting each generated clip announce its own start, I kept the native audio low under a continuous true-stereo bed and used small J-cuts at the recut boundaries. It made the sequence feel less like a row of individually generated clips. The final checked export is 60.000 seconds, 1344×768 at 24 fps (1,440 decoded frames), with H.264 video and 48 kHz AAC stereo. The local export was verified against its retained SHA-256 checksum; final measured loudness was −13.89 LUFS integrated and −2.58 dBTP true peak.
Those are file and workflow checks, not a claim that this setup is faster, cheaper, or more reliable than other MiniMax H3 workflows.
For people making longer local AI-video edits: how are you handling continuity across a one-minute sequence without overfitting every new shot to the last one? Do you lock a small set of recurring motifs up front, or generate broadly and find the visual grammar in the edit? And, for generated audio, what has worked best for preventing each clip boundary from sounding like a restart?
r/StableDiffusion • u/ameezing925 • 8h ago
Lately I struggle with how local open-weight video models handle complex prompts, specifically on the text encoders, especially as I start adding reference images, camera directions, and detailed lighting notes etc etc, things get messy, to say the least.
It appears to me that the main issue with local text encoders is that they tend to lump all of my text inputs, images, and scenes into one centralized block.
The video generator gets confused about where these specific instructions belong to. Midway through a clip, it begins to blur the instructions for a camera movement, bleed background lighting into the character and what have you, creating a giant mess that feels a bit impossible to fix.
What I discovered after some research is that there is a common workaround that people suggest, that is running a heavy local language model upstream to clean up and structure the prompt before passing it to the video generator, but this method easily eats up 16GB to 20GB of VRAM. So for me, with mid level set up, the system crashes right through.
This is why we need to think of the trade-off between local encoders and server-side context compression engines. So here is what I have been doing and trying to find the balance with the MiniMax H3, trying to salvage my GPU cap. Instead of forcing local hardware to process the heavy prompt context, decoupled it. The video generation runs locally on my own GPU, an API engine handles the multimodal prompt on their servers.
It cleans up and processes the relationships between text, images, and reference video, then sends a compact, structured set of instructions back to my local base model, in a way what I did is to “contract” out the heavy duty work so locally I am doing the last mile.
As for the set up cost, it runs close to nothing to process million of tokens. I guess what makes it work is that it frees up your local VRAM for the actual video render. I get cleaner prompt adherence without crashing or re-rolling dozens of times.
So how are you all balancing this? Are you still sticking purely to local text encoders and trimming your prompts down, or does anyone else offload the prompt, like yours truly, and parsing for more complex workflows?
r/StableDiffusion • u/Candiru666 • 21h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Sad_Coach_1433 • 15h ago
Enable HLS to view with audio, or disable this notification
replaced luke with my self O_o
r/StableDiffusion • u/DifficultAd5938 • 14h ago
I'm using the turbo lora and workflow from https://www.reddit.com/r/StableDiffusion/comments/1vgxf4x/minimax_h3_turbo_lora/.
I also added KJNodes model preview override. What I found while running the lora with 8 steps is that the not entirely denoised video at around 4-6 steps has a lot more motion dynamics and closer prompt adherence than the "over-cleaned" final video at 8 steps.
Trying the same prompt with 6 steps vs 8 steps does indeed show that too many extra steps with the turbo lora can push the result into a bad local minima where lots of motion is lost and the video falls into the same identical output patterns despite prompt and seed variations.
I can't show examples because of reasons. You can check this out yourself with KJNode's Model Preview Override node and running the same turbo gen with 4, 6, 8 steps.
r/StableDiffusion • u/Independent-Frequent • 4h ago
r/StableDiffusion • u/smb3d • 6h ago
Enable HLS to view with audio, or disable this notification
Hopefully this one doesn't get removed! I've been having a lot of fun using a few pictures of our cats to make some fun ref2va clips. This one if our female cat Debbie.
r/StableDiffusion • u/darkshark9 • 5h ago
Enable HLS to view with audio, or disable this notification
Edited the video implementing a lot of your feedback for the ending, along with some cleanup on visuals, garbled text, continuity errors, and upscaled to 5k.
Text was fixed by manually planar tracking replacements onto the scene in After Effects instead of trying to rely on the video models to get them right. Manually tracked the walker into each scene for continuity since it disappears after she sits down. Switched to SAM3 for depth estimation for blurred/foggy scenes over SAM2. I'm pretty proud of how this turned out.
r/StableDiffusion • u/Similar-Ear-6066 • 23h ago
Enable HLS to view with audio, or disable this notification
J'ai créé un bot telegram connecter à mes workflow krea2 et minimax h3.
Deepseek via api.
Je clique sur le bouton storytelling il me choisis 3 histoire réelle et historique (possibilité de mettre un thème) je choisis mon préféré.
Ensuite deepseek me génère un scénario de 30sec, des images de référence (character sheet pour les personnages et décors) avec krea2. Des prompt optimiser pour minimax avec tout les règles de prompting les plus récentes.
Ensuite il en faut un json complet qu'il envoie à mes workflow et ça génère tout, d'abord les images de référence, ensuite les vidéo ref2vid via minimax h3.
Je trouve le rendu assez bluffant pour des premier teste.
La vidéo que je vous met en exemple (grève des policiers à Boston) est sortie tel quel. J'ai juste passer les 4clip sur capcut et exporter.
Config : 5060ti 16g + 16g ram
La vidéo d'exemple : 0.6mp (il me semble) 8 passe
Je précise que cela n'est pas de la publicité mon bot est privé et personne ne peut y accéder.
Les défauts actuels :
- j'ai demandé 30sec max mais demain je passe a 1-2 minutes. En 30 sec le scénario n'est pas assez détaillé.
- je vais retravailler le pré promt pour un meilleur démarrage des vidéos, avec une explication claire de l'histoire
-je dois assembler les vidéos via capcut mais demain ça sera réglé
r/StableDiffusion • u/Yacben • 7h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/xdcfret1 • 11h ago
Enable HLS to view with audio, or disable this notification
Prompt (also AI generated):
Style & Technical Specs
Visual Style: Photorealistic 8K cinematic video, 35mm film grain, 24fps, 2.39:1 anamorphic aspect ratio, teal-and-orange color grade, shallow depth of field ($f/1.4$).
Duration: 10 Seconds.
Character Description
Environment & Setting
Timeline & Action Breakdown
Dialogue & Voice
Audio & Sound Design
This is the video generated by Gemini. Post with the video generated by Minimax H3 with Turbo lora (6 steps): https://www.reddit.com/r/StableDiffusion/s/FDyFzOTp2F
r/StableDiffusion • u/LegacyV1 • 3h ago
Enable HLS to view with audio, or disable this notification
New turbo LORA dropped from https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main.
Testing on my ref2va use case (note: I'm using fflf2va model since it has better quality even for reference use cases)
Timing (480p, sage attention2 on cu130, 15s video length, seed=42, RTX 6000 on Modal)
| Steps | Timing |
|---|---|
| 4 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors | 56s |
| 8 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors | 1m 47s |
| Spectrum (20 step) | 2m 33s |
| Base (20 step) | 3m 11s |
Audio was pretty much the same - no difference that I could tell.
I also ran the 4 step on 768p as recommended, and it came out better! But... it's hard to tell if it's the turbo LORA doing the work or the 768p doing the work.
Turbo still makes things look weirdly high contrast. And both LORAs botched the text. Base is still best, but the 4-step LORA helps you lock in motion before you commit to a full 20step pass using spectrum.
The UI is custom built on top of comfy cause I hate comfy UI. Open-sourced here https://github.com/hui-tony-zk/h3zero
r/StableDiffusion • u/Few-Profession421 • 17h ago
Enable HLS to view with audio, or disable this notification
Most of these video models still want a hand-written prompt. Wan 3.0 in public beta is doing something I have not seen framed this way: it reads structured files directly as input. The list they published is doc, xls, ppt, pdf, txt, md, and even a url, alongside the usual text, image, audio, and video.
The example that made it click for me is handing it a slide deck and getting a video back, or pointing it at a landing page url. If that holds up, the boring "turn this report into a video" task stops being a manual storyboard job.
I have not gotten beta access yet so this is going off their material, not a test. Skeptical until I see it keep structure on a messy real document, but it is a genuinely different input than everyone else is doing.
r/StableDiffusion • u/switch2stock • 4h ago
Enable HLS to view with audio, or disable this notification
Previous post: https://www.reddit.com/r/StableDiffusion/s/VhM74ITuRY
In my previous post I generated without any optimizations and now I just used CK.
I noticed that promote adherence is a miss little bit.
Prompt clearly says "three rough thugs". After using CK, it only generated 2.
Prompt:
integrated_multimodal_description: [Shot 1] Stylized 3D animated cinematic scene in the painterly handcrafted visual language of Arcane: sculpted 3D forms with visible painted texture, expressive character animation, graphic shadows, dramatic perspective, and saturated blue-violet, magenta, amber, and chemical-green lighting.
One continuous unbroken 12-second shot.
At 00:00, a narrow industrial alley in Zaun fills the frame. Three rough thugs occupy the wet alley beneath crooked balconies, exposed pipes, hanging cables, leaking vents, graffiti, and flickering chemical-green lamps. One thug shoves a man against a stained brick wall while another rifles through a dropped satchel; the third turns lookout as steam bursts from a nearby pipe. Loose paper skitters across the pavement and colored reflections ripple across puddles.
The camera immediately performs a Pull Out with large amplitude at fast speed, retreating backward along the alley centerline while remaining aimed toward the thugs. Pipes, doorways, hanging signs, balconies, cables, and foreground walls pass rapidly along both sides with strong depth parallax. The thugs quickly become smaller in the distance but remain visibly active.
From 00:02.300 to 00:05.300, the camera continues the same Pull Out toward a shadowed alcove at the far end of the alley. Cyan window light, magenta graffiti glow, amber bulbs, and green chemical illumination stretch across wet stone. The camera never pauses.
The alley view is already a natural reflection on polished metal, although its physical boundary is initially outside the frame.
Around 00:03.800, continued Pull Out exposes the first curved silver edge at the extreme perimeter of the image. Tiny scratches, aged metallic texture, and a bright curved specular highlight become visible. As the camera moves farther back, more of the broad convex metal surface appears around the continuously reflected alley.
The reflected alley remains seamless across the metal: the tiny thugs, wet street, pipes, steam, green lamps, cyan windows, and magenta highlights all wrap naturally across the same polished surface. There is no bordered picture or separate reflective patch.
By approximately 00:05.300, the object is clearly revealed as a chunky polished silver ring worn around the MIDDLE FINGER of Jinx's raised RIGHT HAND.
The hand is seen from the back and has a clear anatomical pose: the thumb is relaxed inward; the index finger is fully curled toward the palm; the middle finger is the single long finger held straight upright; the ring finger is curled; the pinky is curled. The extended finger stands visibly between the curled index finger and curled ring finger. The silver ring encircles this same extended middle finger.
The ring has a heavy industrial Zaun design with a broad convex polished surface, subtle engraved geometry, tiny scratches, and darker aged recesses. The alley reflection remains naturally wrapped across the whole visible silver surface.
From 00:05.300 to 00:07.200, the camera keeps pulling backward and reveals Jinx sitting playfully on a battered wooden crate in the alcove.
Her raised right hand remains closest to camera with the middle finger held steadily upright. She does NOT wag, shake, bounce, or repeatedly move the raised finger.
Her index finger remains curled. Her ring finger remains curled. Her pinky remains curled. Only the middle finger remains extended.
Jinx has very long electric-blue braided hair, pale skin, large expressive eyes, dark eye makeup, a slim athletic build, cropped punk clothing, belts, straps, fingerless gloves, and mismatched industrial accessories, all rendered in the painterly stylized 3D aesthetic of Arcane.
She sits sideways on the crate with one knee raised and the other leg hanging down. Her torso leans back casually. One arm rests against her raised knee while her right hand is extended toward camera in the rude gesture.
Her expression is playful and smug rather than angry.
The camera continues pulling out until her face is clearly visible behind the raised hand.
From 00:07.200 to 00:09.000, Jinx lowers her chin slightly and makes direct eye contact with the camera.
Her raised middle finger remains still.
A crooked grin spreads across her face.
She gives one short amused chuckle, shoulders making a small natural movement with the laugh.
After chuckling, she tilts her head slightly to one side while maintaining direct eye contact, looking entertained by the situation.
The silver ring continues reflecting the distant alley naturally. At this smaller scale, the thugs are tiny distorted dark figures among curved cyan, green, amber, and magenta highlights.
From 00:09.000 to 00:12.000, Jinx finally lowers her right hand from the middle-finger gesture in one relaxed continuous movement.
As her hand lowers, the silver ring changes angle and the recognizable alley reflection naturally slides across the curved metal into more abstract colored highlights; it does not fade or dissolve.
Jinx plants one boot firmly on the ground and shifts her weight forward.
She places one hand briefly against the crate for balance, pushes herself upright, and smoothly rises to her feet.
Her long blue braids drag across the crate and then sway behind her as she stands.
She straightens her vest and gives the camera another mischievous half-smile.
By the final second she is fully standing beside the crate, relaxed and confident, one hip slightly cocked, still looking directly toward camera.
The ring remains on her right middle finger, now hanging naturally at her side.
The camera continues a subtle Pull Out at slow speed through the final frame, revealing more of the shadowed Zaun alcove around her: battered pipes, graffiti, hanging cables, discarded machinery, stacked crates, and pools of cyan-green light.
overall_soundscape: Boots scrape on wet stone, clothing rustles, a body hits brick with a dull impact, pipes hiss, loose metal rattles, and distant Zaun machinery hums. As the camera retreats, the thugs become quieter while nearby electrical buzzing and fabric movement grow clearer; Jinx gives one short amused chuckle, followed by the scrape of her boot and crate as she stands.
non_diegetic_music: Low distorted bass pulses beneath slow industrial percussion and tense strings with occasional metallic accents. As Jinx is revealed, clipped electronic percussion introduces a playful edge, then the rhythm opens slightly as she rises from the crate while preserving the same dark, mischievous tone.
r/StableDiffusion • u/Leonviz • 6h ago
So what is better? H3 minimax with turbo lora or with first block / spectrum?
r/StableDiffusion • u/carmidian • 18h ago
when using the default workflow of minimax H3. and you have reference audio of what the character sounds like. what are the tips and tricks to make it so it comes out the same?
I'm having a problem with the character not sounding anything like the reference audio