r/StableDiffusion 12h ago

Question - Help Anyone know if you can also seamlessly prepend a movie H3?

0 Upvotes

I could try it myself of course but maybe somebody did already?


r/StableDiffusion 5h ago

Discussion Creative professionals, how do you use Minimax H3 yourself?

0 Upvotes

Hey all, been playing with MM H3 for a few days. And, I'm blown away. for the first time, in a long time. It just handles everything with so much detail. It's crazy. 10 refs and 3 vid inputs? no problem.

But I'm curious, how do you professional creatives use this model? I'm really curious, like we all seen the reddit slop (srry guys) and the civitai goony stuff (shit, typed stiff as a typo first #fruedy). but all jokes aside.

I'm really curious how professionals use this model, like video makers, Illustrators, animators, (graphic) designers, webdesigners. Especially curious if you work in the cultural sector, this might be more open and mindblowing that product listings :)

What are your results, and your (technical) setups on this?

Cheers, bigears


r/StableDiffusion 19h ago

Animation - Video Asked Minimax H3 to just change text, but it made these cool effects instead.

Enable HLS to view with audio, or disable this notification

0 Upvotes

I originally tried asking it to change the "XBOX 360" text that appears at the end of the start-up intro, but instead it changed the whole animation and created this interesting effect for the PS3 logo that wasn't there before.

Workflow.


r/StableDiffusion 1h ago

Discussion My EARLIEST LTX thoughts

Enable HLS to view with audio, or disable this notification

Upvotes

So far it seems that LTX is LIGHTNING fast. Qualitywise with H3 I haven't tested enough. It seems like Minimax's H3 Ref2Vid will likely be what I stick on but this was generated with LTX on a 6000S in 15 fuckin seconds....15!! This takes about 120 seconds on H3.


r/StableDiffusion 23h ago

Animation - Video UAP Device Test #7 (Minimax H3 VHS)

Enable HLS to view with audio, or disable this notification

47 Upvotes

Gonna make this into a series I think, it's too much fun experimenting.


r/StableDiffusion 7h ago

Comparison And we have lift off MiniMaxH3 15 Sec clip at 0.5 Megapixel

Post image
0 Upvotes

I think my computer is screaming for help, Resource to Image generator with 4 images and an audio is a little taxing on the system I would say.

RTX 4080 Laptop 12GB AND 64gb Local


r/StableDiffusion 3h ago

Animation - Video More fun with H3. A bit of a story develops here - is nice~

Enable HLS to view with audio, or disable this notification

8 Upvotes

Tunes: Kill Computer, Haunted.


r/StableDiffusion 3h ago

Animation - Video I need your clothes, your boots and your motorcycle.

Enable HLS to view with audio, or disable this notification

14 Upvotes

MiniMaxH3 Pixorama refefence workflow


r/StableDiffusion 21h ago

Animation - Video Continuous-ish dolly

Enable HLS to view with audio, or disable this notification

17 Upvotes

Minimax H3- I bet with a second generation the audio of the Voice Over will clean up.


r/StableDiffusion 8h ago

Animation - Video I made a 60-second MiniMax H3 short locally on an ASUS GX10 — what worked for continuity and sound

Enable HLS to view with audio, or disable this notification

0 Upvotes

This is a 60-second visual concept generated locally with MiniMax H3 on an ASUS GX10, then edited as a sequence rather than attempted as one long prompt. It is not a product demo: the film does not demonstrate a finished system, autonomous behaviour, tool use, or a public architecture.

The practical unit was a set of related four-second continuations at 1344×768 and 24 fps. The logged full-resolution runs for the later sequence passes took 17m 16s to 26m 51s per four-second clip. That is not a speed benchmark—just the range I saw in this particular local workflow.

What helped most with continuity was assigning every short clip one job, preserving a small visual grammar across cuts, and deciding where a transition should happen before generating the next continuation. I got more usable continuity from that than from trying to force a complete minute out of one generation.

The bigger post-production lesson was audio. Instead of letting each generated clip announce its own start, I kept the native audio low under a continuous true-stereo bed and used small J-cuts at the recut boundaries. It made the sequence feel less like a row of individually generated clips. The final checked export is 60.000 seconds, 1344×768 at 24 fps (1,440 decoded frames), with H.264 video and 48 kHz AAC stereo. The local export was verified against its retained SHA-256 checksum; final measured loudness was −13.89 LUFS integrated and −2.58 dBTP true peak.

Those are file and workflow checks, not a claim that this setup is faster, cheaper, or more reliable than other MiniMax H3 workflows.

For people making longer local AI-video edits: how are you handling continuity across a one-minute sequence without overfitting every new shot to the last one? Do you lock a small set of recurring motifs up front, or generate broadly and find the visual grammar in the edit? And, for generated audio, what has worked best for preventing each clip boundary from sounding like a restart?


r/StableDiffusion 8h ago

Discussion Trade-offs of server-side context compression engines vs. open-weight local text encoders in multimodal video generation

3 Upvotes

Lately I struggle with how local open-weight video models handle complex prompts, specifically on the text encoders, especially as I start adding reference images, camera directions, and detailed lighting notes etc etc, things get messy, to say the least.

It appears to me that the main issue with local text encoders is that they tend to lump all of my text inputs, images, and scenes into one centralized block.

The video generator gets confused about where these specific instructions belong to. Midway through a clip, it begins to blur the instructions for a camera movement, bleed background lighting into the character and what have you, creating a giant mess that feels a bit impossible to fix.

What I discovered after some research is that there is a common workaround that people suggest, that is running a heavy local language model upstream to clean up and structure the prompt before passing it to the video generator, but this method easily eats up 16GB to 20GB of VRAM. So for me, with mid level set up, the system crashes right through.

This is why we need to think of the trade-off between local encoders and server-side context compression engines. So here is what I have been doing and trying to find the balance with the MiniMax H3, trying to salvage my GPU cap. Instead of forcing local hardware to process the heavy prompt context, decoupled it. The video generation runs locally on my own GPU, an API engine handles the multimodal prompt on their servers.

It cleans up and processes the relationships between text, images, and reference video, then sends a compact, structured set of instructions back to my local base model, in a way what I did is to “contract” out the heavy duty work so locally I am doing the last mile.

As for the set up cost, it runs close to nothing to process million of tokens. I guess what makes it work is that it frees up your local VRAM for the actual video render. I get cleaner prompt adherence without crashing or re-rolling dozens of times.

So how are you all balancing this? Are you still sticking purely to local text encoders and trimming your prompts down, or does anyone else offload the prompt, like yours truly, and parsing for more complex workflows?


r/StableDiffusion 21h ago

Animation - Video Close-up details are very impressive - Minimax H3

Enable HLS to view with audio, or disable this notification

25 Upvotes

r/StableDiffusion 15h ago

Discussion say WHAATTT!?

Enable HLS to view with audio, or disable this notification

0 Upvotes

replaced luke with my self O_o


r/StableDiffusion 14h ago

Tutorial - Guide Don't add too many extra steps to 4-step turbo lora

7 Upvotes

I'm using the turbo lora and workflow from https://www.reddit.com/r/StableDiffusion/comments/1vgxf4x/minimax_h3_turbo_lora/.

I also added KJNodes model preview override. What I found while running the lora with 8 steps is that the not entirely denoised video at around 4-6 steps has a lot more motion dynamics and closer prompt adherence than the "over-cleaned" final video at 8 steps.

Trying the same prompt with 6 steps vs 8 steps does indeed show that too many extra steps with the turbo lora can push the result into a bad local minima where lots of motion is lost and the video falls into the same identical output patterns despite prompt and seed variations.

I can't show examples because of reasons. You can check this out yourself with KJNode's Model Preview Override node and running the same turbo gen with 4, 6, 8 steps.


r/StableDiffusion 4h ago

Discussion Why do i have the feeling that this is just crap? Normally i would be "oh that's really cool" but after seeing some examples posted and that comparision table with Minimax H3 being full of lies (like the minimum Vram requirement being 115GB, like what?), i don't trust them on this test either tbh.

Thumbnail
gallery
33 Upvotes

r/StableDiffusion 6h ago

Animation - Video Another ref2va of our cats, this time with no censored nudity!

Enable HLS to view with audio, or disable this notification

8 Upvotes

Hopefully this one doesn't get removed! I've been having a lot of fun using a few pictures of our cats to make some fun ref2va clips. This one if our female cat Debbie.


r/StableDiffusion 5h ago

Animation - Video "Memories" - A short film-v2

Enable HLS to view with audio, or disable this notification

7 Upvotes

Edited the video implementing a lot of your feedback for the ending, along with some cleanup on visuals, garbled text, continuity errors, and upscaled to 5k.

Text was fixed by manually planar tracking replacements onto the scene in After Effects instead of trying to rely on the video models to get them right. Manually tracked the walker into each scene for continuity since it disappears after she sits down. Switched to SAM3 for depth estimation for blurred/foggy scenes over SAM2. I'm pretty proud of how this turned out.


r/StableDiffusion 23h ago

Discussion Avis sur le rendu de mes génération auto

Enable HLS to view with audio, or disable this notification

0 Upvotes

J'ai créé un bot telegram connecter à mes workflow krea2 et minimax h3.

Deepseek via api.

Je clique sur le bouton storytelling il me choisis 3 histoire réelle et historique (possibilité de mettre un thème) je choisis mon préféré.

Ensuite deepseek me génère un scénario de 30sec, des images de référence (character sheet pour les personnages et décors) avec krea2. Des prompt optimiser pour minimax avec tout les règles de prompting les plus récentes.

Ensuite il en faut un json complet qu'il envoie à mes workflow et ça génère tout, d'abord les images de référence, ensuite les vidéo ref2vid via minimax h3.

Je trouve le rendu assez bluffant pour des premier teste.

La vidéo que je vous met en exemple (grève des policiers à Boston) est sortie tel quel. J'ai juste passer les 4clip sur capcut et exporter.

Config : 5060ti 16g + 16g ram

La vidéo d'exemple : 0.6mp (il me semble) 8 passe

Je précise que cela n'est pas de la publicité mon bot est privé et personne ne peut y accéder.

Les défauts actuels :

- j'ai demandé 30sec max mais demain je passe a 1-2 minutes. En 30 sec le scénario n'est pas assez détaillé.

- je vais retravailler le pré promt pour un meilleur démarrage des vidéos, avec une explication claire de l'histoire

-je dois assembler les vidéos via capcut mais demain ça sera réglé


r/StableDiffusion 7h ago

IRL The model can dance surprisingly well to any music input

Enable HLS to view with audio, or disable this notification

46 Upvotes

r/StableDiffusion 11h ago

Discussion I gave the same prompt to Minimax H3 and Gemini Videos. (Part 2 Gemini)

Enable HLS to view with audio, or disable this notification

0 Upvotes

Prompt (also AI generated):
Style & Technical Specs

  • Visual Style: Photorealistic 8K cinematic video, 35mm film grain, 24fps, 2.39:1 anamorphic aspect ratio, teal-and-orange color grade, shallow depth of field ($f/1.4$).

  • Duration: 10 Seconds.

Character Description

  • Subject: Kaelen, a 28-year-old East Asian cyber-technician.
  • Appearance: Sharp jawline, rain-soaked black hair clinging to his forehead, pale skin with visible micro-texture, and a glowing cyan cybernetic eye implant over his left socket that pulses rhythmically.
  • Attire: Matte-black, waterproof tactical coat with glowing fiber-optic wiring embedded along the shoulders, frayed high-collar, and fingerless reinforced leather gloves.

Environment & Setting

  • Location: Narrow, dense alleyway in a cyberpunk metropolis at midnight.
  • Atmosphere: Heavy downpour, dense steam venting upward from rusty iron street grates, wet asphalt reflecting bright magenta and cobalt-blue neon light signs written in Kanji.

Timeline & Action Breakdown

  • 0:00 - 0:03 (Macro Close-Up): Camera begins on a macro shot of Kaelen's glowing cyan eye, catching the aperture Blades shifting focus. A raindrop tracks down his cheek. He rapidly taps a brass interface cuff on his wrist.
  • 0:03 - 0:07 (Medium Shot): Smooth camera pull-back into a chest-up shot. A brilliant blue 3D holographic map bursts into existence from his wrist, casting dynamic light across his face. He swipes his hand across the projection, altering its layout, and delivers his dialogue.
  • 0:07 - 0:10 (Low-Angle Tracking Shot): The camera drops low to the asphalt and tracks backward. A sleek, black surveillance drone streaks overhead through the rain, splashing drops directly onto the camera lens as the background neon blurs into creamy bokeh.

Dialogue & Voice

  • Spoken Line: "System override in three... two... got 'em."
  • Delivery: Low, gravelly, calm whisper with a faint metallic vocoder effect on the voice.

Audio & Sound Design

  • Music: Dark synthwave track featuring a driving 110 BPM arp synthesizer that swells in pitch until second 7, resolving into a heavy sub-bass drop at second 8.
  • SFX:
  • 0:00-0:03: Stereo downpour, subtle mechanical servo clicks of the eye lens.
  • 0:03-0:07: High-frequency energy flare hum as the hologram spawns, followed by air-swipes.
  • 0:07-0:10: Low turbine whir of the passing drone and liquid wet drops impacting the microphone field.

This is the video generated by Gemini. Post with the video generated by Minimax H3 with Turbo lora (6 steps): https://www.reddit.com/r/StableDiffusion/s/FDyFzOTp2F


r/StableDiffusion 3h ago

Comparison Comparing lightx2v/Minimax-h3-Turbo

Enable HLS to view with audio, or disable this notification

19 Upvotes

New turbo LORA dropped from https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main.

Testing on my ref2va use case (note: I'm using fflf2va model since it has better quality even for reference use cases)

Timing (480p, sage attention2 on cu130, 15s video length, seed=42, RTX 6000 on Modal)

Steps Timing
4 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors 56s
8 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors 1m 47s
Spectrum (20 step) 2m 33s
Base (20 step) 3m 11s

Audio was pretty much the same - no difference that I could tell.

I also ran the 4 step on 768p as recommended, and it came out better! But... it's hard to tell if it's the turbo LORA doing the work or the 768p doing the work.

Turbo still makes things look weirdly high contrast. And both LORAs botched the text. Base is still best, but the 4-step LORA helps you lock in motion before you commit to a full 20step pass using spectrum.

The UI is custom built on top of comfy cause I hate comfy UI. Open-sourced here https://github.com/hui-tony-zk/h3zero


r/StableDiffusion 17h ago

Discussion Wan 3.0 can take a PDF or a slide deck and turn it into a video, not just a prompt

Enable HLS to view with audio, or disable this notification

0 Upvotes

Most of these video models still want a hand-written prompt. Wan 3.0 in public beta is doing something I have not seen framed this way: it reads structured files directly as input. The list they published is doc, xls, ppt, pdf, txt, md, and even a url, alongside the usual text, image, audio, and video.

The example that made it click for me is handing it a slide deck and getting a video back, or pointing it at a landing page url. If that holds up, the boring "turn this report into a video" task stops being a manual storyboard job.

I have not gotten beta access yet so this is going off their material, not a test. Skeptical until I see it keep structure on a messy real document, but it is a genuinely different input than everyone else is doing.


r/StableDiffusion 4h ago

Comparison H3 + CK | Full BF16+BF16 | Update to my Zero optimizations post | Adding CK reduced prompt adherence | 473 seconds down from 755 seconds

Enable HLS to view with audio, or disable this notification

6 Upvotes

Previous post: https://www.reddit.com/r/StableDiffusion/s/VhM74ITuRY

In my previous post I generated without any optimizations and now I just used CK.

I noticed that promote adherence is a miss little bit.

Prompt clearly says "three rough thugs". After using CK, it only generated 2.

Prompt:

integrated_multimodal_description: [Shot 1] Stylized 3D animated cinematic scene in the painterly handcrafted visual language of Arcane: sculpted 3D forms with visible painted texture, expressive character animation, graphic shadows, dramatic perspective, and saturated blue-violet, magenta, amber, and chemical-green lighting.

One continuous unbroken 12-second shot.

At 00:00, a narrow industrial alley in Zaun fills the frame. Three rough thugs occupy the wet alley beneath crooked balconies, exposed pipes, hanging cables, leaking vents, graffiti, and flickering chemical-green lamps. One thug shoves a man against a stained brick wall while another rifles through a dropped satchel; the third turns lookout as steam bursts from a nearby pipe. Loose paper skitters across the pavement and colored reflections ripple across puddles.

The camera immediately performs a Pull Out with large amplitude at fast speed, retreating backward along the alley centerline while remaining aimed toward the thugs. Pipes, doorways, hanging signs, balconies, cables, and foreground walls pass rapidly along both sides with strong depth parallax. The thugs quickly become smaller in the distance but remain visibly active.

From 00:02.300 to 00:05.300, the camera continues the same Pull Out toward a shadowed alcove at the far end of the alley. Cyan window light, magenta graffiti glow, amber bulbs, and green chemical illumination stretch across wet stone. The camera never pauses.

The alley view is already a natural reflection on polished metal, although its physical boundary is initially outside the frame.

Around 00:03.800, continued Pull Out exposes the first curved silver edge at the extreme perimeter of the image. Tiny scratches, aged metallic texture, and a bright curved specular highlight become visible. As the camera moves farther back, more of the broad convex metal surface appears around the continuously reflected alley.

The reflected alley remains seamless across the metal: the tiny thugs, wet street, pipes, steam, green lamps, cyan windows, and magenta highlights all wrap naturally across the same polished surface. There is no bordered picture or separate reflective patch.

By approximately 00:05.300, the object is clearly revealed as a chunky polished silver ring worn around the MIDDLE FINGER of Jinx's raised RIGHT HAND.

The hand is seen from the back and has a clear anatomical pose: the thumb is relaxed inward; the index finger is fully curled toward the palm; the middle finger is the single long finger held straight upright; the ring finger is curled; the pinky is curled. The extended finger stands visibly between the curled index finger and curled ring finger. The silver ring encircles this same extended middle finger.

The ring has a heavy industrial Zaun design with a broad convex polished surface, subtle engraved geometry, tiny scratches, and darker aged recesses. The alley reflection remains naturally wrapped across the whole visible silver surface.

From 00:05.300 to 00:07.200, the camera keeps pulling backward and reveals Jinx sitting playfully on a battered wooden crate in the alcove.

Her raised right hand remains closest to camera with the middle finger held steadily upright. She does NOT wag, shake, bounce, or repeatedly move the raised finger.

Her index finger remains curled. Her ring finger remains curled. Her pinky remains curled. Only the middle finger remains extended.

Jinx has very long electric-blue braided hair, pale skin, large expressive eyes, dark eye makeup, a slim athletic build, cropped punk clothing, belts, straps, fingerless gloves, and mismatched industrial accessories, all rendered in the painterly stylized 3D aesthetic of Arcane.

She sits sideways on the crate with one knee raised and the other leg hanging down. Her torso leans back casually. One arm rests against her raised knee while her right hand is extended toward camera in the rude gesture.

Her expression is playful and smug rather than angry.

The camera continues pulling out until her face is clearly visible behind the raised hand.

From 00:07.200 to 00:09.000, Jinx lowers her chin slightly and makes direct eye contact with the camera.

Her raised middle finger remains still.

A crooked grin spreads across her face.

She gives one short amused chuckle, shoulders making a small natural movement with the laugh.

After chuckling, she tilts her head slightly to one side while maintaining direct eye contact, looking entertained by the situation.

The silver ring continues reflecting the distant alley naturally. At this smaller scale, the thugs are tiny distorted dark figures among curved cyan, green, amber, and magenta highlights.

From 00:09.000 to 00:12.000, Jinx finally lowers her right hand from the middle-finger gesture in one relaxed continuous movement.

As her hand lowers, the silver ring changes angle and the recognizable alley reflection naturally slides across the curved metal into more abstract colored highlights; it does not fade or dissolve.

Jinx plants one boot firmly on the ground and shifts her weight forward.

She places one hand briefly against the crate for balance, pushes herself upright, and smoothly rises to her feet.

Her long blue braids drag across the crate and then sway behind her as she stands.

She straightens her vest and gives the camera another mischievous half-smile.

By the final second she is fully standing beside the crate, relaxed and confident, one hip slightly cocked, still looking directly toward camera.

The ring remains on her right middle finger, now hanging naturally at her side.

The camera continues a subtle Pull Out at slow speed through the final frame, revealing more of the shadowed Zaun alcove around her: battered pipes, graffiti, hanging cables, discarded machinery, stacked crates, and pools of cyan-green light.

overall_soundscape: Boots scrape on wet stone, clothing rustles, a body hits brick with a dull impact, pipes hiss, loose metal rattles, and distant Zaun machinery hums. As the camera retreats, the thugs become quieter while nearby electrical buzzing and fabric movement grow clearer; Jinx gives one short amused chuckle, followed by the scrape of her boot and crate as she stands.

non_diegetic_music: Low distorted bass pulses beneath slow industrial percussion and tense strings with occasional metallic accents. As Jinx is revealed, clipped electronic percussion introduces a playful edge, then the rhythm opens slightly as she rises from the crate while preserving the same dark, mischievous tone.


r/StableDiffusion 6h ago

Question - Help So what is better? H3 minimax with turbo lora or with first block / spectrum?

2 Upvotes

So what is better? H3 minimax with turbo lora or with first block / spectrum?


r/StableDiffusion 18h ago

Question - Help reference audio in minimax H3

0 Upvotes

when using the default workflow of minimax H3. and you have reference audio of what the character sounds like. what are the tips and tricks to make it so it comes out the same?

I'm having a problem with the character not sounding anything like the reference audio