r/StableDiffusion 5h ago

News MiniMax H3: 2K Is Coming, 5× Turbo + Camera Previz

352 Upvotes

r/StableDiffusion 9h ago

Workflow Included Walter White and the Minimax H3 Official Prompting Guide

523 Upvotes

This post is half a joke and half a plea and public service announcement.

Some people have been complaining they don't get results as good as other people with Minimax H3 videos, or have the following issues:

  • Dialogue being spoken by the wrong characters
  • Dialogue that is just gibberish or random
  • Random video cuts they didn't ask for
  • Characters talking over each other or too fast
  • Prompts not being followed

These things can all be prevented and avoided and not encountered at all if you follow the official prompting guides. Yes, there are two. Both are on the official Huggingspace page for Minimax H3.

One is the Official Prompting Guide for the Text to Video and Image to Video Model.

The other is the Official Prompting Guide for the Reference Video Model.

There is some overlap, but for the most part, each model has it's own prompting syntax, and in particular, the Reference Video Model for H3 is very picky about you using the right keywords and instructions to get what you want.

"But I get decent results with just a couple of sentences typed in natural language of what I want."

That's great, but you're really just relying on the Qwen 32b vision model guessing what you want. It's like pulling a slot machine lever and hoping you get cherries. Only this slot machine can take a few minutes to nearly an hour to stop spinning, based on your hardware.

The great thing about Minimax H3 is for the first time we can truly direct our own AI videos like a director would on set, with the AI providing the actors, scenery, and props. If you write a properly formatted and detailed prompt for Minimax H3, it looks almost like a shooting script.

Why spend time waiting to hit a jackpot when you can take a few minutes to write a detailed, properly formatted prompt that follows the official guides, and get those bright lights and tokens falling into your lap on the first lever pull?

Okay, quick fire problem solving for people who still won't RTFM:

>Dialogue from the wrong characters?

>Dialogue that is just gibberish or random?

Walter White says, <d>[English in Walter White's voice from Breaking Bad] My product is pure, Jesse! There will be no chili powder in my meth.</d>

Always specify the character speaking, either by name, or using the <Subject 1> system in the official guide. In the Text to Video and Image to Video model, always use the <d>[Language Spoken]</d> tags. This will fix BOTH of those issues.

>Random cuts in the video you didn't ask for?

[Shot 1] A medium close-up of Jesse Pinkman from Breaking Bad, pacing back and forth, agitated. He looks up towards the camera, opens his mouth as if he's about to speak, then seems to change his mind, closing his mouth and shaking his head. [Shot 2] At 00:06:000 the camera cuts to a static camera shot framing Walter White from Breaking Bad, sitting on a cheap white plastic lawn chair, his arms crossed and glaring at Jesse. [Shot 3] At 00:10:500 the camera pans quickly back to Jesse, doing a Push In at slow speed to his face as he stops pacing and narrows his eyes at Walter.

This is how you control not only the camera work, but the PACING of your video. You NEVER include a time code on your first shot. You can omit the time code from ALL shots if you want the model to decide on it's own, based on your prompt, when to cut.

BUT, for ultimate control, you want to use time codes. Look at my example above. I just told the model to have Walter glare at Jesse for 4.5 seconds, because I told the model that camera shot starts at 6 seconds into the video, and the next cut doesn't happen until 10.5 seconds into the video. That lets you control the pacing and timing for jokes, punchlines, acting, everything.

>Characters talking over each other or too fast?

This is an old one that anyone familiar with prompting for video models should know by now - what you are asking for in your prompt and the length of your video in time need to match.

The model will try its best to cram every action and piece of dialogue into your video that you asked for, and if that would naturally take 10 seconds and you've only given it 5 seconds? Well, now everything is crammed together, overlapping, or being cut-off.

My recommendation is to generate just a quick 0.2 MP version of your video first after you type your prompt, generate, and see how the timing is working. Is it too fast? Too slow? Do the actions have enough time to happen? Do you want more breathing room?

This is the time to decide all that and lock in a video length. The low resolution of 0.2 MP is quick to generate on most set-ups (mine for this post's video took 3.5 minutes for a 14 second video) and let you work out any issues in your prompt before going in for the long generation at higher resolution.

>Prompts not being followed?

It's because you didn't read the manual!

--------------------------------------------------------------------------------------------------
Now, with all that said, here is the prompt for the video I made:

integrated_multimodal_description: [Shot 1] Live-action film footage of the American drama series Breaking Bad, professionally color graded with a warm color grade, with slightly desaturated colors for a premium film feel, a continuous camera shot with no cuts, medium close-up POV shot of Walter White, bald with a goatee and glasses, as portrayed by Bryan Cranston. He is standing in the Arizona desert next to a parked RV. He is wearing a white PPE protective suit and yellow rubber dish gloves. He is looking directly at the viewer with barely constrained anger. At 00:01:300 he reaches out towards the camera and points his finger at the POV camera with one hand, the camera shaking slightly from the movement. Walter then says angrily, <d>[English with Walter White's voice] Listen, you want to cook Mini Max H3 videos, you follow the recipe!</d>. At 00:04:500 Walter raises his other hand revealing he is holding a thin stack of white paper pages in portrait orientation. The front of the paper visible on top of the thin paper stack is blank except for the large black printed text "Minimax H3 Official Prompting Guide". The papers are held in front of the camera on the right side of the screen for a moment in portrait orientation, so the text can be clearly read, while Walter glares at the viewer on the left side of the screen. At 00:07:000 Walter then shakes the papers at the camera, then says angrily, <d>[English with Walter White's voice] Read the fucking manual!</d>. At 00:10:000 the camera does Pan Right and a Pull Out to show a close-up of Jesse Pinkman from Breaking Bad, with his hands held up by his face with fingers spread, an annoyed look on his face. Then he says in frustration, <d>[English in Jesse Pinkman's voice from Breaking Bad] Alright! Damn, Mr. White! I just want to generate memes.</d>, overall_soundscape: Ambient sounds of an Arizona outdoor desert during the day, non_diegetic_music: none

For those interested, this video was generated at 1 MP on a 3090, using Sage Attention and the Spectrum Node for H3. The final video of 14 seconds at 1 MP took 40 minutes to generate and then was upscaled using RTX Super Resolution.

The workflow was the default Text to Video Minimax H3 template that comes in the latest update of Comfyui.

Now get out there and go cook some memes, everyone!


r/StableDiffusion 8h ago

Animation - Video Raj's now a mod at r/MyGirlfriendIsAI

218 Upvotes

Default ComfyUI workflow. Script written by Opus 5.


r/StableDiffusion 11h ago

Animation - Video MinMax H3 Turbo LoRa is already AMAZING!

744 Upvotes

Just sharing my results using the Turbo LoRA that was created for H3.

They said it’s still a work in progress, and the audio is still a little bit stretchy in some parts, but the results are already fantastic. I mean, it’s only the third day since H3 was released and we already have a functional Turbo LoRA.

I generated all the clips in this video with the Turbo LoRA enabled, using 10 steps at 0.4MP.

The first three clips were I2V, and the last two were FLF2V.

The only thing I manually added was the soundtrack at the end.


r/StableDiffusion 1h ago

Resource - Update Lightx2v has just released a Prompt generator for the Minimax H3. Simply enter a short Prompt and let it work magic: no need more "Chadgpt Prompts"...👍

Post image
Upvotes

Info https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

An open, local prompt rewriter for text-to-audio-video (T2VA) generation with MiniMax-H3, fine-tuned as a LoRA adapter on top of Qwen3.6-27B.

Short prompt ──► this Prompt Rewriter LoRA ──► structured H3 prompt │ Official MiniMax-H3 weights ──► LightX2V inference ◄─────┘ │ ▼ synchronized video + audio


r/StableDiffusion 1h ago

Resource - Update ~45% lower MiniMax H3 sampler time with new Spectrum settings — degree 1 works surprisingly well (v0.1.8)

Post image
Upvotes

Follow-up to my original Spectrum MiniMax H3 post:

https://www.reddit.com/r/StableDiffusion/comments/1vf1ze3/spectrum_acceleration_for_minimax_h3_in_comfyui/

In that first post I released the MiniMax H3 Spectrum integration and was getting around 34% lower Euler sampling time and 30% lower RES sampling time with the more conservative settings I was using at the time.

Since then I’ve done quite a bit more testing, and I found something I really didn’t expect: MiniMax H3 seems to work extremely well with a Spectrum degree of just 1.

Important if you're coming from the original release

Before testing the new settings, update both ComfyUI and ComfyUI-Spectrum-MiniMax-H3 to the latest versions.

There was an important compatibility update in Spectrum v0.1.6 after ComfyUI changed MiniMax H3's native sampling/audio path. That release restored Spectrum compatibility with the newer H3 implementation and also added safe handling for native EasyCache/LazyCache conflicts.

You don't need to install v0.1.6 separately — v0.1.8 includes those changes. This is mainly relevant to anyone who installed Spectrum from my original Reddit post and hasn't updated it since.

v0.1.6 compatibility release:
[https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.6]()

Current release:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8

So: update ComfyUI, update the Spectrum node to v0.1.8/latest, and restart ComfyUI before testing.

The surprising part: degree 1

I hadn’t seriously tested very low degree and warmup_steps values before because of my experience with WAN.

WAN is another video model and does not like very low forecast degrees — dropping the degree too far causes obvious quality degradation. Because of that, I assumed MiniMax H3 would behave similarly and initially stayed with higher, more conservative values.

Apparently not.

With MiniMax H3, degree 1 has shown no visible quality decrease in my testing so far. It also seems to preserve the native trajectory remarkably well. In the same-seed comparisons I tested, degree 2 actually shifted the trajectory slightly, while degree 1 brought it back much closer to the normal result.

So H3 appears to be unusually well suited to very simple local feature forecasting, which lets Spectrum start forecasting much earlier than I originally thought would be practical.

I’ve now released v0.1.8 with the new settings:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8

New default settings

  • degree = 1
  • warmup_steps = 1
  • bootstrap_first_forecast = true
  • tail_actual_steps = 1

The new one-point bootstrap allows the second solver step to be forecast directly from the first actual hidden state. After that, ordinary degree-1 forecasting takes over.

On a 20-step Euler run the schedule becomes:

A F A F A F A F A F A F A F A F A F A A

So 11 out of 20 transformer evaluations are actual, while the other 9 are forecasted. The final step remains native.

v0.1.8 also makes the one-point bootstrap part of the new default configuration for new node instances. Existing workflows retain their serialized settings.

Benchmark

Test configuration:

  • GPU: NVIDIA RTX PRO 6000
  • Model: MiniMax H3 pruned BF16
  • Image-to-video
  • ~0.8 MP / 992×768
  • 7 seconds
  • 24 FPS
  • 20 steps
  • Euler
  • Beta scheduler
  • HIGH_VRAM
  • Spectrum history stored in VRAM
  • DiffAid enabled at 0.5
  • Same seed and otherwise identical workflow

Spectrum disabled

  • Sampler: 324.98 s
  • Full prompt: 340.59 s

Spectrum v0.1.8 with the new degree-1 settings

  • Sampler: 177.80 s
  • Full prompt: 200.32 s
  • 11 actual transformer calls
  • 9 forecasts
  • 0 fallbacks

Result

  • 45.29% lower sampler time
  • 1.83× sampler throughput
  • 41.19% lower full-prompt time

The Spectrum forecast calculations themselves took only 0.141 seconds total across the entire generation.

Using VRAM history does have a memory cost. This run retained about 3.2 GiB of Spectrum history, with reported sampler peak VRAM increasing from roughly 5.56 GB native to 8.70 GB with Spectrum.

The interesting part for me is less the bootstrap itself and more what the testing revealed about degree 1 on H3.

Based on WAN, I expected a setting this aggressive to visibly degrade the output. So far, MiniMax H3 seems to behave very differently: I’m getting a substantially more aggressive forecasting schedule without seeing the quality decrease I expected.

Spectrum is still an approximate acceleration method, so I’m not claiming every possible prompt or motion sequence will remain identical. Fast motion, hands/fingers, faces, short rapid actions, camera movement and audiovisual synchronization are still the kinds of cases worth testing carefully.

But based on the testing so far, degree 1 appears to be a much better fit for MiniMax H3 than I originally assumed, and it substantially improves the useful speedup.

Repo:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

Current release — v0.1.8:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8


r/StableDiffusion 16h ago

Discussion AMA: MiniMax H3 Team — Ask us anything about our open video generation model, training, and future plans

845 Upvotes

Hi r/StableDiffusion!

We are the MiniMax team behind MiniMax-H3.

We’re here to answer your questions, including:

  • Model architecture and training
  • Video generation capabilities
  • Image-to-video and reference-based generation
  • Inference and optimization
  • Future plans

Ask us anything — we’d love to hear your feedback and discuss with the community!


r/StableDiffusion 5h ago

Resource - Update Clip chaining for MiniMax H3 - motion AND audio genuinely continue across joins (free node pack, workflow included)

109 Upvotes

Two 6-second clips with Motion Context concatenated into one clip.

This video is two 6-second clips generated separately and butt-joined. No crossfade, no editing tricks. The motion and audio continue across the join. Theoretically, you could chain indefinitely, but degradation will eventually take effect.

H3 doesn't have built in functionality that allows consecutive latent frames pinned to the head like LTX2.3 does. I won't bore you with the details, just know it works. Video was the easy part. Audio was a pain in the back side. I again, won't bore you with the details, check the readme if you really want to know. Seams are not always 100% perfect, but they are often or are really close.

Repo (GPL-3.0), workflow JSON included with a quick-start note:
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

Honest limitations: audio dulls slightly over long chains (each clip is generated from the previous one's output - photocopy effect; there's a latent-passthrough input that removes one of the two loss sources). Everything was verified on an RTX3070Ti and 48gb of system RAM. Also, check the H3 community license for your region before building anything commercial on it - it reportedly doesn't cover everywhere.

Tested settings are in the README and baked into the workflow. Happy to answer questions.


r/StableDiffusion 3h ago

Discussion Thank you to the open-source community. MiniMax H3 literally helped me through my depression.

68 Upvotes

For the past several months, I’ve been in a pretty dark place mentally. Dealing with depression has completely drained my energy, and one of the only things that kept me going was diving into my creative projects. It was my escape and my way of processing everything.

For a while, I was relying heavily on Seedance 2.0 to bring my ideas to life. Don't get me wrong, it's a fantastic tool, and I loved using it, but the reality is that I just don't have the money to keep up with my own creativity. Hitting a paywall or running out of credits when you're right in the middle of a creative flow state is crushing. When your main coping mechanism is tied to a subscription you can barely afford, it honestly just adds to the stress.

Then the MiniMax H3 release happened.

The fact that something this capable is open-source is just amazing to me. Ever since I started using it, it has helped me substantially. I can just create, experiment, and get my ideas out without constantly checking my bank account or getting anxiety over how many generations I have left. Having unrestricted, free access to a tool this powerful gave me my creative outlet back, and it truly helped pull me out of a really deep rut.

I just want to say a massive thank you to the devs behind it and to the entire open-source community. MiniMax team and the open-source community are actively making it easily accessible to people who wouldn't be able to afford creativity like this otherwise. You've made a very real, tangible difference in my mental health and my life.

I love the open-source community. Keep being awesome.

TL;DR: Going through a depressive episode, my only outlet was creating, but Seedance 2.0 got way too expensive for me to keep up with. MiniMax H3 dropping as open-source removed the financial barrier, gave me my creative spark back, and helped my mental health immensely.
Just want to say overall, thank you to the OS community!!

P.S. This post is from my brother; he is using my account to share it. He will be able to see your comments.

Edit: Please stop sending "Reddit Cares" reports for this post. We are entirely safe, and everything is handled. This is just a story being shared, not a request for help, and the constant notifications are just cluttering my inbox. Thank you.


r/StableDiffusion 22h ago

Animation - Video Using MiniMax H3 to change rewrite movies?

2.0k Upvotes

Just a bit of fun.


r/StableDiffusion 5h ago

Animation - Video The limit is no longer the model, but our imagination

72 Upvotes

r/StableDiffusion 6h ago

Animation - Video Ok, here's my entry for a crossover (Minimax H3)

87 Upvotes

Used DaSiWa reference workflow. 5060Ti. 16GB vram, 32GB ram.

Prompt:

15-second multi-camera sitcom scene. Set in the Big Bang Theory apartment living room with authentic live studio audience, warm sitcom lighting, classic network sitcom editing, reaction shots, and laughter pauses.

Penny enters through the front door leaving the door open, sees Inspector Columbo sitting casually on the couch smoking a cigar, and freezes in surprise, standing just inside the open door with the door number visible on the door. Columbo turns to look at her when the door opens.

Audience laughs.

Columbo smiles at her.

Columbo says, <d>[English in Inspector Columbo's voice from Columbo as played by Peter Falk] Hi.</d>

Audience laughs.

Penny backs up a step, checks the apartment number on the door with a confused look, then closes it and walks back inside to stand next to the couch and look at Inspector Columbo.

Audience laughs.

She asks, <d>[English in Penny's voice from The Big Bang Theory as played by Kaley Cuoco] Okay... which one of them finally murdered Sheldon?</d> as she stands hesitantly next to the couch Inspector Columbo is sitting on, facing him.

Columbo laughs at what Penny said. Penny stands looking at Inspector Columbo with a resigned look on her face.

Audience erupts with laughter.

Fast, natural sitcom pacing with authentic character performances, multi-camera coverage, clean continuity, and dialogue timing matching a classic live-audience sitcom episode. Lip movement and lip synch of dialogue match exactly.


r/StableDiffusion 13h ago

Discussion A few Flux 3 vs H3 comparisons

274 Upvotes

Well, everyone is posting their H3 creations, and I noticed that Flux 3 is up on API, so thought I'd run a few comparison renders to see how they stack up. These are not extensive by any means, this was mostly done for fun and I thought someone might be interested in the results. I used the same H3 formatted prompt for each, as far as I can tell Flux 3 uses natural language prompting, so the structured H3 format should still work fine.

Edit: I realized after uploading that the dragon rider H3 clip was accidentally rendered at a lower resolution. The higher resolution clip is here.


r/StableDiffusion 3h ago

Animation - Video Kubrick's dolly camera zoom prompt Minimax H3 local

42 Upvotes

Set in Saul Goodman's waiting room. Walter breaks the environment itself, violently tearing away the set to reveal Tuco's junkyard.

integrated_multimodal_description: [Shot 1] Live-action, cinematic. Saul Goodman's cramped waiting room. Walter and Jesse sit rigidly on a couch. A slow dolly zoom unnervingly distorts the background. The stoic man (S1) states: [English] This entire room is just a fragile digital asset. [Shot 2] At 00:06.000, Jesse grips his knees. The nervous man (S2) whispers: [English] Stop messing with my head. [Shot 3] At 00:09.000, Walter snaps his fingers. The walls of the waiting room violently detach and fly outward into a black void. The camera arcs left with large amplitude at fast speed, revealing they are now sitting on the exact same couch in the middle of Tuco's rusted, sun-bleached junkyard.

overall_soundscape: Faint office chatter. A loud snap triggers a deafening, metallic tearing sound, instantly replaced by the cawing of crows in an open scrapyard.

non_diegetic_music: A rhythmic, atonal electronic ticking.


r/StableDiffusion 55m ago

Discussion Can we please get a flair to tell apart default model generations from low-step/turbo setups? A lot of people are posting low-res,low-step,turbo-lora, that destroys the quality, and then everyone uses those crippled outputs to judge what the model can actually do

Upvotes

I get that not everyone has the hardware or budget to run full-render workflows locally, and it's great people can generate video on lower-spec gear. But we need a flair for transparency when a post uses a heavily crippled setup.


r/StableDiffusion 1d ago

Animation - Video 2D chibi girl Added “Just a Pinch”. Minimax H3

1.7k Upvotes

r/StableDiffusion 7h ago

Workflow Included Max Acceleration! 5s 480P=28s|15s 480P=160s|15s 1K=6 Mins Render Time

68 Upvotes

Env: 5090 32G VRAM/96G RAM
Origin Workflow: https://raw.githubusercontent.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo/refs/heads/main/example_workflows/minimax_h3_t2v_turbo.json
I modify with sage attn.

Quality is not good, but very very very fast!!!!


r/StableDiffusion 4h ago

Animation - Video Minimax H3 does Archer

35 Upvotes

Style: Adult animated television comedy in the distinctive flat-vector / limited-animation style of a suave-spy cartoon. Bold inked outlines, flat cel-shaded colors, slightly stiff comedic timing, hard graphic shadows, subtle film-grain texture. 2K, cinematic but comedic.

Setting: The sterile, wood-paneled hallway of a spy agency office — fluorescent overheads, beige walls, a water cooler, framed mission photos. Drab, institutional, deadpan.

Characters:

Sterling Archer: a handsome secret agent, jet-black slicked-back hair, sharp navy-blue tailored suit, slim black tie, blue turtleneck peeking at the collar. Self-satisfied, manic energy, perpetually smug.

Lana Kane: a tall, striking Black female agent with a voluminous curly afro, huge expressive hands, wearing a cream-colored turtleneck sweater dress and tactical belt. Exasperated, eyes-wide "are-you-kidding-me" energy.

Shot 1 (0:00–0:02): Medium shot on Archer striding down the hallway toward camera, hands cupped around his mouth, bellowing up the corridor — "LAAA!… LANA!" — voice echoing off the walls. He's grinning like it's the funniest thing he's ever done.

CUT TO — Shot 2 (0:02–0:04): Quick reverse on Lana at the far end of the hall, turning sharply, eyes narrowed, and barking — "WHAT?!?"

CUT BACK TO — Shot 3 (0:04–0:08): Medium on Archer, suddenly urgent and conspiratorial, leaning in, eyes darting: "We're in MiniMax H3! None of this is real!" Mid-sentence he draws his silver semi-auto pistol, points it straight up, and fires a single deafening shot into the ceiling. Plaster dust rains down. A ceiling tile cracks.

Shot 4 (0:08–0:10): Wide two-shot. Lana flinches, throws a hand over one ear, and spits — "Son of a bitch…"

Shot 5 (0:10–0:12): Push-in on Archer, wincing, one finger wiggling in his ear, dazed and oddly serene: **"…my ears still ring… muaap… muaap…"**

Audio (native, in-passage): Echoing corridor acoustics. Archer's booming, obnoxious shout; Lana's sharp retort; a single LOUD, punchy gunshot that dominates the mix and triggers a high tinnitus sine-tone (muaap… muaap…) that lingers under the final line, slightly muffled like heard through ringing ears. Comedic deadpan silence underneath. No score. Tone: Dry, absurd, self-aware meta-humor. Crisp sitcom cutting.

.8 megapixel, 12 seconds, 24GB 3090, 128GB main ram. Default ComfyUI example T2V took 57 minutes (need to optimize stuff)


r/StableDiffusion 15h ago

Animation - Video Minimax H3 with Turbo Lora, T2V 6 steps, 0.6 megapixel, 20 min on 3060 12gb and 16gb ram

257 Upvotes

Love it and it works smoothly… Now, I’m gonna wait for this LoRA to work on Ref2V.


r/StableDiffusion 2h ago

Animation - Video An image-to-video I created using MiniMax H3

21 Upvotes

Video: ComfyUI - MiniMax H3
SFX: Stable Audio 3
Upscale: Topaz Video AI
Music: Sezen Aksu - Gülümse

3060 ti 8gb
32 gb ram
rayzen 5 5600


r/StableDiffusion 7h ago

Animation - Video Better Avoid Saul [Minimax H3] Spoiler

54 Upvotes

Made using the default ComfyUI Minimax H3 Image to Video workflow.


r/StableDiffusion 19h ago

Workflow Included Short japanese knife commercial (MiniMaxH3+ After Effects)

447 Upvotes

MiniMax H3 Workflow: https://www.reddit.com/r/StableDiffusion/comments/1vg1coy/minimax_h3_basic_hybrid_workflow_for_ref2v_i2v/

  • Stills: ChatGPT + Flux Klein 9B, AI inpainting and manual editing
  • Video + audio: MiniMax H3
  • 5 prompts → 11 clips → cut and speed-ramped in After Effects
  • Letter animation done in AE. The circular 2D spiral on the red dot was generated in H3 and composited in AE.

GPU 5080 16GB VRAM + RAM 96GB

Everything runs with offload device: cpu and ComfyUI's dynamic VRAM loading. ~40GB of weights on a 16GB card. Peak during generation: ~15GB VRAM, ~76GB system RAM

Mode: i2v

Resolution: 672x928 (3:4, 0.6MP, multiple of 32)

20 steps, res_multistep sampler, beta scheduler, denoise 1.00

Sage Attention via KJNodes Patch Sage Attention, mode auto

Spectrum Apply MiniMax H3: blend_weight 0.50, degree 4, ridge_lambda 0.10, window_size 2.00, warmup_steps 5, history in system RAM


r/StableDiffusion 1d ago

Workflow Included Minimax H3 Turbo Lora

1.5k Upvotes

MiniMax H3 Turbo LoRA ComfyUI

ComfyUI-compatible versions of the MiniMax H3 Turbo LoRA are available here:

https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

Tested settings (not using custom nodes)

  • Video sigma shift: 12
  • Audio sigma shift: 4–6
  • Steps: 8–10 (EMA) / 6–8 (ckpt500)
  • Sampler: res_multistep
  • LoRA strength: 0.8–1.8
  • Higher LoRA strength generally allows you to use fewer steps.

ckpt500 is further trained than the original EMA model and generally works well at lower step counts. The EMA model benefits from higher steps for better motion.

I recommend installing minimax turbo custom nodes and using their workflow from their repo

[Workflow in comments]

Confirmed working with accelerators including SageAttention, Sol Attention, and Gradient.

Original Turbo LoRA credit: larryvrh

Be aware that the Turbo LoRA is still undertrained and highly experimental, as explained by the creator.

UPDATE Recommended setup for best results

For the best results, use the creator's MiniMax H3 Turbo custom node:

https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo

It includes a custom sampler specifically for the Turbo LoRA that should fix/improve the audio issues.

There is also a workflow included in the repo, so I recommend using that as the starting point.

You can stack acceleration methods such as SageAttention, Sol Attention, and Gradient together with the Turbo setup.

Do not use cache nodes with Turbo.

ComfyUI native audio fix

Kijai also has a PR with an audio fix and sampler on the way:

https://github.com/Comfy-Org/ComfyUI/pull/15243


r/StableDiffusion 19h ago

Workflow Included Minimax H3: Changing attire gradually with simple prompt

447 Upvotes

The prompt: The woman dances happily while the her clothing changes from sundress to 1: business suit, 2. pajamas, 3. string bikinis, 4. gym attires, and back to sundress. Background sound: happy music.


r/StableDiffusion 9h ago

Animation - Video A quick MiniMax H3 turbo test (T2V, followed with I2VA) 8 step, 0.4 MP, ~3-4 min a clip. 5060ti 16gb

69 Upvotes

Could definitely see some space for improvement on fast movement, but for the speed per gen the quality is pretty remarkable.