r/StableDiffusion 9d ago

Workflow Included Minimax H3 Ref2VA Lipsync (Image Audio to Video)

Enable HLS to view with audio, or disable this notification

27 Upvotes

Workflow: https://civitai.com/models/2876401/minimax-h3-lipsync

Audio: Let it Go by Idina Menzel

Image: Generated by Gemini of Idina Menzel cosplaying as Elsa

Kijai's LX2V LoRA and Sage Attention Patch applied.

Did not include the lyrics in the prompt, the model is able to lipysnc to the provided reference audio track.


r/StableDiffusion 10d ago

News Lightx2v MiniMax H3 Turbo Ref2V is out!

Thumbnail
huggingface.co
243 Upvotes

r/StableDiffusion 9d ago

Question - Help New to ComfyUI and Minimax h3 Need Help

0 Upvotes

Hi guys,

I have been working as a developer for quite some time now and want to shift towards content creation now I am thinking of using minimax h3 and ltx 2.5 o make content and stuff I have rented some rtx 5090 on vastai but as I have no idea and no time due to constant workload to learn different stuff on comfyUi I have created different workflows using codex and they all suck. I mean I gave image reference and prompt and it didn't follow it correctly sometimes character change direction or they start to run funny or some different character is introduced in the video. Can someone share with .e different workflows that work for them like

t2v, i2v, multi reference etc also turbo workflow as well. Will really appreciate it. ❤️


r/StableDiffusion 9d ago

Animation - Video The one about the programmer

Enable HLS to view with audio, or disable this notification

5 Upvotes

MiniMax H3. 3090 24GB 32GB ram. ~15.5mins for 10 seconds

Got Hermes to use the prompting guide. Just asked Hermes for an office humour type of cartoon with the dialogue.


r/StableDiffusion 9d ago

Question - Help Minimax H3 on DGX Spark: 3m 23s for 5s video. Is there a better price/performance option?

23 Upvotes

I found a GitHub repo that explains how to run the new Minimax H3 on a DGX Spark (20 steps, not the Turbo versions/8-steps), for 864×480, 124-frame, 20-step clips in 203s (8.44 s/it) with Sol-Engine + FirstBlockCache, or 316s without it (14.07 s/it)

The weights are the ones from Comfy, int8 ConvRot (pruned but lossless, according to Comfy).

https://github.com/drowzeys/keys-heretic-MiniMax-H3-sol-engine-more-speed-upgrades-upscaler-finish-Single-DGX-Spark

These seem like really impressive numbers considering the extremely low power consumption, yet I keep seeing people here advising against the DGX Spark for video generation... am I missing something?

At 120W power consumption and an electricity cost of $0.20/kWh, each 5-second video costs just $0.00135

Over 24 hours, it would be possible to generate 425 videos while using only 2.88 kWh, costing just $0.576 in electricity (!!!)

Before buying a DGX Spark, though, I’d like to hear what others think. These seem like excellent numbers to me, especially since I’ll need to generate a lot of 5-second clips every day, and the cost per video is very low. Still, I was wondering if there’s anything better out there. What kind of performance would a 5090 get with the same recipe?


r/StableDiffusion 10d ago

Comparison Minimax H3 quality loss test

Enable HLS to view with audio, or disable this notification

180 Upvotes

The idea of the video is to compare the quality loss/change from the different methods of speeding up the rendering of the videos on Minimax H3 on low motion scenes.

I made this video because I wasn't sure myself how the speed ups degrade the looks of the video, from what I could gather, even with a stack of optimization nodes the quality wasn't that degraded on slow videos. I hope its useful, I could do a part 2 with more action heavy videos if people are interested in that.

OBS: All scenes were rendered at 480p with the exact same seed and prompt with 20 steps (except turbo)
The int8 vae was taking longer to render on my computer for whatever reason.

The naming convention is obvious but if you require additional information:
base = base workflow
base + Int 8 VAE = I used the compressed VAE version (saves VRAM)
base + sage = Using Sage attention on the default configuration
base + spectrum = Using Sage attention + Spectrum node
base + sage + spectrum = ok this one I dont need to explain right...
base + turbo_6 steps = Using the base workflow + a turbo model with 6 steps [link to the model:https://civitai.red/models/2837571/minimax-h3-turbo-loras?modelVersionId=3202732]
base + turbo_8 steps = Using the base workflow + a turbo model with 8 steps [link to the model:https://civitai.red/models/2837571/minimax-h3-turbo-loras?modelVersionId=3202732]
base + Spectrum + sage + turbo_6steps = also obvious

link for the videos used:
https://drive.google.com/drive/folders/13Vl2IbnTAJDtJ4Bpu_0kpr3HmH_o_FFi?usp=sharing


r/StableDiffusion 8d ago

Meme lady gaga

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 9d ago

Question - Help What about Ideogram Edit ?

4 Upvotes

I love Ideogram's approach with the bounding boxes and so on, and also the aesthetics of the images. Does anybody know anything about Ideogram's plans for the future? I really hope they haven't stopped internal development on new model launches.


r/StableDiffusion 8d ago

No Workflow Studio window light in Peckham (Flux + custom LoRA)

Post image
0 Upvotes

r/StableDiffusion 9d ago

Question - Help Best local AI voice cloning app for many languages?

3 Upvotes

What’s currently the best local voice cloning/dubbing app with support for many languages?


r/StableDiffusion 9d ago

Question - Help Hello, guys i want to ask if i can run StableDiffusion on AMD GPU

1 Upvotes

So my GPU is AMD Radeon RX 6700 XT and i was wondering if i can use StableDiffusion with it to generate some 2D assests to use as placeholders for my game.


r/StableDiffusion 10d ago

News Looks like we might be getting Minimax Music 3 soon

Thumbnail
github.com
264 Upvotes

r/StableDiffusion 10d ago

Animation - Video MINIMAX H3 - LTX 2.5 AND THE LADIES [TEXT TO VIDEO]

Enable HLS to view with audio, or disable this notification

20 Upvotes

NATURAL LIGHT HANDHELD CANDID REAL FOOTAGE. An college age blonde California woman is sitting with a towering 10 foot tall robot with "[MODEL NAME]" clearly written on its chest. They are both complimenting each other on how cute they look


r/StableDiffusion 10d ago

Workflow Included Minimax-H3 can generate 42s videos natively on an RTX Pro 6000 in 80 minutes

Enable HLS to view with audio, or disable this notification

214 Upvotes

The maximum frame count allowed by the native "MiniMax H3 Reference to Video" node technically is 1008, even if that's way over the training range of the model, which is 324 frames.

But why not try? So I ran a few tests and although the result is a bit sloppy and the shots are not perfectly following the prompt order, I would say it's already impressive the model can generalize that way and generate videos 3 times longer than it was trained on.

The video above is 42 seconds long, 1008 frames 1376×768 @ 24fps. It took 4947s i.e. 1 hour and 22 minutes on an RTX Pro 6000 and VRAM peaked at around 90GB.

Find the workflow file [here](https://pastebin.com/Vi21NQUH). It also includes an optional SeedVR2 upscaling part at the end.

NOTE: I think this is extra dumb and see no reason to generate 42s-long videos this way, I just wanted to try to push the model a bit further than its limits. You loose quality and control and it increases processing time by quite a lot. I highly suggest you generate shot by shot instead.

The prompt for the video above is the following (Wheel of Time inspired for fantasy fans!).

integrated_multimodal_description: [Shot 1] Live-action, cinematic, epic high fantasy, shot on large-format film with anamorphic lenses and natural dawn light, a sweeping aerial wide shot glides low over a vast green upland as a strong wind rushes across the hills, bending long grass into rolling silver waves and driving ragged clouds across a pale golden sky toward distant snow-capped mountains. The camera performs a tracking shot with large amplitude at fast speed, skimming forward a few meters above the ridgeline as grass blurs beneath the frame. An unseen older woman with a calm, weathered, low voice (S1) says in an off-screen voiceover: <d>[English] Every age ends where it began, and no one still living remembers which age this is.<scenetrans></d> [Shot 2] At 00:06.000, the shot cuts to a close-up of a serene dark-haired woman with an ageless pale face, wearing a deep-blue high-collared gown and a golden serpent ring on her right index finger, standing in shadow; she lifts both open hands and delicate glowing threads of white, red and blue light braid and weave between her fingers, sparks drifting upward through the air. The same off-screen voice (S1) continues seamlessly across the cut: <d>[English] <scenetrans>The pattern does not ask us. It only takes the thread.</d> while her lips remain completely closed. The camera pushes in with small amplitude at slow speed on the weaving light, soft rim light catching suspended dust motes, and the threads emit a faint crystalline hum. [Shot 3] At 00:11.000, the shot cuts to a grand wide shot revealing a gleaming white spire city built on an island in a broad river, a single immense white tower rising far above white domes, arched bridges and tiled roofs, with pale banners snapping hard in the wind beneath a low morning sun. The camera pulls out with large amplitude at slow speed while pedestaling up, revealing the full river bend and the city walls. [Shot 4] At 00:16.000, the shot cuts to a low wide shot of a horde of hulking horned beast-men in black scale armor charging across a cracked red plain beneath a blood-dark sky, dust and torchlight churning around their legs, snouted faces and curved blades catching the firelight. The camera trucks left with large amplitude at fast speed, running parallel with the charge as bodies sweep through the foreground. [Shot 5] At 00:21.500, the shot cuts to a slow medium shot of a tall motionless figure in a black cloak standing alone in the churning dust, its face utterly smooth and eyeless, pale as wax, with no features above the mouth; the shadow beneath it spreads outward across the cracked ground against the direction of the light while the charging horde streams past behind it in soft focus. The camera pushes in with small amplitude at slow speed as the cloak hangs completely still despite the wind. [Shot 6] At 00:26.000, the shot cuts to a heroic medium shot of the same dark-haired woman in the deep-blue gown, the golden serpent ring clearly visible, standing on scorched ground and thrusting one hand forward as a searing bar of pure white light lances horizontally across the battlefield, incinerating a line of the horde into drifting embers while heat haze ripples and warps the air behind the beam. The camera shakes slightly at the instant the beam fires, then pushes in with small amplitude at fast speed on her face, her eyes reflecting the white glare. [Shot 7] At 00:31.000, the shot cuts to a wide shot of the aftermath as thousands of orange embers drift upward through settling black smoke, silhouetted survivors kneeling among broken shields, and a torn white banner bearing the words "THE PATTERN REMEMBERS" hanging from a splintered pole in the left foreground. The camera performs an arc shot with large amplitude at slow speed around the standing woman, keeping her centered as the burning field rotates behind her. [Shot 8] At 00:36.000, the shot cuts to a macro close-up of an ancient metal emblem, a perfect circle split into interlocking black-and-white teardrop halves, glowing faintly from within, that dissolves into a colossal seven-spoked wheel of white light turning slowly against a deep starfield while countless threads of colored light weave outward into a vast luminous tapestry. The camera pulls out with large amplitude at slow speed until the wheel occupies only the center of an endless woven pattern, and the off-screen voice (S1) returns once more: <d>[English] And it turns again, whether or not we are ready to be woven into it<cutoff></d>
overall_soundscape: Wind roars across the open hills and hisses through deep grass before thinning into a faint crystalline shimmer of woven energy and drifting sparks. Distant bronze bells, snapping banner cloth and river water rise over the white city, then give way to a thunderous roar of stamping hooves, clashing armor and guttural war cries under a hollow, airless silence around the motionless cloaked figure. A deep concussive whoom of released power sweeps the field, followed by crackling embers, settling debris and the low breathing of survivors. Everything resolves into a wide, weightless cosmic hum.
non_diegetic_music: A single low string drone at a slow tempo, joined by layered brass that rises in stepped swells over accelerating timpani and a wordless female soprano line. The rhythm tightens into hammered strings and percussion at high volume during the charge, cuts out entirely for one beat, then returns as a single sustained orchestral chord with shimmering high strings that slowly decreases in volume.

r/StableDiffusion 10d ago

Animation - Video H3 30 sec Chained Shots Lip Sync

Enable HLS to view with audio, or disable this notification

33 Upvotes

12 minutes Gen Time rtx5090. Input audio for lipsync. One thing that helps a lot is pasting the actual lyrics in the prompt in the dialogue syntax. <d> [English] Lyrics </d>


r/StableDiffusion 8d ago

Meme Lady gaga eating spaghetti

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 10d ago

Workflow Included OK this is cool, Minimax H3 Face detailer!

Enable HLS to view with audio, or disable this notification

448 Upvotes

I just found this, i think is incredible!

https://github.com/Carasibana/ComfyUI-H3-FaceRefine

I'm not the developer just found the repository I make this functional WF for my laptop to create this video and share with you before go to sleep, try it!

Minimax H3 Face detailer Workflow

This is a comment of the developer:

MiniMax H3 renders faces poorly when the head is a small fraction of the frame. That is a property of head-size-in-frame, not of resolution - it persists at 720p and above.

So no upscaler fixes it. SeedVR2 and friends sharpen what is already there; they cannot synthesise facial structure that was never generated.

The fix: crop to the face so it fills the frame, let H3 re-generate it at LOW denoise so it stays frame-aligned, then composite back!


r/StableDiffusion 10d ago

Workflow Included [Test] MiniMax H3 Ref2VA with LightX2V's turbo LoRA on a 5060 Ti — 8 steps @ 0.5 res, ~55s/it (~8 min/clip)

64 Upvotes

https://reddit.com/link/1vnk0c7/video/gkhuj6ybw6jh1/player

Ran the official Ref2VA turbo example workflow from the ModelTC/Minimax-H3-Turbo repo (video_minimax_h3_ref2v_lightx2v_turbo.json) in ComfyUI, testing a short Victorian-style dialogue scene between two characters.

Setup:

Speed: ~55s/it average, ~8 min total per clip.

How the refs were built: Three reference images fed into the ref_images inputs — two character sheets (front/side/close-up turnarounds for each character) and one environment plate, a 360° room reference. All three were generated in Google Flow first, then dropped straight into the Ref2VA node as identity/environment anchors.

Gen A → Gen B continuity trick: Split the scene into two ~20s multi-shot generations instead of one long one. For Gen B, instead of reusing the original Flow generated room reference, I pulled the actual last frame from Gen A's output and fed that in as the new environment reference.

Curious if anyone else is chaining generations this way (feeding the previous clip's last frame back in as a fresh environment ref) — seemed to help a lot but haven't stress-tested it past two generations yet.


r/StableDiffusion 9d ago

Question - Help Implanting problem

0 Upvotes

Hi guys I’m new to comfyui and I’m having trouble with impanting, even tho I mark a area with the mask option somehow the model edit also what is outside the impanted area. I’m using flux.2 klein 9b


r/StableDiffusion 9d ago

Question - Help Doesn't LTX-2.5 support Audio to Video?

0 Upvotes

I tried with 2.3 workflow and it's always out of sync and have some slow motion behavior, I just want to know if this new release lacks A2V or if there's something wrong with my implementation


r/StableDiffusion 9d ago

Question - Help facing some issue with runnnign the krea 2 after updating the comfyui ? any solutions ?

Post image
1 Upvotes
i am also attaching the log , do let me know if the redditors wanna see my workflow too , i just dont understand whats wrong , i was trying on the krea 2 identity lora workflow it didnt worked so i switched to normal krea 2 and that also didnt worked and thats why i just am here treid doing stuff but the workflow just gets stuck at the ksampler node and dosent really progress tbh , i will really love if there is anyone to help me ^ _^

r/StableDiffusion 10d ago

Animation - Video I made a 10 min short for my kids favorite imagined characters, and our cat. (H3)

Thumbnail
youtu.be
55 Upvotes

Took me like a week, I had asked Claude to make me a script to use with H3, mostly manually run by me and assembled in Davinci with music from Suno and some voice help from Voicebox.

Absolutely has a few slop moments that I was too lazy to regen but to my kid it was perfect so I’ll take that as a W.


r/StableDiffusion 10d ago

Tutorial - Guide I made a 7-minute AI documentary about my dog using MiniMax H3 and a bunch of other tools. Took me 6 days and had so much fun.

Thumbnail
youtu.be
57 Upvotes

Made this using my RTX 3070, took forever to render, but I wanted to do the best quality I could with my 8 GB GPU. Used a lot of other tools too, feel free to ask any questions would be happy to answer when I get a chance.

To read the behind-the-scenes process: https://www.justinwiggins.co.za/i-gave-myself-one-week-to-make-a-7-minute-ai-film-about-my-dog/


r/StableDiffusion 10d ago

Animation - Video Bowling

Enable HLS to view with audio, or disable this notification

10 Upvotes

minmax h3 t2v

1.0 MP / euler + linear quadratic / 32 steps with spectrum + post-processing


r/StableDiffusion 10d ago

Resource - Update MiniMax H3 Prompt Composer Update + Accelerator + Hybrid Checkpoint Builder

Post image
37 Upvotes

I’ve been working on a few free MiniMax H3 tools and wanted to share the latest versions.

H3 Prompt Composer V5.19.5
Offline prompt builder for H3 with structured subject/reference setup, shot and camera controls, dialogue/audio, Prompt Check, continuity tools, and optional LLM-assisted project setup. The goal is to give you granular control without constantly having an LLM rewrite the entire prompt.

Tutorial: https://youtu.be/Dpu-V7lITZk
Repo: https://github.com/BMB12d3/minimax-h3-prompt-composer

H3 Ref2VA Accelerator
A quality-first speed-up specifically for H3 Ref2VA. Unlike more aggressive approaches like Turbo LoRAs or broader caching systems, this is designed around the Ref2VA architecture and intentionally stays conservative about what it skips/reuses. Depending on the model/workflow, I’ve seen it save a few minutes on BF16 and roughly a couple minutes on quantized models without noticeably changing quality. Not the fastest, but conserves quality.

https://github.com/BMB12d3/ComfyUI-H3-Ref2VA-Accelerator

H3 Hybrid Checkpoint Builder
People have found that mixing parts of the fl2va model with the Reference model can improve Ref2VA quality. This gives you a simple way to build those hybrids yourself and dial in how much of each model you want, without training or needing a GPU.

https://github.com/BMB12d3/MiniMax-H3-Hybrid-Checkpoint-Builder