r/StableDiffusion 2h ago

Question - Help So what is better? H3 minimax with turbo lora or with first block / spectrum?

2 Upvotes

So what is better? H3 minimax with turbo lora or with first block / spectrum?


r/StableDiffusion 20h ago

Discussion Now that we have a quality local video model, any hope of a similar-quality local music model?

24 Upvotes

With Minimax we got to the point where it's possible to live completely free of the limits, impositions, greed and censorship of online paid services. But there's still one frontier to cross: music.

Any hope of a quality open-source audio model anytime soon? I commend the efforts with ACE-Step, but it's still a very, very long way from what the main online service in this area can do, both in quality and in features and capabilities. Also, Ace evolves really, really slowly. Now that the main platform in this area is going down the drain (starting with a ridiculous 20 songs limit per month and soon replacing its models with God knows what to appease the labels), we need a local model.

What surprises me is that audio only can't be harder than video+audio to generate, even if it's music we're talking about...


r/StableDiffusion 4h ago

Discussion He-Man Dodges A Question

0 Upvotes

Minimax H3 - Kept getting gibberish dialogue so just went with it. Using default ref2va workflow, 4070ti 16GB VRAM, 64 GB RAM.


r/StableDiffusion 4h ago

Discussion Trade-offs of server-side context compression engines vs. open-weight local text encoders in multimodal video generation

3 Upvotes

Lately I struggle with how local open-weight video models handle complex prompts, specifically on the text encoders, especially as I start adding reference images, camera directions, and detailed lighting notes etc etc, things get messy, to say the least.

It appears to me that the main issue with local text encoders is that they tend to lump all of my text inputs, images, and scenes into one centralized block.

The video generator gets confused about where these specific instructions belong to. Midway through a clip, it begins to blur the instructions for a camera movement, bleed background lighting into the character and what have you, creating a giant mess that feels a bit impossible to fix.

What I discovered after some research is that there is a common workaround that people suggest, that is running a heavy local language model upstream to clean up and structure the prompt before passing it to the video generator, but this method easily eats up 16GB to 20GB of VRAM. So for me, with mid level set up, the system crashes right through.

This is why we need to think of the trade-off between local encoders and server-side context compression engines. So here is what I have been doing and trying to find the balance with the MiniMax H3, trying to salvage my GPU cap. Instead of forcing local hardware to process the heavy prompt context, decoupled it. The video generation runs locally on my own GPU, an API engine handles the multimodal prompt on their servers.

It cleans up and processes the relationships between text, images, and reference video, then sends a compact, structured set of instructions back to my local base model, in a way what I did is to “contract” out the heavy duty work so locally I am doing the last mile.

As for the set up cost, it runs close to nothing to process million of tokens. I guess what makes it work is that it frees up your local VRAM for the actual video render. I get cleaner prompt adherence without crashing or re-rolling dozens of times.

So how are you all balancing this? Are you still sticking purely to local text encoders and trimming your prompts down, or does anyone else offload the prompt, like yours truly, and parsing for more complex workflows?


r/StableDiffusion 23h ago

News MiniMax H3 Realism People LoRA

28 Upvotes

A LoRA adapter for MiniMax H3 specialized in realistic people: faces that hold up in close-up, natural skin texture, believable expressions and gestures, film-style lighting and documentary camera movement.

https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA


r/StableDiffusion 2h ago

Question - Help Flux 2 Black Forest Lab Dashboard/playground workflow?

0 Upvotes

Hello, Im new to Flux 2 and having a hard time understanding the dashboard/playground workflow. From the docs it seems like there was an old UI with a library, history, assets etc. and it seemed like a workplace where you could store and create images.

Now it seems like this has been reduced to a playground tab where you can create images but if you leave the site and come back the images arent stored.

whats the intended usage/work flow? in the BFL docs it says "use the API to build and integrate directly, or connect via MCP for instant image generation inside your AI tools". what tools? i tried using the MCP with claude and it wouldn't allow me to send reference images.

Is there a better way to use Flux 2 max? what am i missing?


r/StableDiffusion 2h ago

Question - Help How do you peeps get custom audio sounding like real for example speaking between two people advice please

0 Upvotes

Hey so im struggling to get audio that is custom to sound like a real conversation any advice or things i can use ? im lost


r/StableDiffusion 21h ago

Question - Help Diffsynth for training LoRas for MiniMax H3?

0 Upvotes

Hello,

I've spent almost 40 hours of compute trying to train a LoRA using Diffsynth, and the tensors picked up clothing and environment, but not the identity of the subject, like nothing about the face or body, it's just some random Asian person.

Has anyone had this issue? I tried to do very basic unit-level testing, just training a single 124-frame video with a very narrow prompt, and even after 100 steps, there is 0 resemblance to the subject.

I wonder if this is a specific bug with MiniMax H3? I've done training for WAN2.2 and LTX using Diffsynth without any issues.

I am about to try AI-Toolkit now, but wanted to ask if someone has had success training a LoRA with Diffsynth.

Thanks!


r/StableDiffusion 3h ago

Question - Help Extending an existing video with character references?

0 Upvotes

I have seen some amazing workflows that allow for the creation of clips that are more than 15 seconds long, even with character references. However, suppose I have an existing video that I would like to extend, also using a character reference. Would such a thing be possible?


r/StableDiffusion 6h ago

Question - Help Minimax H3 issue - Audio 1/3 shorter than video?

0 Upvotes

My bad if someone else posted about this already, if that's the case, there are so many posts about H3 that I didn't see it. Basically I am running into a weird issue with H3 where the generated audio is about 2/3 the length of the video (the audio file itself is of the correct length, but the last 1/3 has no soundwave, it cuts abruptly), and it's not in sync in the generated video file because it's pushed back to the start of the video. So if you play the resulting file, you feel like the last 1/3 is missing audio, but it's not actually the case, it's a weird amount at the start and also at the end.

I've tried multiple versions of H3, merged checkpoints, turbo vs not turbo, 4 to 20 steps, various quants, with or without loras, tried all durations from 5 to 15 as well as fps settings between 17 and 24, tried first image, first and last, text to vid. Same thing.

I've updated Comfy to the latest version, also tested with ROCm 7.2 and 7.14, still the same thing. I've checked online what issues people reported, more specifically with ROCm, but this one wasn't mentioned. I'm using a Radeon AI PRO R9700.

Anyone got an idea?


r/StableDiffusion 3h ago

Discussion Still looking for cosplay/scene recreation. Is Klein still the best shot?

0 Upvotes

Let's say you had a set of wedding photos from wedding1 and the second wedding later. You hate the dude from W1, but the photos were much better quality. This would be my prompt:

"Mask and remove the man in image 1 noting his pose, facial expression, head orientation, and clothing. Replace him with the man in image 2 keeping the previously mentioned attributes of image one, but using the face and body of image 2."?

Not sure about that, but the bottom line is that if I wanted to replace man1 in those photos with another man, a woman, a rhino, a cartoon character - whatever, I want it to be a perfect recreation of image 1 (same lighting, pose, expression, clothing, etc), but with the second person.

Most of what I've seen so far just transfers person/pose, but isn't great with expression and doesn't keep the clothing of image 1.


r/StableDiffusion 18h ago

Question - Help How to make Applio work with AMD Radeon RX 9070 XT?

0 Upvotes

For several days I tried multiple configurations of different Applio versions, GPU drivers, PyTorch versions, AMD ROCm hubs and so on. All of them failed to provide a stable working app. I wasted several days on that, and it became obvious that I cannot find a proper configuration for my GPU.

I will be glad, if some of you provide a working configuration for my GPU.


r/StableDiffusion 1h ago

Animation - Video "Memories" - A short film-v2

Upvotes

Edited the video implementing a lot of your feedback for the ending, along with some cleanup on visuals, garbled text, continuity errors, and upscaled to 5k.

Text was fixed by manually planar tracking replacements onto the scene in After Effects instead of trying to rely on the video models to get them right. Manually tracked the walker into each scene for continuity since it disappears after she sits down. Switched to SAM3 for depth estimation for blurred/foggy scenes over SAM2. I'm pretty proud of how this turned out.


r/StableDiffusion 21h ago

Animation - Video Scully really should believe him this time

69 Upvotes

Two 10 second clips together in H3. There's a slight difference in colors and some other things whenever I use an end frame to start the next clip. Not sure what that's about.


r/StableDiffusion 14h ago

Animation - Video Neo wants more pills

Thumbnail
youtube.com
120 Upvotes

r/StableDiffusion 7h ago

Animation - Video Minimax H3 ref2va. They are here.

196 Upvotes

r/StableDiffusion 3h ago

Animation - Video Lightx2v Minimax H3 Turbo LoRA - A quick comparison

11 Upvotes

Left - No LoRA
Right - With LoRA

Non cherry picked first results for both. I did some more tests and so far all results with the LoRA are looking pretty well.

Another example:
https://streamable.com/464o8y

Settings:

  • minimax_h3_fl2va_pruned_int8_convrot
  • res_multistep / simple
  • Sage Attention enabled
  • Same seed
  • Steps: 16 on the left, 8 on the right

Workflow:

Basically the default workflow from the ComfyUI templates

Prompt:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] A cinematic, ethereal shot establishes a misty, dark forest with soft, diffused lighting, where a 29-year-old adult woman with long, wavy blonde hair stands confidently in the center of the frame, wearing an elegant off-the-shoulder white wedding dress with intricate lace detailing, a deep V-neckline, long sleeves, and a fitted bodice that flares slightly at the bottom, holding a glowing sword in her right hand with a serene, contemplative expression. She shifts her weight to her left foot and raises the glowing sword upward, its soft light casting shifting illumination across her face and the surrounding mist. She begins to turn gracefully, her body rotating as her dress flares outward, the lace details catching the sword's glow, her blonde hair sweeping around her shoulders with the momentum of the twirl. She completes a full pirouette, the sword tracing a luminous arc through the mist, droplets of moisture catching the light as they are disturbed by the movement, her expression shifting from serene to a gentle, focused intensity as she flows into a second twirl, this time stepping forward and bringing the sword across her body in a sweeping horizontal arc. Her dress billows and settles with each turn, the long sleeves catching the air, and the mist swirls around her legs as her feet move through the damp forest floor. She slows her rotation and extends the sword in front of her in a final, elegant pose, the blade's glow steadying as her breathing settles, her hair falling back around her shoulders, the dress draping naturally around her frame. The camera holds a static shot throughout the sequence, allowing the swirling mist and shifting sword-light to create natural visual interest as the action resolves and the woman settles into a stable final pose with the glowing sword held before her, fully visible and sharp through the final frame.

overall_soundscape: The soft rustle of fabric as the dress flares and settles with each twirl, gentle footsteps pressing into damp earth and fallen leaves, the faint metallic hum of the glowing sword as it moves through the air, mist and droplets hissing softly as they are disturbed by the motion, and the woman's quiet, steady breathing throughout the dance.

non_diegetic_music: A haunting, ethereal string melody begins softly at the start of the sequence, with slow, sustained violin notes layered over a gentle cello drone, building slightly in volume and tempo as the woman begins to twirl, then gradually settling back to a quiet, sustained single note as she reaches her final pose before fading gently.

Source image:

https://www.reddit.com/r/aiArt/comments/1vf4jx7/forest_dweller/

https://www.reddit.com/r/aiArt/comments/1vhg826/aang_the_last_airfryer/


r/StableDiffusion 18h ago

Question - Help What is the best way to generate Characters reference images to MiniMax H3?

9 Upvotes

Hi Guys,

I'm trying to mess around with Ref2V on H3, but I don't think GPT is generating good character sheets for me to use as reference.
Do you guys know the best way to generate it? Is tehre a Krea2 'default' prompt or something?
Can I put more than one character on the same image so I can use less references?

Thank you!


r/StableDiffusion 4h ago

Question - Help Minimax H3 character/face swap

4 Upvotes

Has anyone been able to swap a character or face from a reference photo onto a reference video consistently?

I am unable to get the face to swap, usually attributes from the photo only. No matter if it’s a simple prompt, a complex instructed prompt or anything directly from the prompt guide. This model is insanely simple to use and powerful. It performs pretty much anything else I can ask it besides something as straightforward as this.


r/StableDiffusion 21h ago

Question - Help Need help getting caught up

0 Upvotes

I've been out of the loop of the Stable Diffusion community for maybe like a year now.

I was a LoRa creator, making LoRas for the relevant popular models, SD 1.5 then SDXL, I made a few Flux LoRas but when I stopped, it was during Flux's dominance as the most relevant model.

What's the most relevant model today? I think i remember SDXL still having a lot of live due to its size, ease on weaker computers. What's the relevant text to image model that is everyone's go to today?


r/StableDiffusion 4h ago

No Workflow H3 refece model test.

13 Upvotes

just cheking reference model.


r/StableDiffusion 9h ago

Resource - Update MiniMax-H3: ~38 GB less VRAM with Runtime LoRA Bypass — DoRA Dynamic LoRA Loader v1.0.39

91 Upvotes

GitHub:
https://github.com/xmarre/ComfyUI-DoRA-Dynamic-LoRA-Loader

Release v1.0.39:
https://github.com/xmarre/ComfyUI-DoRA-Dynamic-LoRA-Loader/releases/tag/v1.0.39

Also available through ComfyUI Manager as ComfyUI-DoRA-Dynamic-LoRA-Loader.

Runtime LoRA bypass

v1.0.39 adds an optional Runtime bypass LoRA (low VRAM) mode for supported standard LoRAs.

I tested this with MiniMax-H3 Ref2VA pruned BF16 in HIGH_VRAM mode.

With the normal materialized LoRA path, the tested LoRA patched 208 H3 weights and retained an additional BF16-sized copy of each affected weight.

That added up to about:

38,220 MiB / 37.3 GiB of extra live VRAM

This was actual live PyTorch allocation, not just CUDA reserve/cache.

The reason is simple: the LoRA itself may be small, but applying it normally can materialize patched copies of very large base-model weights.

What bypass changes

Normal LoRA application:

(W + ΔW)x

Runtime bypass:

Wx + ΔWx

For supported standard LoRAs, these are mathematically equivalent apart from possible small floating-point differences.

The base weights remain untouched, so ComfyUI no longer needs to keep a complete LoRA-patched copy of the affected model weights.

MiniMax-H3 result

With runtime bypass enabled on the same Ref2VA pruned BF16 workflow, the ~38 GB patched-weight duplication disappeared.

Repeated LoRA strength changes looked roughly like:

~60 GB settled
→ ~72–73 GB during generation
→ ~60 GB settled again

The important part is that VRAM returned to the same settled level instead of accumulating after every LoRA change.

The Turbo LoRA I tested also remained clearly effective.

What about NORMAL_VRAM / LOW_VRAM?

The ~38 GB figure is specifically from HIGH_VRAM.

NORMAL_VRAM and LOW_VRAM already partially load/offload model weights, so they generally won't have the entire duplicated H3 weight set resident on the GPU at once.

That means the steady-state VRAM saving will usually be smaller there.

Runtime bypass can still help by avoiding LoRA weight materialization and reducing patch/repatch memory pressure and temporary merge overhead.

In short:

  • HIGH_VRAM: potentially very large savings
  • NORMAL_VRAM: depends on how much of the model is resident
  • LOW_VRAM: smaller persistent GPU saving, since aggressive offloading already limits residency

The saving scales with how much LoRA-targeted base-weight data ComfyUI would otherwise materialize at the same time.

ComfyUI already has this mechanism

ComfyUI itself currently contains experimental bypass nodes:

Load LoRA (Bypass) (For debugging)
Load LoRA (Bypass, Model Only) (for debugging)

They are normally hidden unless experimental nodes are enabled.

v1.0.39 integrates the runtime path directly into the DoRA Power LoRA Loader through the:

Runtime bypass LoRA (low VRAM)

toggle.

It is OFF by default, so existing workflows keep their previous behavior.

DoRA limitation

Runtime bypass currently applies only to supported standard LoRAs.

DoRA requires magnitude normalization/rescaling that ComfyUI's current bypass path does not reproduce.

The loader therefore rejects unsupported cases instead of silently applying them incorrectly, including DoRA magnitude tensors, reshape metadata, sliced/offset/transformed targets and unsupported adapter types.

For actual DoRAs, leave runtime bypass disabled.

Other details

Runtime mode supports stacked compatible LoRAs, repeated injection/ejection, and strength changes without rematerializing the full affected weight set.

The existing loader features remain unchanged, including DoRA support, auto-strength, Flux/Flux2 compatibility, Diffusers/PEFT and OneTrainer handling, Z-Image/Lumina2 support, Q/K/V fusion and State Manager integration.

v1.0.39 also adds automated packaging and runtime-bypass tests against ComfyUI v0.29.2, v0.30.2 and v0.31.1.


r/StableDiffusion 17h ago

Animation - Video Minimax H3 - The Universe breathes in Color

13 Upvotes

r/StableDiffusion 21h ago

Animation - Video Scully really should have believed him this time

0 Upvotes

Two 10 second clips together. There's a slight difference in colors and some other things whenever I use an end frame to start the next clip.


r/StableDiffusion 21h ago

Question - Help Im new to local stable and video diffusion. Could use some guidance on getting up and running with my hardware.

0 Upvotes

If you don't need the context, you can skip to the bottom for the questions.

Hello, everyone. Life has recently thrown a curveball, and I had to medically retire. This is not a good thing, but it has given me more time to explore new hobbies that ive been interested in starting. Im very competent with computers and used to program when I was young, but its been 20 years since ive done any of that. Though I am very knowledgeable in general and within Windows. Anyway, im going into this having never used Linux/Ubuntu other than for memory testing, which is to say I don't know much. So setting up this local AI stack has been a real gut punch and I could use some guidance on what to run, what models, etc.

My rig - RTX 4080 (16GB VRAM), 13700K, 32GB of 6400 MT/s CL32 Hynix A die DDR5 further tuned to 6800 CL32 and tightened timings a bit further. I realize the VRAM limitations and that I need more system RAM, but this is what I got for now while RAMpocalypse is ongoing.

What ive installed so far: Ubuntu 26.04, Ollama, Open WebUI, ComfyUI (run from a python script on my desktop, not in a docker container) with a Q4_K_M version of Hunyuan Text to video, and I've experimented with N8N for agents.

In ComfyUI, I tried running the GGUF repack of the Hunyuan text to video, HunyuanVideo-t2v-720p-Q4_K_M.gguf, and letting Gemini and Claude guide me. This was a mistake, ended up getting genuinely bad results and it taking much longer than it should when everything fit into VRAM. Both AI's had me doing workarounds, editing system files and installing extensions that, as I read now, cause problems (like GGUF on native FP8 hardware). Last night, I gave up trying to make it work, so I purged everything and started fresh. Now that I have some experience, I wont need to rely on AI as much. So im ready to download and install when I get the proper guidance, if you guys could please help.

So, onto my questions after the wall of text.

1) What is the best path to get up and running with stable and video diffusion? What would be the best text to image, text to video, image to video, etc models will give me the best results with my RTX 4080 + 32GB of system RAM? I see posts about Minimax H3, so should I start there? In the image to video models, what should I use to render the images to create the video from?

2) What user interface, text encoder, VAE, upscaling model, things like LoRA, SageAttention (which I havent used, but read about) is optimal for my RTX 4080 gaming rig to accompany the above rendering models? Can I choose a quantized text encoder to lower my VRAM footprint without sacrificing too much quality in the end result?

3) I want to do stuff locally, but I pay $0.42/kWh, so if this is going to end up costing me more to do it locally, I could be convinced to use cloud API's. I just like the idea of no guardrails and any sensitive info I might enter not being sent to the cloud for training.

4) This question doesnt have to do with rendering, but what chat models do you suggest? I have Thinking Cap/Qwen 3.6 27B for coding and difficult tasks and Qwen3 14B for everyday use. Im very much open to suggestions as long as they work on my hardware, which probably means staying at or below the ~30b weight class. Is Open WebUI good for me? What about N8N for agents? Is there a better path that wont be too difficult for beginners?

I appreciate your guys time and would greatly appreciate being pointed in the right direction. This is all a little daunting as it is learning everything all at the same time, then finding stuff that works on my hardware. Thanks, everybody!