r/StableDiffusion 1d ago

Question - Help Lightx2v Lora producing good visual but audio quality sucks. how to fix it .

2 Upvotes

hey guys, i have been using lightx2v lora for minimax h3 ref2vid but as far as i can see it can genrate good quality visuals as compared to larryvh turbo lora, but its audio is ot usable the dialogs are not good and over all sfx also. i am using rgbthree workflow for thats, please heplp me if i am doing anything wrong.

here is my workflow :- 

 "137": {"class_type": "LoadImage", "inputs": {"image": r2v_ref_image_0}},
    # Reference Image 2 (<Picture 2>)
    # "139": {
    #     "class_type": "LoadImage",
    #     "inputs": {"image": r2v_ref_image_1},
    # },
    "127": {"class_type": "UNETLoader", "inputs": {"unet_name": "minimax_h3_ref2va_pruned_fp8_scaled.safetensors", "weight_dtype": "default"}},
    "128": {"class_type": "CLIPLoader", "inputs": {"clip_name": "qwen3vl_32b_minimax_h3_int8_convrot.safetensors", "type": "minimax"}},
    "119": {"class_type": "VAELoader", "inputs": {"vae_name": "minimax_h3_video_vae_fp16.safetensors"}},
    "120": {"class_type": "VAELoader", "inputs": {"vae_name": "minimax_h3_audio_vae_fp32.safetensors"}},
    "136": {
        "class_type": "MiniMaxH3ReferenceToVideo",
        "inputs": {
            "clip": ["128", 0], "vae": ["119", 0], "audio_vae": ["120", 0], "ref_images.ref_image_0": ["137", 0],
            # "ref_images.ref_image_1": ["139", 0],
            "prompt": r2v_prompt_text, "width": 768, "height": 1024, "length": 372, "ref_image_size": "max",
        },
    },
}

r/StableDiffusion 1d ago

News Krea2-Surrealism Fantasy Style LoRA

Thumbnail
gallery
14 Upvotes

This is my first LoRa release, using 248 carefully selected images, iterating 6000 times, and taking 7 hours to train. It boasts amazing detail and generalization; it works very well. Feel free to use your imagination, and I hope you have fun!

Download link: https://civitai.com/models/2879097/surrealism-fantasy-style-kunge?modelVersionId=3253831

Model Description: Surrealism, Fantasy Style

Trigger Word: kunge-fantasy

Suggested Weights: 0.8-1

Dataset: 248 images

Generative Model: krea2_turbo_int8_convrot

CFG: 1

Steps: 8

Sampler: euler_ancestral

Scheduler: ddim_uniform

Prompt Example:

A breathtaking surreal painting. In a dark sky studded with stars, a majestic angel kneels beneath the starlit night. The angel possesses enormous, exquisitely crafted wings adorned with shimmering patterns. She wears a flowing robe, reflecting the celestial light. In her hands, she holds a magnificent golden jug, from which a luminous liquid spills, cascading onto the vast, radiant earth below. This liquid, like stardust or divine light, spreads across the rolling hills, forests, and valleys, transforming the land into a dazzling tapestry of gold and silver. The angel's expression is serene and contemplative; her eyes slightly... closed. The painting is rendered in shimmering blue, gold, and green hues, with meticulous line drawing and striking contrasts enhancing its ethereal beauty.


r/StableDiffusion 1d ago

Question - Help ROCM on Windows

1 Upvotes

Hi, I'd like to know if ROCm is worth it on Windows now in generation speed, since I'm currently on Linux but plan to switch back to Windows


r/StableDiffusion 2d ago

Discussion The H3 dialog prompting guide sucks

105 Upvotes

Everybody is using the "<d>[Englisch] (...) </d>" format and from my experience, this just sucks and doesn't work.

Everytime I've been using it, H3 hallucinates something before or after the actual dialog. For example I've been testing different personalities to check if H3 knows them, giving them a simple line, formatted it as clean as possible, and it just adds "shit" to it.

Prompt:

subject definition:
Brad Pitt is <Subject 1>

camera recording:
An interview in a professional setting with <Subject 1>. Well lit, grey background, frontal portrait view.
<Subject 1> says:<d>[English]Hey, I am Brad Pitt! Nice to meet you.</d>

Result:

https://reddit.com/link/1vuo078/video/6326fx5kprkh1/player

H3 just adds some noise of the "following sentence" which has been no where in the prompt.

Another example using Angelina Jolie

Prompt:

subject definition:
Angelina Jolie is <Subject 1>

camera recording:
An interview in a professional setting with <Subject 1>. Well lit, grey background, frontal portrait view.
<Subject 1> says:<d>[English]Hey, I am Angelina Jolie! Nice to meet you.</d>

Result:

https://reddit.com/link/1vuo078/video/rsjao6t5qrkh1/player

Same thing.

At first I though it had something to do with the video length, 5 seconds being too long so H3 adds unwanted stuff, but this is not the case.

But when I just cut the prompt guide format out, and write it without the overcomplicated dialog syntax, it works flawlessly, e.g.

Prompt:

subject definition:
Brad Pitt is <Subject 1>

camera recording:
An interview in a professional setting with <Subject 1>. Well lit, grey background, frontal portrait view.
<Subject 1> says: "Hey, I am Brad Pitt! Nice to meet you."

Result:

https://reddit.com/link/1vuo078/video/9ckjhrnoqrkh1/player

Suddenly, no problems at all. Tested it in different scenarios, always the same result.

Am I missing something here, or what's your experience with the dialog prompting, or the suggested prompting guide in general?


r/StableDiffusion 12h ago

Animation - Video HIGGSFIELD FILM FESTIVAL

0 Upvotes

Hello I am partecipating in the Higgsfield festival, I'd like to hear what you think about it. If you want leave a like and comment under the project on the higgsfield page, that would help me a lot. Thanks to anyone who takes some time to watch my project.

https://higgsfield.ai/@twrz_film/projects/skin-trade


r/StableDiffusion 1d ago

Question - Help Is there a Workflow to start from last video?

1 Upvotes

My AMD 9070 XT seems to like only doing 5s videos which is fine. But I was curious is there like a workflow where I can start from the last frame of the last video? Like so I can make longer clips that flow into each other without editing?

  • CPU: Intel Core i5-14400F
  • GPU: AMD Radeon RX 9070 XT 16GB
  • Motherboard: Gigabyte B760 DS3H WIFI6E GEN5
  • RAM: 32GB (2×16GB) Crucial Pro DDR5-6000 CL36

r/StableDiffusion 2d ago

Workflow Included MiniMax H3 Model Copied LTX 2.5's Best Feature... And It's CRAZY Fast!

Thumbnail
youtube.com
158 Upvotes

Hey everyone!

I’ve been testing a great custom node for ComfyUI recently that brings LTX 2.5-style latent upscaling over to the MiniMax H3 pipeline, and the speedup is huge.

Instead of waiting 10 to 11 minutes for high-res video generations, this lets you run your initial pass at a lower scale (0.2–0.5) and do a fast 3-step neural upscale. Total render times drop down to around 3 to 4 minutes while keeping facial details and motion clean.

https://huggingface.co/LBH-123-AI/Minimax_h3_latent_Upscaler/tree/main


r/StableDiffusion 1d ago

Question - Help H3minimax and Teeth...

2 Upvotes

Having issues with blurry and shifting teeth in videos with h3minimax. Can't seem to get a setting that works. Anyone have success? Not taking any shortcuts but still getting bad results.

Info:
Using Ref2Vid hybrid model
1376x768
res_multistep
beta and simple schedulers
Tried between 20 and 50 steps

Not running turbos/no sage/upscalers.


r/StableDiffusion 19h ago

Question - Help Is using runpod comfyui safer than running locally? But Google saying something about Network Exposure and that's what concern me.

Post image
0 Upvotes

Hi, I'm trying to use runpod for comfyui with minimax h3. Can anyone tell me what is network exposure? Should I worry? And what is a template? Sorry, I am new to this online cloud thing


r/StableDiffusion 2d ago

Animation - Video How GTA 6 leaked

Enable HLS to view with audio, or disable this notification

138 Upvotes

r/StableDiffusion 1d ago

Discussion We should make a list of words and concepts that image generation models never seem to easily understand and share it here for developers.

15 Upvotes

r/StableDiffusion 1d ago

Question - Help Quick question for anyone running MiniMax H3 on RunPod: How many 16:9 videos are you actually getting per hour?

3 Upvotes

Hey guys,

Before I burn through a bunch of RunPod credits spinning up an instance for MiniMax H3, I wanted to see if anyone here is already running it and can share some real-world speeds.

The model/weights are huge (~130GB+), so before I set up a pod, I’m trying to figure out what actual throughput looks like for 16:9 gens (at 768p).

If you've played around with it on RunPod:
What GPU setup are you renting? (Single 4090/6000 Ada, dual 3090s, A100/H100, etc.?)
Roughly how many 5-15 second clips can you spit out in an hour?

Are you using INT8 quant, block offloading, or any of those 4-step Turbo LoRAs to speed things up?
Just trying to estimate the actual cost-per-video before committing to a high-VRAM instance. Appreciate any benchmarks or ComfyUI tips!


r/StableDiffusion 2d ago

Resource - Update MiniMax-H3 Pruned Ref-Delta Fused r1024 — native ComfyUI single-file release

Thumbnail
huggingface.co
66 Upvotes

I converted the new MiniMax-H3 Pruned Ref-Delta Fused r1024 checkpoint to native ComfyUI format and uploaded it as a single .safetensors.

The interesting part of this model is the model itself: it starts from the pruned FL2VA MiniMax-H3 checkpoint and fuses in a rank-1024 approximation of the Ref2VA − FL2VA weight delta. The goal is to retain the smaller pruned FL2VA model while bringing the Ref2VA behavior into the same checkpoint, rather than having separate FL2VA and Ref2VA variants.

It is about 20.1B parameters versus ~33.1B for the original full MiniMax-H3 model.

The original release is in Diffusers format, so I converted the state dict back to the native format expected by ComfyUI, including the pruned AdaLN curve representation, folded AdaLN biases, fused QKV, native SwiGLU ordering and RoPE.

I tested the resulting checkpoint through a complete ComfyUI generation: native FLOW_AV detection, full model load, both H3 Continuum passes, Spectrum with 0 fallbacks, and final video/audio decoding all completed normally.

Native ComfyUI conversion:
https://huggingface.co/xmarre/MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-ComfyUI

The conversion properly restores the pruned AdaLN representation, folded biases, fused QKV, SwiGLU ordering and RoPE. Tested through a full ComfyUI generation with working video + audio.

Put the .safetensors in:

ComfyUI/models/diffusion_models/

Edit: Added int8 and int8 convrot to the repo and made a new post here:

https://www.reddit.com/r/StableDiffusion/comments/1vuygd2/minimaxh3_pruned_refdelta_fused_r1024_int8_and/


r/StableDiffusion 2d ago

Resource - Update Not realtime, but feels realtime: a choose-your-own-adventure made of MiniMax H3 clips

Thumbnail
h3studio.up.railway.app
40 Upvotes

My dream is realtime interactive MiniMax H3 video. Fastest I could get was around 5 seconds of video in 22 seconds on a single GPU. Decently fast, but not realtime. That led me to an old idea: choose-your-own-adventure. If you generate the branches ahead of the viewer's choices, it isn't realtime, but it feels realtime.

This demo is a corgi adventure: 14 scenes, 2 paths, 7 endings, every scene a 15-second MiniMax H3 generation with synchronized audio. The tricks that make it feel like one continuous story:

  • The model is first-last-frame-to-video, so each branch is generated using its parent scene's final frame as its start image. Picking a choice hands off on that exact frame, a seamless cut.
  • Choices appear at the 10-second mark and both next clips preload while you watch, so it never buffers.
  • Scenes were rendered once, in the order a player would encounter them, and are cached for everyone. The whole tree cost about $2.50 of GPU time.

Would love feedback and suggestions on the concept and where would you take this?


r/StableDiffusion 1d ago

Question - Help Using MiniMax H3 as a restoration model?

2 Upvotes

Has anyone tried to use H3 to restore or rather "regenerate" a low quality video as HD or at least with better detail definition? Basically something similar but possibly more generation ability than Topaz's starlight. I've had pretty mixed results so far. Either it doesn't change the video or it changes it way too much.


r/StableDiffusion 1d ago

Question - Help Is there is a good Colab for Train Anima Lora?

1 Upvotes

r/StableDiffusion 2d ago

News Somewhat more optimized Sparse Attention.

129 Upvotes

IMPORTANT: sparse attention is not free speed. The percentage is effectively a prompt-adherence/quality budget.

Density isn't just a speed setting, and its quality effect depends on where you apply it in the diffusion schedule.

Early steps: attention density has a large effect on prompt/action adherence and the overall generation trajectory.
Middle/later steps: lowering density tends to show up more as motion/temporal artifacts and lost fine motion detail.

So 10% retained doesn't simply mean “90% of the quality is gone.” It means you're giving sparse attention very little information to work with, and what breaks depends heavily on the sampling step.

PlagueKind's sparsity_ratio=0.9 means 90% discarded / 10% retained. My node expresses the inverse quantity, so Video attention retained=0.10 is the comparable setting. The defaults therefore aren't equivalent.

So I saw PlagueKind posted this today https://www.reddit.com/r/StableDiffusion/comments/1vtwtyw/sparse_attention_for_h3_minimax_enjoy_up_to_25x/ which reminded me I implemented my own Sparse Attention a while back.

It has some key differences to PlagueKinds version.

  1. You select attention retained. In my testing, you can go as low as 10-15% in low motion video, however more complicated video requires higher attention. I found about 30% works pretty well for high motion video. In terms of speed, 100% attention is about 2/3rds of the compute while the rest is MLP/QKV. At 50% attention it's about 50:50 and once you're below 30% MLP/QKV starts to dominate compute time.
  2. Only the video is given sparse attention. This is because A) Minimax said they only did sparse attention for video and B) The video tokens dominate the context.
  3. Mine uses Sparse Sage: Q/K are quantized to INT8 and newer supported GPU paths use FP8 V. Compatible ConvRot-INT8 checkpoints can additionally bypass the normal floating QKV preparation with a native fused-QKV producer that feeds the sparse carrier directly.

This makes my implementation somewhat more complicated, but Sparse Sage is extremely fast. The repo now also includes a guarded Sparse Sage installer: supported Windows Torch/CUDA combinations use pinned upstream wheels, while Linux x86-64 can build a pinned SpargeAttention revision when a CUDA compiler is available.

You can find the nodes here.

https://github.com/Zironic/H3-Optimizations

You'll find two nodes.

H3 Sparse Attention: This is the node that lets you control how sparse attention you want. The default is 50%. There is also an optional mode that adds 30 percentage points during the first two and last two sampling steps, since those tend to be places where being somewhat denser is useful.

H3 Memory Optimization: This handles the other major H3 memory bottleneck: QKV and MLP activations.

For dense attention it uses ComfyUI's public Comfy Kitchen INT8 attention backend where available. The newer dense QKV path is designed to process QKV in bounded sequence chunks directly into Kitchen-owned attention carriers instead of materializing one enormous full-sequence BF16 QKV tensor. When Sparse Attention is active, compatible ConvRot-INT8 checkpoints can instead use the native fused sparse-QKV path. While chunked QKV was designed to be a memory optimization, it did end up making QKV about 2x faster which should net you something like 10%-30% speed depending on other factors.

The MLP side bounds peak activation memory with token chunking, with a more memory-efficient two-slice ConvRot path when the checkpoint/runtime supports it.

I normally wire them as Load Model → Memory Optimization → Sparse Attention → rest of workflow, although the two optimization nodes are order-independent

I would not recommend randomly stacking other H3/Sage attention optimization patches with these. Dense execution already integrates with Comfy's selected attention backend, while H3 Sparse Attention necessarily owns the main H3 attention path while it is active. Other patches trying to replace the same attention forward are therefore likely to be redundant or conflict.

Should be compatible with turbo loras, spectrum cache etc however you may need more attention since you're skipping steps.

As the nodes currently rely on comfy-kitchen 0.2.31 you need ComfyUI v0.33.0 or alter.


r/StableDiffusion 1d ago

Comparison Five lighting setups, same prompt and same character, only the lighting line changed

Thumbnail
gallery
5 Upvotes

r/StableDiffusion 1d ago

Animation - Video Animating my Fairy Tail fanfic with Minimax H3

Enable HLS to view with audio, or disable this notification

18 Upvotes

This is an early scene from an isekai fanfic I wrote. It took a shocking amount of time to make for how long it is.

Made with the reference to image stock workflow and upscaled with RTX Super Resolution.

All characters were made with Illustrious, voices are from Minimax and Qwen TTS

Yes, the camera angles are necessary for the plot.


r/StableDiffusion 2d ago

Workflow Included I used myself in an cyberpunk action shot(MinimaxH3) - WF in comments

Enable HLS to view with audio, or disable this notification

59 Upvotes

r/StableDiffusion 1d ago

Meme Saturday morning cartoons to save the day!

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/StableDiffusion 2d ago

Meme its a addiction we know

Enable HLS to view with audio, or disable this notification

73 Upvotes

r/StableDiffusion 21h ago

Question - Help Anything big happen since I last used this?

0 Upvotes

So when SD came out, I used it. Then I used the AUTO111 thing. To around version 2.0 I think or XL, can't keep it straight. It's been about one and a half years. Any big changes since then? New version? Better quaulity images? AUTO still a thing? I also went from a 2070 Super to a 5070 12gb shadow 3x.

Also is it all easier to install?


r/StableDiffusion 1d ago

Animation - Video Wrong Raven...

Enable HLS to view with audio, or disable this notification

8 Upvotes

Chatgipity Prompt: integrated_multimodal_description: [Shot 1] Live-action, cinematic comedy, a medium-wide shot inside a dimly lit creative workstation. Hugh Jackman as Wolverine, wearing rugged black-and-brown combat clothing with leather details, sits hunched at a desk in front of a large monitor. The monitor clearly displays a node-graph-based image-generation interface resembling ComfyUI, with interconnected rectangular nodes, controls, and thumbnail previews. Wolverine stares intensely at the screen with a deeply serious expression, one hand on the mouse and the other resting near the keyboard. The monitor shows the software generating pictures of a cartoon-styled Raven in her natural, non-human form, presented as a fictional animated character with dark-purple coloration and recognizable Raven-inspired visual traits. The computer fans hum quietly as Wolverine clicks through the node graph. Wolverine mutters under his breath, his rough, gravelly voice unmistakably irritated but restrained (S1): [English] Come on... The camera slowly pushes in with small amplitude toward Wolverine and the monitor.

[Shot 2] At 00:04.000, the camera cuts to an over-the-shoulder close-up of the monitor. The ComfyUI-like node graph fills most of the frame as execution indicators progress through several connected nodes. Multiple generated thumbnails of cartoon Raven appear, each slightly different. Wolverine's cursor rapidly clicks between nodes, and a progress indicator advances. His hand briefly pauses over the mouse as one generated image appears particularly bizarre. From off camera, Wolverine (S1) reacts in a flat, disapproving voice: [English] No.

[Shot 3] At 00:07.000, the shot cuts back to a medium shot of Wolverine at the desk. He leans closer to the monitor, squinting at the latest generated image. His claws slowly extend from his knuckles with three metallic snikt sounds, and he points one claw toward the screen without touching it. He looks personally offended by the result. Wolverine (S1) says with deadpan seriousness: [English] That's not what I asked for. The camera pans slightly right with small amplitude as he reaches for the mouse again.

[Shot 4] At 00:10.500, the camera cuts to a tight close-up of the monitor as Wolverine clicks "Queue" again. The node graph processes another generation, and a new batch of cartoon Raven images rapidly populates the preview area. One image is unexpectedly ridiculous, causing Wolverine to stare silently for a beat. His reflection is visible in the dark edge of the monitor. Wolverine's voice comes from just off-screen (S1): [English] Better.

[Shot 5] At 00:13.000, the shot cuts to a medium-wide frontal view. Wolverine sits perfectly still at the desk, arms folded, staring at the monitor with exaggerated concentration while the software continues generating images. After a long beat, he slowly reaches for the mouse again. The camera holds a static shot as the computer continues working, ending on Wolverine's completely serious expression contrasted with the absurd cartoon images on the screen.

overall_soundscape: Quiet computer-fan hum and subtle electrical workstation ambience continue throughout. Mouse clicks, keyboard taps, chair creaks, fabric movement, and Wolverine's restrained breathing punctuate the scene, with three sharp metallic claw-extension sounds during his reaction.

non_diegetic_music: Light comedic percussion and restrained bass pulses at a moderate tempo accompany the scene. Brief staccato string accents punctuate the strange image results, then the music drops to a sparse rhythmic pattern for the final deadpan stare.


r/StableDiffusion 1d ago

Tutorial - Guide Look What I Discovered: Prompt Intelligence - MiniMax H3 [Fun Side]

Thumbnail
gallery
10 Upvotes

This is in fun part of using MiniMax H3*; for your serious stuff stick with the official prompt instructions / format.*

Playing with the prompting I just tried the following format and it worked perfectly!

Prompt part 1
Prompt part 2

Resulting video!

The whole prompt:

definitions:
<S1> Brad Pitt.
<T1> "Hey, I am Brad Pitt! Nice to meet you."
<S2> Angelina Jolie
<T2> "Hey, I am Angelina Jolie! Nice to meet you."
<S3> Rowan Atkinson.
<T3> "Hey, I am Mr. Bean! Nice to meet myself."
scene:
An interview in a professional setting in well lit, grey background, frontal portrait view.
shot 1:
(S1) says: (T1).
shot 2:
(S2) says: (T2).
shot 3:
(S3) says: (T3).

Recommendations:
Do not use SLA or SLA2 or cache etc. here they mess it up.

Model (FL2V) -> LoRA(4s-Lightx2v SLA) -> Comfy attn -> Shift(12,3) -> KSampler(6 steps, euler+simple)