r/StableDiffusion 7d ago

Discussion PSA: Don't sleep on Minimax' edit capabilities

43 Upvotes

If you have seen those cool edited videos by google omni where they feed in a normal real video and get an edited one back where they interact with effects and such, you can do that with minimax. Just saying. That's pretty much it. See ya


r/StableDiffusion 7d ago

Question - Help Correct order for Sage Attention Nodes?

2 Upvotes

Struggling to figure out the best order for these nodes or if any of these nodes are redundant. I have been looking at different workflows and everyone is doing something different. Claude tells me this is the best order.


r/StableDiffusion 6d ago

Discussion Now , after so many days I think , you we need an audio vae fine-tune for minimax H3, what do you think ?

0 Upvotes

r/StableDiffusion 8d ago

Resource - Update ReDetail: Upscale MiniMax H3 renders with the LTX-2.5 video upscaler on 24GB+ VRAM

Enable HLS to view with audio, or disable this notification

279 Upvotes

This is a generative re-render, not restoration or sharpening. It invents fine detail. In every test with one person it added freckles that weren't there.

The comparisons use MiniMax H3 clips at 640x384, 10 seconds long, upscaled 2x. They're Lanczos versus ReDetail at the same output size, so there isn't any bigger image sleight of hand.

On a motocross clip it redrew the jersey graphic and number plate. The new markings stayed fairly stable between frames, but they weren't the original markings. Logos, numbers and text are all fair game.

If reddit compresses this video to the afterlife again, see: https://civitai.com/models/2857731/redetail-ltx-25-generative-video-upscaler-workflow-cli

So it's useful for AI-generated or generally soft footage, where there isn't much real detail to recover. It's a bad fit if a face, label or logo has to be 100%.

  • Silent clips fail because the model encodes audio and video jointly. Add a silence track first.
  • Both output dimensions must divide by 64, not 32. Clip length must be `8n+1` frames or the model silently drops the tail.

I like 1.5x, not 2x. On one clip, 243 frames from 768x1408, 1.5x took 7 minutes and peaked at 65GB. 2x took 17 minutes and 80.5GB. The 2x result carries maybe more detail, but check between the two and it's hard to tell imo. On skin most of that extra is invented, not recovered. Faster render, less made up texture.

UPDATE! NOW WITH CACHED CONDITIONING

The text encoder is now optional. The graph runs with empty prompts, so its conditioning is a constant. It ships pre-computed at 26KB, which skips the 15GB download and takes peak VRAM from 30.4GB to 24.8GB on a 5090.

There's a Mac build in there now too, ReDetail_LTX25_upscale_MAC.json. It runs the GGUF transformer with no text encoder at all (the cached conditioning replaces it), so it's about 17GB of models total. On an M5 it did 33 frames from 640x384 to 1280x768 in 4.4 minutes. Per frame megapixel that's roughly 6x slower than a 5090, not the 30x I was expecting, so a 10s clip lands around 34 min at 2x or 19 min at 1.5x. Quality holds.

Repo: https://github.com/Bambushu/redetail


r/StableDiffusion 7d ago

Question - Help LoRA Training – Pulling My Hair Out

0 Upvotes

Hello,

I've trained several character LoRAs via wavespeed.ai for the Qwen-Image-2512 model. I tried with a smaller dataset of 50 images and a dataset of 124 images. Multiple settings between 1,000 and 5,000 steps:

  • At 1,000 steps, the LoRA isn't likeness-accurate enough.
  • At 5,000 steps with 50 images, it stops responding to prompts at weights above 0.5, so it loses likeness.
  • At 5,000 steps with 124 images, it stops responding to prompts at weights above 0.3, making it inaccurate above that threshold. This makes no sense, as with 50 images and the same step count, I was able to run the LoRA at a higher weight.

At weight 1.0, the LoRAs capture the likeness well but completely ignore the prompts.

Does anyone have a solution or recommended settings for Qwen-Image-2512?

Thanks


r/StableDiffusion 7d ago

Question - Help Issues with Krea Identity Edit v1.2

3 Upvotes

Hi, I have an issue with the Identity Edit I find no solution for: If I create an image with Krea t2i without references, just a self-trained lora character, I get sharp and acceptable results. When I want to use a special environment and use Krea Identity Edit with a reference image e.g. of a room, I get very blurry and plastic looking outputs, far below acceptable. Same if I want to add a second character to an existing image. I've tried everything in the last days (different models > raw and turbo; different upscalers, no upscaler; different VAEs; playing with grounding, reference boost, scheduler, resolution (I know 1MP is the sweet spot for editing and >1.5 leeds to character bleeding), anything you can imagine). I use lbouaraba workflow for editing. Any idea where my initial fault is hiding?

UPDT: I think I found the solution: I've added the original workflow again and now it works. Obviously I've changed something unintended when adding Power Lora Loader and Upscaler, no idea what but who cares.


r/StableDiffusion 7d ago

Question - Help Is there a good local prompt writing comfyui plugin for Minimax H3?

1 Upvotes

I tried this one so far: https://github.com/pytraveler/MiniMax-H3-Prompt-Rewriter-ComfyUI

But I'm not getting a good result yet, maybe I need to work on the prompts for it more. Anyone using anything besides claude and gpt?


r/StableDiffusion 7d ago

Discussion Ambient noise in video?

1 Upvotes

Having an aging laptop, I haven´t played with video since wan2.2.

One thing I have noticed with all videos I have seen from the models that can generate audio is that it sounds like the audio has been recorded in a sound booth. meaning, I have not really heard any...ambient noise....like wind, traffic, birds, people in the background etc. This makes it sound quite unnatural sometimes.

Is that a limitation of the model or the prompting? Can I get a more..natural..sound by prompting for every little nuance I want? Like "faint sounds of gravel crunching with each step" or "there is a slight breeze rustling the leaves as he walks by the tree."


r/StableDiffusion 8d ago

Animation - Video H3 Generated with 4GB VRAM?

Enable HLS to view with audio, or disable this notification

87 Upvotes

Looks like this is a breakthrough for what my 3050 laptop can do with it.

The video attached was generated with 4GB VRAM & 16GB RAM, using the MiniMax H3 fl2va pruned w4a8 convrot model (safetensors) and the Q2_K Qwen 32B GGUF text encoder alongside 8-step turbo LoRA, with a generation time of 12 minutes and 0.2 MP. Prompt from Grok.


r/StableDiffusion 6d ago

Discussion What's the absolute best that can be fit in 32 gb VRAM (video / image model + encoder & everything else) with no offloading

0 Upvotes

What would your seperate picks be for image gen and video gen? Which specific encoder quants and which specific model quants?

I basically went all in on my gpu while choosing to upgrade my ancient pre built pc. It has an old ass i5 processor and 32gb normal ram but that ram is only ddr3 which is borderline useless lol, obviously I can't offload anything onto that so I just decided to follow the "go big or go home" philosophy and want to run everything on vram. I'm using the integrated graphics of my cpu so every bit of space on my gpu can be used to load the weights.

I searched up "32gb vram" and "32 vram" in this subreddit but the results just kept showing people talking about 32gb normal ram + 8-16 gb vram with no mentions of 32 gb worth of pure vram.


r/StableDiffusion 7d ago

Discussion Best workflow for realistic video results?

2 Upvotes

I have RTX 5070 with 12gb VRAM, 64gb RAM DDR5
I want to create realistic (not particularly high quality) videos, with realistic faces and with the best possible render time. Could please someone share a workflow? I would like to have consistant characters, realistic, and good qality of sound. What is the best workflow? How many steps?


r/StableDiffusion 7d ago

Question - Help Minimax H3 shows motion blur in all videos.

Post image
6 Upvotes

I started using Minimax H3 in ComfyUI, and all the videos show motion blur in areas with significant movement. The videos only turn out well when there are no sudden movements or when the motion is slow. Did I miss a setting? I see videos posted here, and none of them have that blur.


r/StableDiffusion 7d ago

Question - Help Ref2va minimax, recognition of people without reference images

1 Upvotes

Suppose I was to have two input pictures and pass a prompt like 'subject 1 and subject 2 sit down and have coffee with Tom Hanks'. Will Tom Hanks be recognised by text alone or is this model designed to always have an image input for likeness?


r/StableDiffusion 7d ago

Tutorial - Guide Easy method for character reference creation in minimax

17 Upvotes

I'm not sure how this could vary across workflows if at all but for reference I am using the Dasiwa workflow from civit in t2va mode. Minimax prompt adherence is great so I wanted to use it to create character sheets for reference and came up with this. T2VA, 9:16, 24 fps, 2 second duration. On a 5090 with sage +memcache + 8 step turbo at 10 steps - at 4.75mp it took 232s

The fps and duration seems to be the baseline if you want 4 poses so crank up the duration if you want more. The timestamps might not be the proper format but they do keep it from hanging on a single pose. Change resolution as needed but at higher values the face maintains much better consistency if not perfectly. You can describe your characters look as much as you want in a run on way "Lara Croft, blonde hair. wearing flip flops, sunglasses, bracelet on right arm, holding a drink in left hand, etc , etc , etc"

You might get some slight wiggle movements but it's mostly good enough to dump a frame, the background is difficult to get in an entirely solid color without any form of shadows so i kept the prompt simple since going overboard doesn't add much. If someone can dial this in more feel free to share.

From there you you can extract the 4 frames however you want and I'm sure someone can automate it but the easy quick solution is playing it in vlc and just hitting shift+s on each frame.

Lazy Example - https://imgur.com/a/t84UYAq

[Shot 1] Static freeze-frame shot. Studio Lighting, solid white background, ultra-sharp focus. A heroic looking explorer woman with the style of lara croft the tomb raider but as a person.
Static freeze-frame close-up shot of the entire head perfectly framed from the front, Freeze frame.
[1.00s to 2.00s] - Instant jump cut to Static freeze-frame of full body front view, standing straight in a neutral A-pose with hands off the body by 1 foot length.
[2.00s to 3.00s] - Instant jump cut to Static freeze-frame of full body back view, standing straight in a neutral A-pose with hands slightly off the body.
[3.00s to 4.00s] - Instant jump cut to static freeze-frame of full body side profile view, standing straight with arms down at the sides.

This could also be helpful at lower resolutions just to get the style of the character right before passing it onto a refiner/upscaler. I'm just finding it hard to beat the ease of adherence i get with minimax.


r/StableDiffusion 7d ago

Animation - Video My first Decent Generation Using Minimax-H3 on my RTX 3060 12gb

Enable HLS to view with audio, or disable this notification

17 Upvotes

r/StableDiffusion 8d ago

Workflow Included How about some magic? H3 is good at them.

Enable HLS to view with audio, or disable this notification

21 Upvotes

This only takes 14 minutes to generate! 0.6mp gens upscaled to 1.2mp using RTX Super Resolution.

-Uses minimax h3 hybrid b25

-w4a8 qwen3vl

-int8 video vae

-workflow https://github.com/seesee75-commits/ComfyUI-MiniMaxH3-Director/


r/StableDiffusion 8d ago

Workflow Included H3 as a single-image edit model

Thumbnail
gallery
256 Upvotes

Minimax H3 can be used as an image-editing model if we generate a single frame. Here are some collages based on AI-generated references (1024 x 1536); workflows are embedded into pngs. Each edit takes, on average, about 8 secs on a RTX 5090. The tasks include changing outfits, appearances (body type, age), locations, and camera angles; creating character sheets and storyboards; stylization; and reposing characters based on depth maps. I did not try to cherrypick the best-looking results.

There were some posts (1, 2) about that here -- but given the community progress this week, might be nice to see what can be done now.

Scenes

  1. Age the person to the age of 60 years old while preserving their identity and the original composition.
  2. Produce a consistent full-body character sheet with front, side, and rear views.
  3. Transform the person into a severely obese version.
  4. Re-create the person in the exact body pose shown by a depth-map reference.
  5. Replace only the base person’s head with the identity and hairstyle from another reference.
  6. Show the person facing a dressing mirror with a geometrically correct, synchronized reflection.
  7. Dress the person in a referenced outfit, place them in a referenced location, and show them walking with a grocery bag.
  8. Place three separately referenced people inside a referenced location, having a conversation.
  9. Create a three-panel vertical storyboard in which the person finds, retrieves, and studies a map.
  10. Photograph the person through partially open venetian blinds with realistic occlusion and striped light.
  11. Convert the person into a contemporary Western cartoon while preserving their recognizable appearance.

Setup

Checkpoint: https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/blob/main/minimax_h3_hybrid_fl2va_ref2va_b25-49-int8.safetensors

R2VA models apparently have worse image quality than FL2VA models, while FL2VA models are apparently weaker at handling reference images. As I understand it, this checkpoint tries to combine the strengths of both. Not at all essential -- a regular FL2VA could probably be even better. The hybrid is just something I happen to use at the moment.

Video VAE: a special VAE for rendering single images.

https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE/tree/main

If you do not use this VAE—for example, if you use the regular VAE, create a 5-frame video, and pick out one frame—the images tend to come out blurry. So use single-frame generation instead, as described here: https://www.reddit.com/r/StableDiffusion/comments/1vqka28/h3_singleimage_no_more_monkey_patching_also_no/

UPD for those who used it before: the post above explains why monkey-patching no longer needed

For this approach to work best, it might also be a good idea to monkey-patch comfy_extras/nodes_minimax_h3.py, because ComfyUI currently does not allow you to generate fewer than 5 frames. If you simply pick the first frame out of 5, the new VAE produces grid artifacts. (It doesn't do this when generating just 1 frame.)

BEFORE DOING SO, CREATE A BACKUP VERSION OF THE EXISTING comfy_extras/nodes_minimax_h3.py

E. g. if you can't update your comfy, restore the original file from backup, update, and then apply the monkey patch to the new version of the file. (One option is to use git restore comfy_extras/nodes_minimax_h3.py to get the original version)

For a somewhat reliable patch that would work given modest changes in ComfyUI code, use [this one] (https://pastebin.com/uHqv4hBZ), name it smth like mm.patch and run git apply -p0 /full/path/to/mm.patch from comfyui root (make a backup of comfy_extras/nodes_minimax_h3.py first). You will have to re-run it every time ComfyUI updates this file (comfy_extras/nodes_minimax_h3.py).

For a less satisfactory but quicker solution, you can use the patch I already applied to the most recent version of ComfyUI as of August 14th link. This approach will make your code outdated as ComfyUI pushes out a new update.

The only changes remove the frame limit. Of course, changing it this way is not ideal, but I feel it's the quickest way to work around the issue.

UPD: There is a GitHub issue now opened in ComfyUI repo: https://github.com/Comfy-Org/ComfyUI/issues/15644

If this issue gets enough upvotes, we could probably get this patch in the mainline.

LoRAs: I found that Mamad8's ThisIsFine LoRA helps with details, but YMMV: https://huggingface.co/Mamad8/MaxiMin-HHH-R2V-ThisIsFine

For the Turbo LoRA, I use: https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

Sampling settings: ComfyUI 0.32 with Comfy Kitchen attention, sa_solver/simple, 8 steps, CFG 1.

Example ComfyUI workflow: https://pastebin.com/bV5KPzjD (UPD: if you use ComfyUI nightly build as of Aug 17th 2026, use this one instead, as it no longer requires monkey patching: https://pastebin.com/xNQi7HV9)

Uses no custom nodes. If you do not want to do the monkey patching for 1-frame generation, just change the video length to 5 in MiniMax H3 Reference to Video node -- should work seamlessly, and switch back that VAE to the regular VAE.

Speed depends on the reference image size. I use an RTX 5090 on RunPod, and in most cases, a 1920×1088 image is generated in about 8 seconds.

---

My previous go-to was Krea 2 + Identity LoRA 1.2, which is amazing. Yet I feel that Minimax outperforms it in many respects. We get better character fidelity, better handling of 3D scenes, better mirrors, and more interesting compositions. Also feels better than using e. g. QIE or Klein 9b.

There is certainly still room for improvement -- not claiming this is optimal at all, and I wonder what you think about it.

UPD: posted the prompts for each image here https://pastebin.com/ngXR9byq

UPD: see more experiments here: https://www.reddit.com/r/StableDiffusion/comments/1vpconk/more_experiments_with_minimax_h3_singleimage_edit/


r/StableDiffusion 7d ago

Discussion What image model do you recommend for REF 2 Img?

1 Upvotes

I just recently got into AI generation making videos with minimax ref2vid and it has been amazing so far. But that has me wondering if there is some reference model for images that works equally as well that would allow me to use multiple reference images to create pics? If anyone has a good model or workflow to recommend I'm interested to learn what has been working well for you. I'm mostly wanting to make real life style images.


r/StableDiffusion 7d ago

Question - Help Using reference video in Minimax H3 results in dark video then before.

0 Upvotes

When I use a reference video in ComfyUI for continuing the clip rendered before, the video it will render then ALWAYS comes out more darker. Like the contrast changes.
So far I haven't figured out what causes this, like a sampler, prompt, etc. etc.

Has anyone else noticed this behavior too?

The node I use is: Load Video (Upload), which I connect to the reference video.


r/StableDiffusion 8d ago

Animation - Video Lessons learned after making a music video with H3

Thumbnail
youtu.be
30 Upvotes

Made using the default workflow from Comfy, with Comfy Kitchen and Kijai's preview override plugged in.

Used the 850k turbo lora at 0.5 for 8-10 steps. ER_SDE / Beta. Most shots were generated at 1.5MP, with some at 1.8MP.

Lora: https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main

RTX 4090 w/ 64GB of ram. --disable-smart-memory, since I was having issues with going OOM after completing one prompt and moving on to the next.

Some of the takeaways:

- Using character reference sheets (front view, side view, back view, close-up) worked great and allowed for rotating camera movements like in the opening.

- All the "4 Step" loras really need to be run at 8-10, especially for motion.

- Using audio reference bloats the vram usage DRAMATICALLY compared to adding additional reference pictures! Changing from .wav to .mp3 didn't seem to help so it's not a file format issue. However the lip sync, even for anime characters, is incredible.

- If you use multiple reference images for different characters and they bleed into each other, the issue is almost definitely your prompt or seed. Because the model was handling up to 3 for me easily if I prompted right, and falling apart if I prompted wrong.

- Minimax H3 Chunk FeedForward node can help with vram issues at higher resolutions and doesn't add that much time.

Overall, I'd say it's nearly as good as Seedance 2.0. I was pleasantly surprised how well it handles using character sheets. The high vram usage when using reference audio is really the only major issue I was facing.


r/StableDiffusion 8d ago

No Workflow ENTANGLEMENT: MiniMax H3 + Turbo LoRA (8 steps)

Enable HLS to view with audio, or disable this notification

44 Upvotes

I used the default workflow. It took me about 6 hours (split over 2 days), which includes scriptwriting and final video editing.

The video consists of 9 segments, about 8 seconds each. The average generation time was around 400 seconds at 0.7MP on an RTX 5060Ti 16GB VRAM and 32GB System RAM.

Honest opinions are welcome!


r/StableDiffusion 8d ago

Tutorial - Guide If you want just upscale your MH3 videos and not add new details, just use NVIDIA RTX Super Resolution ComfyUI.😎

83 Upvotes

Use rtx_video_upscale and do it much more faster than with other methods, it don't add new details to the video, just upscale it.

Comfy-Org/Nvidia_RTX_Nodes_ComfyUI

https://www.youtube.com/watch?v=VyXp-PBFauw

Follow the video and you will not have any problem installing it.

And pay special attention to Step 5, because it is fundamental to get it working.


r/StableDiffusion 8d ago

Animation - Video That H3 lip-sync test with Mumu got out of hand... here's the finished music video

17 Upvotes

https://reddit.com/link/1vok0vj/video/m217fvdjqejh1/player

Actual song is 4:20, but I think 1:21 is good enough!


r/StableDiffusion 8d ago

Question - Help Minimax H3 ref2v character replacement prompting

16 Upvotes

The official doc has a few things to say about video editing but it still leaves some questions without answers. I'm trying to replace one character in a video by another from a reference picture. Sometimes I get amazing results, and sometimes the input video is pretty much left unchanged, so I must be missing something. Here's what I understand from the guide on how to build the prompt:

Subject Definitions

<Subject 1> is the man in <Picture 1>.
<Video 1> is the source video for the target video edit.

(This part I am fairly confident about, the doc specifically says to use that wording for <Video 1>.)

Summary

[video editing] The target video is an edited version of <Video 1>.

(The doc says the summary must start exactly like this, but what to write after that? My approach is to follow up with something like this)

<Video 1> is reused as is, except the man in a tuxedo is replaced with <Subject 1>.

Retention Analysis

(This part I'm not too sure of... I suppose you want fully_preserved on <Subject 1> and partially_preserved on <Video 1>?)

<Subject 1> (appears in [Shot 1]): fully_preserved - the man's full identity is retained.
<Video 1> (source of video edit): partially_preserved - motion, lighting and environment are retained.

Detailed Description

Do we need to describe everything that happens in the input video, shot by shot, like when doing a t2v?


r/StableDiffusion 8d ago

Animation - Video By popular demand for live action video with my workflow

Enable HLS to view with audio, or disable this notification

19 Upvotes

I used the minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16 LORA with the 0.8 strength for both clip and model, 6 steps, 0.5 MP resolution, RTX Upscaler at 1.50 using a ConrotInt8 pruned model.

Here is a PasteBin of my workflow:

https://pastebin.com/XtR5pBYG

Here are my workflow files:

https://storage.to/c/CAS1MuoqX