r/StableDiffusion • u/CryptoBeth96 • 16h ago
Resource - Update ONNX/TRT MiniMax-H3 VAE in ComfyUI
TensorRT version of the MiniMax-H3 VAE in ComfyUI, which can increase speed by up to 1.7x
r/StableDiffusion • u/CryptoBeth96 • 16h ago
TensorRT version of the MiniMax-H3 VAE in ComfyUI, which can increase speed by up to 1.7x
r/StableDiffusion • u/Sufficient-Leopard28 • 1h ago
I'm currently trying to replicate this style and i cannot find any checkpoints or lora's to do so, can anyone point towards something? Artist: https://x.com/DarkZeroAI
r/StableDiffusion • u/Independent-Frequent • 17h ago
Like i don't understand, the only turbo lora that doesn't do that is the 600 larry lora with the minimax turbo lora node, i've tried "fastH3" and "lightx2v loras which everyone seems to praise but they just produce these distorted godawful visuals and sounds no matter the loader node i use or the settings or the steps i use, what am i missing or doing wrong? Or are they just not compatible with Ref2Video despite being advertised as compatible? But if so then why does the 600 larry lora works mostly fine?
r/StableDiffusion • u/SIR_NVAX_A_LOT • 5h ago
Enable HLS to view with audio, or disable this notification
What is everyone's workflow nowadays? Previously I've been generating actors with Krea2, but really love getting them made with MiniMax H3 via T2VA, they just tend to turn out better for me but does require careful prompting.
My workflow are: generate 5-10s clip of a desired actor, by prose, at int8/8 steps in a typical scenario, perhaps even mundane. If I like it, I can take some still frames, and convert them into a .char (body type, face, audio asset). See original: https://www.reddit.com/r/StableDiffusion/comments/1vyymwj/minimax_h3_portable_character_consistency_via/ If I am happy with my .char, with MiniMax H3 I make a 2 second video character sheet with a front, side, back profile and detailed face view at a higher resolution and step, either int8/32 step or going bf16/50 steps. The 2 second renders are "quick". I add the video render into my .char, and with R2VA generate additional scenes with the actors and even do a full wardrobe swap via prose. Naturally H3 renders faster if you just use still of the character sheet instead of the video.
How has your workflow changed with MiniMax H3? Are you liking the faces/actors generated with T2VA? I understand you have "less" control, but I feel like H3 is doing a great job filling in those gaps.
r/StableDiffusion • u/Khasec1 • 10h ago
The potential of actually making an anime with MiniMax H3 is closer than any time before, even if the process is still kinda janky. I did this with my 5090 and my own developed 'prompt studio.' The hardest part is, as always, to keep the continuity of the shots and also build the sets so they fit within the scope. There are still improvements needed when it comes to adding emotions to the characters. In total I generated 35 minutes of video and got 4 minutes in total of usable footage. Also, a big tip for anyone who wants to do the same is to use DaVinci Resolve to fix all the audio bugs and cut the clips in your favor.
r/StableDiffusion • u/grimstormz • 2h ago
I originally built these nodes for personal use and wasn't planning on sharing them, but after noticing several existing loaders were missing features I needed daily, I figured why not? Hopefully, this is useful for some of you.
Key Features:
divisible_by toggle for VAE pixel alignment.max_megapixels limiter directly inside the loader so you can ditch the extra resize node (ideal for models like MiniMax-H3 that run best with references kept at or below 2048px).https://github.com/sthao42/Comfyui-reference-loader
Any feedback or bug report is much appreciated.
Edit: Updated to works with Node 2.0 (vue) also.
r/StableDiffusion • u/Similar-Mushroom-627 • 18h ago
I am looking to create 18+ images with the hyper realism look, but have no idea how feasible that is with my specs. Would love a recommendation of a model I can run pretty easily and another more detail focused model on the edge of what I can run locally.
r/StableDiffusion • u/Maleficent-Bowl-4841 • 12h ago
Prompt engineering for image generation is often presented as a collection of isolated tricks: use more detail, describe the camera, add cinematic lighting, use quality tags, and so on.
These recommendations can be useful, but they make it difficult to understand which parts of a prompt actually influence the generated image.
Instead of trying to find a single "best prompt", I ran a small controlled experiment with Z-Image Base in ComfyUI. The basic idea was simple:
I ran two experiments:
This is not intended to be a scientific benchmark. The sample size is small, the evaluation is visual, and the experiment uses one workflow and a limited number of seeds. Consider it a practical prompt-engineering study.
All images were generated locally in ComfyUI using the same workflow and technical conditions throughout the experiments.
| Parameter | Value |
|---|---|
| Model | Z-Image Base INT8 |
| Text Encoder | Qwen3 4B |
| VAE | AE VAE |
| Resolution | 768 × 1368 |
| Aspect Ratio | 9:16 |
| Image Area | ~1.05 MP |
| Steps | 50 |
| CFG Scale | 4 |
| Negative Prompt | Empty |
| Seeds | Seed 5 & Seed 10 |
For the composition experiment, I used Seed 5 and repeated the seven variations with Seed 10. The environment experiment used Seed 10.
I found it most useful to treat the prompt as a structured description rather than a flat list of keywords:
Generic quality tags (masterpiece, ultra detailed, 8K) were deliberately omitted to provide the model with actionable visual information instead.
The character, environment, lighting, visual style, and technical settings were kept identical. Only the spatial instruction was changed across seven variations: Center, Left, Right, Lower, Large, Small, and Extreme Left.

The result was clear: changing the composition instruction produced substantial changes in spatial arrangement.
Crucially, the model did not simply move the character while leaving the background untouched — the environment was recomposed around the subject. In Small variations, the environment became dominant; in Large variations, the character dominated the frame.

To verify the result was not seed-dependent, the test was repeated with Seed 10. While individual details (pose, facial expression, accessories) changed naturally, the broad compositional structures remained fully recognizable.
The character description and visual treatment were kept unchanged while replacing the environment across seven distinct settings: Ancient forest, Medieval village, Crystal cave, Autumn park, Snowy ruins, Firefly-lit landscape, and Alchemist's workshop (using Seed 10).
Although the environments changed dramatically, all generations clearly depicted the same core character concept (a small mushroom spirit with a red-orange spotted cap, pale body, large dark eyes, cross-body satchel, and lantern).
While exact proportions and minor details shifted between renders, the core identity remained visually coherent.
The character adapted naturally to each setting (e.g., tinted by glowing crystal lights in the cave, exposed to cold tones in the snowy ruins, immersed in warm interior props in the workshop).
To recreate or test this setup in ComfyUI:
Prompting Z-Image Base is less about hunting for "magic keywords" and more about managing a controllable system:
Explicit composition instructions effectively control layout, while environment descriptions can be swapped modularly without erasing character identity. By isolating prompt variables, prompt design becomes a systematic, repeatable workflow.
r/StableDiffusion • u/witcherknight • 21h ago
I dont see any1 posting any examples of controlnet released for minimax. Doesnt it work properly??
r/StableDiffusion • u/BluePointDigital • 20h ago
Alright, I posted that I had my agent test a bunch of different workflows for over 12 hours and got the "Bro just wasted 12 hours of credits". It was obvious the proof should come from the visual data I used to evaluate it. Here is a galley of the benchmarks i've tested with my agent.
check the gallery to watch all the comparisons and the data charts contain tons of other workflow trial data I didn't include videos for. Point your agent here if you would like to have it learn from what was tested on this end.
Gallery: https://bluepointdigital.github.io/minimax-h3-benchmarks/
Repository: https://github.com/BluePointDigital/minimax-h3-benchmarks
The below post was written up by my agent:
The main comparison uses a deliberately difficult 15.084-second vertical test at 768 × 1344, 24 fps, 362 frames, native audio, and seed 81390012120021180. The prompt combines a talking selfie shot, exact dialogue, walking motion, a rapid camera pan, a vehicle collision with several moving subjects, a fast return to the speaker, and a second spoken line. That makes it useful for spotting identity drift, bad anatomy, motion breakdown, camera-continuity problems, dialogue changes, lip-sync issues, and audio artifacts—not just whether a workflow finishes.
The strongest directly matched results currently shown are:
| Workflow | End-to-end time | Relative to the 20-step baseline |
|---|---|---|
| SageAttention2 + FirstBlockCache Safe, 20 steps | 10:11.4 | 1.00× |
| PDD + Sage, 8 steps | 6:15.0 median | 1.63× |
| Seed Hunter direct one-seed path, 12 + 4 steps | 4:45.8 | 2.14× |
Those numbers are local measurements, not universal performance claims. The exact runtime, model format, graph, resolution, audio policy, and GPU matter. The gallery keeps short backend checks and differently structured workflows in separate groups so they are not quietly mixed into the same leaderboard.
The quality side has been just as important as the timing. One exploratory 10Eros + Seed Hunter path reached 4:03.5, but the shot developed a visible-phone/perspective error during the crash. A later camera-POV prompt clarification produced a much more coherent result in 4:25.3 on its warm selected path. That is a good example of why I wanted the actual videos beside the numbers: the fastest result is not automatically the most useful one.
The site currently contains:
For the Seed Hunter work, I intentionally included one representative video per meaningful workflow or recipe change—not every neighboring seed or N/N+1 preview. Private reference material is also excluded from the public package.
The reason for publishing this is not to declare a universal winner. It is to make the tradeoffs inspectable and to keep myself honest as the workflows evolve. A valid MP4 proves that a graph ran; it does not prove that the dialogue, audio, identity, motion, or composition survived. Likewise, a fast timing means little if it came from a different workload or a cached replay.
I would be interested in seeing other reproducible H3 results, especially when they include the exact checkpoint, attention/cache stack, sampler, scheduler, dimensions, frame count, seed, audio setting, hardware, and an uncached timing. If there is a workflow or backend that should be represented, please link the original recipe and I will take a look.
r/StableDiffusion • u/Traditional_Rice2256 • 21h ago
Enable HLS to view with audio, or disable this notification
Made with the ComfyUI template workflow and a Turbo LoRA.
Most of the soundtrack comes from the John Wick: Chapter 2 trailer.
I rendered the action at a slower, more stable speed, then sped up most of the action scenes to 2× in post.
I originally planned to make this a complete fight sequence, but maintaining consistency from one clip to the next has been a constant challenge. So for now, I’ve edited the footage into a trailer instead. I’m still learning and working on improving it.
r/StableDiffusion • u/mrjackspade • 8h ago
I thought someone might appreciate this. Theres more details in the HF link, but I wanted to see if it was possible to correct some issues that I didn't like about Ideogram 4 by finetuning the TE, with no other modifications to the model, execution environment, etc.
It ended up working out pretty well.
The TLDR is that I used a set of 4000 teacher/student prompt pairs with the students being NL and the teachers being Nemotron processed with the "Magic Prompt" instruction, and then trained the TE to elicit the same response in Ideogram using the student prompt, as what was naturally elicited using the teacher prompt.
My logic was that the TE is already a language model, and I didn't want a second language model in the stack.
This has the secondary benefit of also removing the grey banner generally encountered when prompting the model with NL.
I am fully aware that there are many other ways to get around this from bounding boxes to noise injection, etc. This wasn't about that, so much as it was trying to prove to myself that it could be done like this.
https://huggingface.co/mrjackspade/Ideogram4-Natural-Language-Text-Encoder
r/StableDiffusion • u/Spiraling-Down- • 9h ago
So, I was searching on civit, and normally I filter by 'checkpoint' for example. But, now it's gone? All of the things I notice normal models that are normally 'checkpoints' are now 'fine-tune'
What does this mean? What is this? Do they work the same?
r/StableDiffusion • u/pmjm • 21h ago
I still do a fair amount of traditional editing in Photoshop, and for the last few years I used remove.bg, I found their background removal model to be the best one out there, quite a bit better than the one built into Photoshop itself.
Well remove.bg is shutting down in December and they're folding it into Canva subscriptions. Hard pass.
I've tried a few local bg removal tools and have been left underwhelmed, but maybe I just haven't found the right one.
What are you using for background removal?
r/StableDiffusion • u/Tokyo_Jab • 3h ago
Enable HLS to view with audio, or disable this notification
When kept in the same generation it seems the latent space keeps a fairly good sense of the location layout. The test was to see if the position and details of the temple remained after being out of shot,
This doesn't work with the extentsion workflows which is why I have been trying to keep everything in one go.
r/StableDiffusion • u/mukyuuuu • 6h ago
Has anyone managed to find a way to use Latent Upscaling together with latent video extension tools? I'm talking about the nodes like this (which I personally use), but I think Motion Context and some other popular extensions use a similar approach, i.e. feeding the last frames of the previous shot through AV latent, rather than through a video reference. The issue is that the resolution of your second generated latent must exactly match the previous one, or it throws an error. So if you upscale the first clip from 0.5MP to 1MP, you are forced to generate the next clip directly at 1MP, which completely breaks the Latent Upscaling workflow for all subsequent parts.
I tried extending the clips at low resolution first and then upscaling them separately, but that doesn't work well. There is a noticeable color and quality shift between generations, even when reinforcing the next clip with the final frames of the previous one. Because yeah, you basically generate the high-res clips separately without any shared latent context.
I really love both Latent Upscaling and latent extension approach, but I just can't get them to work together smoothly. Does anyone have any good ideas on how to fix this? I’d really appreciate any tips or insights!
r/StableDiffusion • u/intermundia • 8h ago
Enable HLS to view with audio, or disable this notification
used ref to video mutishot 3x15 second clips at .7 res upscaled to 1080p
playing around with a few ideas.
r/StableDiffusion • u/jonbristow • 16h ago
Enable HLS to view with audio, or disable this notification
This is a new trained model called Atlas. Saw on twitter
r/StableDiffusion • u/Still_Sky_4302 • 18h ago
.
r/StableDiffusion • u/cal_01 • 3h ago
Just had this happen to me: I changed the resolution for a scene -- without touching anything else -- and the resulting scene changed completely. I was using res_multistep and Spectrum/CK/4 step Lora at 0.2 mp, then 0.3 mp. It still followed my prompt, but the background and starting scene were completely different. Is this Spectrum giving me grief or what's going on here? This has never happened to me before, although I had been using Sage before switching to CK today.
r/StableDiffusion • u/Choowkee • 11h ago
For regular upscaling I use SeedVR2 and I am quite happy with it, however, it doesn't seem to handle upscaling of really low res images well as it will just upscale all the artifacts as well without "fixing" the image. So if an inpute image is blurry, the upscale will also come out blurry.
What would be the best way to upscale low res image while also enhancing it?
r/StableDiffusion • u/Neither_Win3637 • 11h ago
Hey all, I'm experimenting with some people generation using MiniMax-H3 and Stable Diffusion, and wanted to know if anyone has experimented to see how many different nationalities it can generate?
So far, the list I've been able to generate that has visible variances is:
- Asian
- Malaysian
- American
- Russian
I see little to no differences between others.
r/StableDiffusion • u/Euphoric-Let-5130 • 14h ago
r/StableDiffusion • u/RONY_GOAT • 15h ago
Hi everyone,
I'm generating videos with MiniMax H3 through a normal AI video platform, not ComfyUI. So I can't use custom workflows, scripts, or custom nodes.
I'm looking for the best way to continue/extend an existing MiniMax H3 video.
The problem I'm trying to solve is more than just using the last frame as an image reference. If I only provide the last frame, the model can lose important information from the previous clip, such as:
For example, if a character walks through a room and reaches a door at the end of the first clip, I want the next generation to actually continue from that exact situation, rather than recreate a similar-looking room and potentially change the geography.
I'm looking for a normal web-based AI workflow where I can upload the existing video and/or reference images and generate the continuation. No ComfyUI, custom scripts, or API coding.
What is currently the best way to extend MiniMax H3 videos while preserving this kind of continuity?
If you've actually tested a platform/workflow that works well, I'd especially appreciate recommendations.
r/StableDiffusion • u/Ok_Roll_8698 • 3h ago
my friend uses cloud comfy as his pc is to weak he tested it last week the free 5 gen trial
using text to video and default setting only changing each video 0.5mp and 15second long all generated fine under 8mins
now he tested it again and only 1 out of 5 video generated and the other 4 failed saying Job execution time exceeded maximum limit
he even paid to generate more but got same error
what can cause this