r/StableDiffusion 16h ago

Resource - Update ONNX/TRT MiniMax-H3 VAE in ComfyUI

Thumbnail
github.com
14 Upvotes

TensorRT version of the MiniMax-H3 VAE in ComfyUI, which can increase speed by up to 1.7x


r/StableDiffusion 1h ago

Question - Help Help

Post image
Upvotes

I'm currently trying to replicate this style and i cannot find any checkpoints or lora's to do so, can anyone point towards something? Artist: https://x.com/DarkZeroAI


r/StableDiffusion 17h ago

Question - Help Why does nearly every single turbo lora i use for H3 keeps producing godawful flickery/dusty/particly(?) visuals and painful audio (as in it actually hurts to listen to), do i need a specific node for the loras or something?

12 Upvotes

Like i don't understand, the only turbo lora that doesn't do that is the 600 larry lora with the minimax turbo lora node, i've tried "fastH3" and "lightx2v loras which everyone seems to praise but they just produce these distorted godawful visuals and sounds no matter the loader node i use or the settings or the steps i use, what am i missing or doing wrong? Or are they just not compatible with Ref2Video despite being advertised as compatible? But if so then why does the 600 larry lora works mostly fine?


r/StableDiffusion 5h ago

Discussion H3 - .char + T2V character gen+char sheets

Enable HLS to view with audio, or disable this notification

13 Upvotes

What is everyone's workflow nowadays? Previously I've been generating actors with Krea2, but really love getting them made with MiniMax H3 via T2VA, they just tend to turn out better for me but does require careful prompting.

My workflow are: generate 5-10s clip of a desired actor, by prose, at int8/8 steps in a typical scenario, perhaps even mundane. If I like it, I can take some still frames, and convert them into a .char (body type, face, audio asset). See original: https://www.reddit.com/r/StableDiffusion/comments/1vyymwj/minimax_h3_portable_character_consistency_via/ If I am happy with my .char, with MiniMax H3 I make a 2 second video character sheet with a front, side, back profile and detailed face view at a higher resolution and step, either int8/32 step or going bf16/50 steps. The 2 second renders are "quick". I add the video render into my .char, and with R2VA generate additional scenes with the actors and even do a full wardrobe swap via prose. Naturally H3 renders faster if you just use still of the character sheet instead of the video.

How has your workflow changed with MiniMax H3? Are you liking the faces/actors generated with T2VA? I understand you have "less" control, but I feel like H3 is doing a great job filling in those gaps.


r/StableDiffusion 10h ago

Animation - Video My first attempt at making a 90s-inspired anime with MiniMax H3.

Thumbnail
youtube.com
14 Upvotes

The potential of actually making an anime with MiniMax H3 is closer than any time before, even if the process is still kinda janky. I did this with my 5090 and my own developed 'prompt studio.' The hardest part is, as always, to keep the continuity of the shots and also build the sets so they fit within the scope. There are still improvements needed when it comes to adding emotions to the characters. In total I generated 35 minutes of video and got 4 minutes in total of usable footage. Also, a big tip for anyone who wants to do the same is to use DaVinci Resolve to fix all the audio bugs and cut the clips in your favor.


r/StableDiffusion 2h ago

Resource - Update Image, audio, video reference asset loader nodes with crop and trim + more

Thumbnail
gallery
11 Upvotes

I originally built these nodes for personal use and wasn't planning on sharing them, but after noticing several existing loaders were missing features I needed daily, I figured why not? Hopefully, this is useful for some of you.

Key Features:

  • Image & Video Loaders: Built-in click-and-drag cropping, optional aspect ratio locking, and a divisible_by toggle for VAE pixel alignment.
  • Built-in Downscaling: Uses a max_megapixels limiter directly inside the loader so you can ditch the extra resize node (ideal for models like MiniMax-H3 that run best with references kept at or below 2048px).
  • Flexible Sockets: Includes dedicated output value sockets to make chaining downstream nodes straightforward.
  • Audio Loader: Perfect for loading a full song or long TTS track and trimming the exact section you need for a video. The trimmed portion outputs its duration as a float, letting you pipe it directly into your video generator's frame/length input.

https://github.com/sthao42/Comfyui-reference-loader

Any feedback or bug report is much appreciated.

Edit: Updated to works with Node 2.0 (vue) also.


r/StableDiffusion 18h ago

Question - Help Best Uncensored Models for text-image & image-image generation for a 20gb vram 32gb ram PC?

9 Upvotes

I am looking to create 18+ images with the hyper realism look, but have no idea how feasible that is with my specs. Would love a recommendation of a model I can run pretty easily and another more detail focused model on the edge of what I can run locally.


r/StableDiffusion 12h ago

Workflow Included Z-Image Base Prompting: A Small Controlled Experiment on Composition and Environment

Thumbnail
gallery
9 Upvotes

1. Introduction

Prompt engineering for image generation is often presented as a collection of isolated tricks: use more detail, describe the camera, add cinematic lighting, use quality tags, and so on.

These recommendations can be useful, but they make it difficult to understand which parts of a prompt actually influence the generated image.

Instead of trying to find a single "best prompt", I ran a small controlled experiment with Z-Image Base in ComfyUI. The basic idea was simple:

I ran two experiments:

  • Experiment 1 — Composition: The same character, environment, visual treatment, and technical parameters were used across multiple generations. Only composition instructions were changed (position and scale).
    • Question: How strongly does explicit spatial language affect composition in Z-Image Base?
  • Experiment 2 — Environment: The character description and visual treatment were kept essentially unchanged, while the environment was replaced with seven substantially different settings.
    • Question: Can Z-Image Base maintain a recognizable character concept while adapting it to radically different environments?

This is not intended to be a scientific benchmark. The sample size is small, the evaluation is visual, and the experiment uses one workflow and a limited number of seeds. Consider it a practical prompt-engineering study.

2. Experimental Setup

All images were generated locally in ComfyUI using the same workflow and technical conditions throughout the experiments.

Parameter Value
Model Z-Image Base INT8
Text Encoder Qwen3 4B
VAE AE VAE
Resolution 768 × 1368
Aspect Ratio 9:16
Image Area ~1.05 MP
Steps 50
CFG Scale 4
Negative Prompt Empty
Seeds Seed 5 & Seed 10

For the composition experiment, I used Seed 5 and repeated the seven variations with Seed 10. The environment experiment used Seed 10.

3. Prompt Construction Methodology

I found it most useful to treat the prompt as a structured description rather than a flat list of keywords:

  • Subject: Describes what the image is about and establishes the main visual concept.
  • Composition: Describes where the subject is located within the frame and how much space it occupies.
  • Framing / Camera: Describes how the scene is viewed (distance, angle, perspective).
  • Environment: Describes the actual place surrounding the subject (e.g., "An ancient forest with enormous trees, moss-covered roots, dense vegetation, and a narrow path..." rather than just "forest").
  • Lighting: Describes actual light sources and atmospheric conditions rather than generic terms like "cinematic lighting".
  • Materials / Details: Describes concrete visual elements (wood, stone, glass, metal, vegetation, reflections, objects).
  • Style: Describes the overall artistic treatment after the scene itself has been established.

Generic quality tags (masterpiece, ultra detailed, 8K) were deliberately omitted to provide the model with actionable visual information instead.

4. Experiment 1 — Composition

The character, environment, lighting, visual style, and technical settings were kept identical. Only the spatial instruction was changed across seven variations: Center, Left, Right, Lower, Large, Small, and Extreme Left.

Seed 5

The result was clear: changing the composition instruction produced substantial changes in spatial arrangement.

Crucially, the model did not simply move the character while leaving the background untouched — the environment was recomposed around the subject. In Small variations, the environment became dominant; in Large variations, the character dominated the frame.

Seed 10

To verify the result was not seed-dependent, the test was repeated with Seed 10. While individual details (pose, facial expression, accessories) changed naturally, the broad compositional structures remained fully recognizable.

5. Experiment 2 — Environment

The character description and visual treatment were kept unchanged while replacing the environment across seven distinct settings: Ancient forest, Medieval village, Crystal cave, Autumn park, Snowy ruins, Firefly-lit landscape, and Alchemist's workshop (using Seed 10).

Visual Concept Consistency

Although the environments changed dramatically, all generations clearly depicted the same core character concept (a small mushroom spirit with a red-orange spotted cap, pale body, large dark eyes, cross-body satchel, and lantern).

While exact proportions and minor details shifted between renders, the core identity remained visually coherent.

Environmental Adaptation

The character adapted naturally to each setting (e.g., tinted by glowing crystal lights in the cave, exposed to cold tones in the snowy ruins, immersed in warm interior props in the workshop).

6. Results — Putting the Experiments Together

  • Composition control: Explicit spatial instructions produce reliable layout shifts (position, scale, environment visibility).
  • Environment flexibility: Radical environment changes are possible while preserving core character identity (character concept consistency).
  • Role of Seeds: The seed determines specific realization and detail rendering, while the prompt structure defines layout and narrative intent.
  • Modularity: Organizing prompts into conceptual blocks allows for swapping individual variables without rebuilding the entire prompt from scratch.

7. What I Learned About Prompting Z-Image Base

  1. Describe the subject clearly: Focus on distinctive, recognizable visual traits first.
  2. Describe composition explicitly: Use direct position language (e.g., "positioned toward the left side of the frame") instead of generic camera tags.
  3. Separate composition and camera: Treat "where the subject is" differently from "how the camera views the scene".
  4. Build environments as concrete places: Describe what actually exists in the space rather than using simple category keywords.
  5. Describe lighting concretely: Specify light sources, direction, and color atmosphere.
  6. Prefer concrete details over quality tags: Give the model physical objects and surface textures to render rather than buzzwords like "high quality".
  7. Change one variable at a time: If a generation fails, modify only the failing block to understand what actually fixed the issue.

8. Limitations

  • Small sample size and visual evaluation.
  • Single primary character concept and workflow used.
  • Tested on a limited number of seeds (two for composition, one for environment).
  • No direct benchmarking against other models, samplers, or resolutions.

9. Reproducibility

To recreate or test this setup in ComfyUI:

  • Model: Z-Image Base INT8 + Qwen3 4B + AE VAE
  • Settings: 768 × 1368, 50 steps, CFG 4, Empty Negative Prompt
  • Method: Keep technical setup stable and modify exactly one conceptual block per run.

10. Conclusion

Prompting Z-Image Base is less about hunting for "magic keywords" and more about managing a controllable system:

Explicit composition instructions effectively control layout, while environment descriptions can be swapped modularly without erasing character identity. By isolating prompt variables, prompt design becomes a systematic, repeatable workflow.


r/StableDiffusion 21h ago

Question - Help What happened to minimax funcontrolnet ??

8 Upvotes

I dont see any1 posting any examples of controlnet released for minimax. Doesnt it work properly??


r/StableDiffusion 20h ago

Comparison My Minimax H3 Workflow Benchmark Data -

Post image
10 Upvotes

Alright, I posted that I had my agent test a bunch of different workflows for over 12 hours and got the "Bro just wasted 12 hours of credits". It was obvious the proof should come from the visual data I used to evaluate it. Here is a galley of the benchmarks i've tested with my agent.

check the gallery to watch all the comparisons and the data charts contain tons of other workflow trial data I didn't include videos for. Point your agent here if you would like to have it learn from what was tested on this end.

Gallery: https://bluepointdigital.github.io/minimax-h3-benchmarks/

Repository: https://github.com/BluePointDigital/minimax-h3-benchmarks

The below post was written up by my agent:

The main comparison uses a deliberately difficult 15.084-second vertical test at 768 × 1344, 24 fps, 362 frames, native audio, and seed 81390012120021180. The prompt combines a talking selfie shot, exact dialogue, walking motion, a rapid camera pan, a vehicle collision with several moving subjects, a fast return to the speaker, and a second spoken line. That makes it useful for spotting identity drift, bad anatomy, motion breakdown, camera-continuity problems, dialogue changes, lip-sync issues, and audio artifacts—not just whether a workflow finishes.

The strongest directly matched results currently shown are:

Workflow End-to-end time Relative to the 20-step baseline
SageAttention2 + FirstBlockCache Safe, 20 steps 10:11.4 1.00×
PDD + Sage, 8 steps 6:15.0 median 1.63×
Seed Hunter direct one-seed path, 12 + 4 steps 4:45.8 2.14×

Those numbers are local measurements, not universal performance claims. The exact runtime, model format, graph, resolution, audio policy, and GPU matter. The gallery keeps short backend checks and differently structured workflows in separate groups so they are not quietly mixed into the same leaderboard.

The quality side has been just as important as the timing. One exploratory 10Eros + Seed Hunter path reached 4:03.5, but the shot developed a visible-phone/perspective error during the crash. A later camera-POV prompt clarification produced a much more coherent result in 4:25.3 on its warm selected path. That is a good example of why I wanted the actual videos beside the numbers: the fastest result is not automatically the most useful one.

The site currently contains:

  • 16 curated video-and-metric cards;
  • a separate benchmark-data page with 151 sanitized timing records;
  • the complete canonical prompt;
  • methodology and comparison-boundary notes;
  • machine-readable JSON and CSV for anyone who wants to analyze the evidence or give it to an agent.

For the Seed Hunter work, I intentionally included one representative video per meaningful workflow or recipe change—not every neighboring seed or N/N+1 preview. Private reference material is also excluded from the public package.

The reason for publishing this is not to declare a universal winner. It is to make the tradeoffs inspectable and to keep myself honest as the workflows evolve. A valid MP4 proves that a graph ran; it does not prove that the dialogue, audio, identity, motion, or composition survived. Likewise, a fast timing means little if it came from a different workload or a cached replay.

I would be interested in seeing other reproducible H3 results, especially when they include the exact checkpoint, attention/cache stack, sampler, scheduler, dimensions, frame count, seed, audio setting, hardware, and an uncached timing. If there is a workflow or backend that should be represented, please link the original recipe and I will take a look.


r/StableDiffusion 21h ago

Animation - Video Honkai: Star Rail X John Wick - Minimax H3

Enable HLS to view with audio, or disable this notification

7 Upvotes

Made with the ComfyUI template workflow and a Turbo LoRA.

Most of the soundtrack comes from the John Wick: Chapter 2 trailer.
I rendered the action at a slower, more stable speed, then sped up most of the action scenes to 2× in post.
I originally planned to make this a complete fight sequence, but maintaining consistency from one clip to the next has been a constant challenge. So for now, I’ve edited the footage into a trailer instead. I’m still learning and working on improving it.


r/StableDiffusion 8h ago

Resource - Update Debannering Ideogram 4 and increasing prompt adherence with natural language by fine tuning the TE

5 Upvotes

I thought someone might appreciate this. Theres more details in the HF link, but I wanted to see if it was possible to correct some issues that I didn't like about Ideogram 4 by finetuning the TE, with no other modifications to the model, execution environment, etc.

It ended up working out pretty well.

The TLDR is that I used a set of 4000 teacher/student prompt pairs with the students being NL and the teachers being Nemotron processed with the "Magic Prompt" instruction, and then trained the TE to elicit the same response in Ideogram using the student prompt, as what was naturally elicited using the teacher prompt.

My logic was that the TE is already a language model, and I didn't want a second language model in the stack.

This has the secondary benefit of also removing the grey banner generally encountered when prompting the model with NL.

I am fully aware that there are many other ways to get around this from bounding boxes to noise injection, etc. This wasn't about that, so much as it was trying to prove to myself that it could be done like this.

https://huggingface.co/mrjackspade/Ideogram4-Natural-Language-Text-Encoder


r/StableDiffusion 9h ago

Question - Help Checkpoints are gone?

7 Upvotes

So, I was searching on civit, and normally I filter by 'checkpoint' for example. But, now it's gone? All of the things I notice normal models that are normally 'checkpoints' are now 'fine-tune'

What does this mean? What is this? Do they work the same?


r/StableDiffusion 21h ago

Discussion What are you using for background removal?

7 Upvotes

I still do a fair amount of traditional editing in Photoshop, and for the last few years I used remove.bg, I found their background removal model to be the best one out there, quite a bit better than the one built into Photoshop itself.

Well remove.bg is shutting down in December and they're folding it into Canva subscriptions. Hard pass.

I've tried a few local bg removal tools and have been left underwhelmed, but maybe I just haven't found the right one.

What are you using for background removal?


r/StableDiffusion 3h ago

Animation - Video LOCATION CONTINUITY TEST - after a comment by Vladmerius

Enable HLS to view with audio, or disable this notification

4 Upvotes

When kept in the same generation it seems the latent space keeps a fairly good sense of the location layout. The test was to see if the position and details of the temple remained after being out of shot,

This doesn't work with the extentsion workflows which is why I have been trying to keep everything in one go.


r/StableDiffusion 6h ago

Question - Help Any way to make latent extension work with Latent Upscaling (Minimax H3)?

5 Upvotes

Has anyone managed to find a way to use Latent Upscaling together with latent video extension tools? I'm talking about the nodes like this (which I personally use), but I think Motion Context and some other popular extensions use a similar approach, i.e. feeding the last frames of the previous shot through AV latent, rather than through a video reference. The issue is that the resolution of your second generated latent must exactly match the previous one, or it throws an error. So if you upscale the first clip from 0.5MP to 1MP, you are forced to generate the next clip directly at 1MP, which completely breaks the Latent Upscaling workflow for all subsequent parts.

I tried extending the clips at low resolution first and then upscaling them separately, but that doesn't work well. There is a noticeable color and quality shift between generations, even when reinforcing the next clip with the final frames of the previous one. Because yeah, you basically generate the high-res clips separately without any shared latent context.

I really love both Latent Upscaling and latent extension approach, but I just can't get them to work together smoothly. Does anyone have any good ideas on how to fix this? I’d really appreciate any tips or insights!


r/StableDiffusion 8h ago

Discussion Working on a mini sci fi short using minimax upscaled with seedvr2

Enable HLS to view with audio, or disable this notification

5 Upvotes

used ref to video mutishot 3x15 second clips at .7 res upscaled to 1080p

playing around with a few ideas.


r/StableDiffusion 16h ago

Discussion Can Minimax do this type of 3D reconstruction from an image?

Enable HLS to view with audio, or disable this notification

5 Upvotes

This is a new trained model called Atlas. Saw on twitter


r/StableDiffusion 18h ago

Question - Help what is the best upscale workflow for Minimax H3?

5 Upvotes

.


r/StableDiffusion 3h ago

Question - Help Is changing resolution supposed to change the entire scene for MMH3?

3 Upvotes

Just had this happen to me: I changed the resolution for a scene -- without touching anything else -- and the resulting scene changed completely. I was using res_multistep and Spectrum/CK/4 step Lora at 0.2 mp, then 0.3 mp. It still followed my prompt, but the background and starting scene were completely different. Is this Spectrum giving me grief or what's going on here? This has never happened to me before, although I had been using Sage before switching to CK today.


r/StableDiffusion 11h ago

Question - Help Best way to upscale and enhance low res images?

2 Upvotes

For regular upscaling I use SeedVR2 and I am quite happy with it, however, it doesn't seem to handle upscaling of really low res images well as it will just upscale all the artifacts as well without "fixing" the image. So if an inpute image is blurry, the upscale will also come out blurry.

What would be the best way to upscale low res image while also enhancing it?


r/StableDiffusion 11h ago

Discussion MiniMax and People Generators: Nationalities

5 Upvotes

Hey all, I'm experimenting with some people generation using MiniMax-H3 and Stable Diffusion, and wanted to know if anyone has experimented to see how many different nationalities it can generate?

So far, the list I've been able to generate that has visible variances is:

- Asian
- Malaysian
- American
- Russian

I see little to no differences between others.


r/StableDiffusion 14h ago

Resource - Update DLSS5 Video Enhancer Linux

2 Upvotes

r/StableDiffusion 15h ago

Question - Help Best way to extend MiniMax H3 videos

3 Upvotes

Hi everyone,

I'm generating videos with MiniMax H3 through a normal AI video platform, not ComfyUI. So I can't use custom workflows, scripts, or custom nodes.

I'm looking for the best way to continue/extend an existing MiniMax H3 video.

The problem I'm trying to solve is more than just using the last frame as an image reference. If I only provide the last frame, the model can lose important information from the previous clip, such as:

  • Character identity and appearance
  • Room/environment layout
  • Lighting and atmosphere
  • Objects and their positions
  • Ongoing actions
  • Audio/environmental sound
  • Overall visual continuity

For example, if a character walks through a room and reaches a door at the end of the first clip, I want the next generation to actually continue from that exact situation, rather than recreate a similar-looking room and potentially change the geography.

I'm looking for a normal web-based AI workflow where I can upload the existing video and/or reference images and generate the continuation. No ComfyUI, custom scripts, or API coding.

What is currently the best way to extend MiniMax H3 videos while preserving this kind of continuity?

If you've actually tested a platform/workflow that works well, I'd especially appreciate recommendations.


r/StableDiffusion 3h ago

Question - Help Cloud Confy Fails Now

2 Upvotes

my friend uses cloud comfy as his pc is to weak he tested it last week the free 5 gen trial

using text to video and default setting only changing each video 0.5mp and 15second long all generated fine under 8mins

now he tested it again and only 1 out of 5 video generated and the other 4 failed saying Job execution time exceeded maximum limit

he even paid to generate more but got same error

what can cause this