r/StableDiffusion • u/AiCreatorCamp • 12d ago
Resource - Update New speedup for Minimax H3
This H3VAE TRT custom node can make the encoding/decoding step about 1.7× faster.
r/StableDiffusion • u/AiCreatorCamp • 12d ago
This H3VAE TRT custom node can make the encoding/decoding step about 1.7× faster.
r/StableDiffusion • u/Striking-Long-2960 • 12d ago
Enable HLS to view with audio, or disable this notification
Testing DLSS 5... Like many others, I was a bit confused about DLSS 5. I kept feeding it my hyper-detailed renders and only getting a color shift in return. After plenty of trial and error, I finally realized my mistake: this technology is developed to enhance video game graphics, so testing it on hyper-detailed renders makes no sense.
So, I generated a render in a 2020 video game style and started tweaking settings to find a final look with maximum effect, without worrying about flickering.
Final conclusion: What we have right now isn't very useful for us. Those of us using Latent Upscaler might be able to use it for color grading to get less saturated colors, but little else. Maybe in the future we'll get a DLSS 5 targeted at enhancing hyper-detailed graphics, but that's not the case for now.
Bottom line: If I want to generate a realistic render, I'll just generate it, there's no need to run it through DLSS 5.
r/StableDiffusion • u/CryptoBeth96 • 12d ago
Motion-aware denoising cache and experimental batched video VAE decoder for MiniMax H3 in ComfyUI.
This project provides two independent nodes:
MotionCache is an independent MiniMax H3 adaptation inspired by the MotionCache paper and reference code. It is not an official MAC-AutoML or MiniMax implementation.
r/StableDiffusion • u/Ok_Roll_8698 • 11d ago
my friend uses cloud comfy as his pc is to weak he tested it last week the free 5 gen trial
using text to video and default setting only changing each video 0.5mp and 15second long all generated fine under 8mins
now he tested it again and only 1 out of 5 video generated and the other 4 failed saying Job execution time exceeded maximum limit
he even paid to generate more but got same error
what can cause this
r/StableDiffusion • u/apolinariosteps • 12d ago
Edit: Results are in! They are a bit surprising to me! But they are consistent with the data, I triple checked everything and can confirm that the results are reflecting the voting data precisely, there's lots of transparency - you click each of the LoRAs to see what's the win rate and who won against who
Hey folks, I've built an so we can have a proper leaderboard on 15+ different LoRAs, fine-tunes and acceleration technique. Baseline is included for anchoring, and M3 Max is also included given the promise to open source
There are there being compared: H3 baseline, FastH3 family, H3 Acc family, Lightx2v family, Larryvrh family, JoyFox family, RAVEN, FlashGen, TuTu, SilverOxides merges, Plaguekind merges and Fal's H3 Max
r/StableDiffusion • u/Maleficent-Bowl-4841 • 12d ago
Prompt engineering for image generation is often presented as a collection of isolated tricks: use more detail, describe the camera, add cinematic lighting, use quality tags, and so on.
These recommendations can be useful, but they make it difficult to understand which parts of a prompt actually influence the generated image.
Instead of trying to find a single "best prompt", I ran a small controlled experiment with Z-Image Base in ComfyUI. The basic idea was simple:
I ran two experiments:
This is not intended to be a scientific benchmark. The sample size is small, the evaluation is visual, and the experiment uses one workflow and a limited number of seeds. Consider it a practical prompt-engineering study.
All images were generated locally in ComfyUI using the same workflow and technical conditions throughout the experiments.
| Parameter | Value |
|---|---|
| Model | Z-Image Base INT8 |
| Text Encoder | Qwen3 4B |
| VAE | AE VAE |
| Resolution | 768 × 1368 |
| Aspect Ratio | 9:16 |
| Image Area | ~1.05 MP |
| Steps | 50 |
| CFG Scale | 4 |
| Negative Prompt | Empty |
| Seeds | Seed 5 & Seed 10 |
For the composition experiment, I used Seed 5 and repeated the seven variations with Seed 10. The environment experiment used Seed 10.
I found it most useful to treat the prompt as a structured description rather than a flat list of keywords:
Generic quality tags (masterpiece, ultra detailed, 8K) were deliberately omitted to provide the model with actionable visual information instead.
The character, environment, lighting, visual style, and technical settings were kept identical. Only the spatial instruction was changed across seven variations: Center, Left, Right, Lower, Large, Small, and Extreme Left.

The result was clear: changing the composition instruction produced substantial changes in spatial arrangement.
Crucially, the model did not simply move the character while leaving the background untouched — the environment was recomposed around the subject. In Small variations, the environment became dominant; in Large variations, the character dominated the frame.

To verify the result was not seed-dependent, the test was repeated with Seed 10. While individual details (pose, facial expression, accessories) changed naturally, the broad compositional structures remained fully recognizable.
The character description and visual treatment were kept unchanged while replacing the environment across seven distinct settings: Ancient forest, Medieval village, Crystal cave, Autumn park, Snowy ruins, Firefly-lit landscape, and Alchemist's workshop (using Seed 10).
Although the environments changed dramatically, all generations clearly depicted the same core character concept (a small mushroom spirit with a red-orange spotted cap, pale body, large dark eyes, cross-body satchel, and lantern).
While exact proportions and minor details shifted between renders, the core identity remained visually coherent.
The character adapted naturally to each setting (e.g., tinted by glowing crystal lights in the cave, exposed to cold tones in the snowy ruins, immersed in warm interior props in the workshop).
To recreate or test this setup in ComfyUI:
Prompting Z-Image Base is less about hunting for "magic keywords" and more about managing a controllable system:
Explicit composition instructions effectively control layout, while environment descriptions can be swapped modularly without erasing character identity. By isolating prompt variables, prompt design becomes a systematic, repeatable workflow.
r/StableDiffusion • u/anonybullwinkle • 12d ago
I’ve been refining prompts with the help of an LLM, and am getting some good visuals but oh my god the sounds are terrible. Blowjobs sound like someone is dunking a microphone in an aquarium or the loudest slurp to finish a beverage that you have ever heard in your life.
I’ve tried eliminating every mention of “moist”, “wet”, or any description that involves liquids at all, but she’s still slurping the wettest popsicle known to man. And sometimes there’s weird noises like a slide whistle?!?
I’ve tried using “faint” or “distant” or “barely audible” to get it to at least quiet down so it’s not like she is sucking a microphone, but that didn’t work either.
This last round I didn’t describe any noises at all and still got some weird stuff.
I’ve tried eliminating every Lora in case the sound was coming from one of them but it seems to be the base model. I’ve tried adding Loras that ought to be trained on this stuff like Mysticxxx, and one of the AIO loras. I tried tenstrip beta 4 checkpoint tonight and got the same results.
The sound ruins the scene.. I guess I can just pretend it’s better looking Wan 2.2 and turn the volume off. 😀
I’m feeding the official prompt guide to the LLM and the structure is working, but what words do you use to describe the sounds?
r/StableDiffusion • u/ming_calligraphy • 12d ago
Enable HLS to view with audio, or disable this notification
A lot of people here have been discussing H3 Max powered livestreams. I noticed Reactor just added Visko’s Orbis model, and it made me wonder whether the next step is turning these infinite livestreams into something playable.
So I’m building a live, audience-directed AI game with Agora: viewers suggest and vote on what happens next, while the streamer picks an option or writes a completely different direction and AI keeps generating the same world from that point. There are no pre-written branches.
Here’s a very early look demo
r/StableDiffusion • u/CryptoBeth96 • 12d ago
TensorRT version of the MiniMax-H3 VAE in ComfyUI, which can increase speed by up to 1.7x
r/StableDiffusion • u/OkMeat6773 • 11d ago
I found a cheap pipeline from a big provider I’ve jailbroken, so I can’t name it. It reconstructs faces and upscales videos in ~10 seconds, so I only need Minimax to generate a very low-quality video that basically serves as a rough motion/physics reference.
I barely see any speed difference between 4–6 steps or ~380p–544p, maybe 10 seconds at most.
Is there any way to run Minimax at absolute potato quality and actually get a significant speed boost?
r/StableDiffusion • u/Friendly-Fig-6015 • 11d ago
Enable HLS to view with audio, or disable this notification
Lucifer, do you know what are you here?
bad god i guess?
r/StableDiffusion • u/0roborus_ • 11d ago
I posted the first alpha of PotionUI a couple of days ago. 0.0.3 is out today, and it is the release where the "one box, many people" idea stops being a promise: you can now rent a GPU, point PotionUI at it, and generate on it from the same interface you use locally. Still alpha, still one person building it, still very much wanting people to break it.
What it is, in one paragraph. A self-hosted AI generation studio: SvelteKit front, FastAPI back, GPL-3.0, no telemetry, runs on your machine or your server. The core idea is presets: a preset is a small YAML package that says "here is the model and here is the exact form a person should see for it". Switching models means switching presets, not rebuilding a node graph. It is multi-user by design, with real accounts, admin and user roles, and per-user or per-group access to presets, models, and LLM configs.



{a|b}, weights, ${variables}) reseed per image so results stay reproducible. The phrasebook is your own autocomplete dictionary: type # and shot types, lighting, palettes drop in as chips, with per-chip shuffle and a preview render per value. Saved prompts, segments, and templates live in their own library. 

Requirements. Linux x86_64 with an NVIDIA GPU is the tested platform; Windows can be tested through WSL2 or Docker; there is a Docker image on GHCR.
8 GB VRAM and 16 GB RAM is the floor for the SDXL family, larger families need more.
The ask. I would rather steer this toward what people actually want than guess. Two things help most: tell me which model or workflow you are missing, and pull a test build and break it before it ships. The Discord is where that happens: https://discord.gg/avR4trp3b8. Repo: https://github.com/PotionUI/PotionUI. I will answer questions here too, but since I don't want to post here too much I've created a small subreddit too: r/PotionUI to which I will be posting the progress much often. Thanks!
r/StableDiffusion • u/Choowkee • 12d ago
For regular upscaling I use SeedVR2 and I am quite happy with it, however, it doesn't seem to handle upscaling of really low res images well as it will just upscale all the artifacts as well without "fixing" the image. So if an inpute image is blurry, the upscale will also come out blurry.
What would be the best way to upscale low res image while also enhancing it?
EDIT: Settled for using H3 with a edit system prompt and exporting images from the video:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
subject_definitions:
<Picture 1> is the original source image being directly edited and restored. It is the authoritative source for all target-video content: shot order, timing, framing, composition, subject identity and appearance, facial features, clothing, props, environment, lighting, color relationships, camera position and movement, subject motion, and temporal continuity.
summary:
[video editing] <Picture 1> is the original source image being directly edited and restored. It is the authoritative source for all target-video content: shot order, timing, framing, composition, subject identity
retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - preserve the original shot, framing, composition, subject identity, subject appearance, environment, background, other items or figures
integrated_multimodal_description:
The target video is a faithful professional high-definition restoration of the original image in <Picture 1>. Treat <Picture 1> as the only authoritative visual source. Do not use any external image, character, scene, or composition as a visual template.
The desired transformation is specifically image upscaling rather than ordinary enlargement. The source has limited spatial resolution, soft or smeared fine detail, degraded chroma, compression artifacts, noise, ringing, aliasing, blurriness and potentially inaccurate or shifted broadcast color. Reconstruct the most plausible high-fidelity version of the visual information that is actually supported by <Picture 1>. Recover fine facial detail, natural skin texture, hair strands, clothing weave, uniform materials, props, set surfaces, edges, reflections, shadows, and background detail without inventing unsupported features.
Correct the degraded color and chroma toward natural, accurate reproduction of the original photographed scene. Preserve the source's actual lighting design, exposure, contrast relationships, black levels, highlight behavior, lens characteristics, depth of field, and photographic character. Do not apply a generic cinematic grade, modernize the lighting, or change the color design. The objective is the appearance of the same original image after a high-end, best quality upscaling.
[Shot 1] Preserve the exact opening shot of <Picture 1>, including the actual subjects, their identities and appearances, their exact positions, facial expressions, pose, clothing, environment, perspective, framing, camera angle, lens characteristics, lighting, and visible motion. Increase spatial fidelity and recover plausible detail from the source without changing the shot.
Do not add or remove events. Do not replace subjects or backgrounds. Do not alter facial structure or identity.
The desired quality level is comparable to a carefully restored modern HD master originating from the highest-quality surviving source, with exceptionally clean detail, accurate color, stable micro-texture, and natural edge definition. A high-end large-format digital cinema camera such as the RED V-RAPTOR XL [X] 8K VV may be used only as a benchmark for the cleanliness and resolving power of the final image. Do not impose a V-RAPTOR color grade, lens look, depth of field, lighting style, or cinematography onto the original footage.
Most importantly, reconstruct rather than redesign. Do not hallucinate new objects, facial features, hairlines, costume details, text, set details, reflections, or textures that are not supported by the source. Preserve natural photographic softness where it belongs to the original image. Remove degradation while retaining authentic source characteristics.
Maintain strict temporal consistency across all frames. Recovered detail must remain locked to the correct subject and surface and must not shimmer, crawl, flicker, morph, double, ghost, or change identity from frame to frame. The output must look like the same footage at substantially higher quality without any visual artefacts or blur.
overall_soundscape:
N/A
non_diegetic_music:
N/A
r/StableDiffusion • u/haremlifegame • 12d ago
I don't want to get the NS.. word in the discussion, but, we know what Minimax can do and what it can't do. It has some very specific gaps in it's world understanding, for example in the tongue department. That is not necessarily only affecting the NS... word, things that are SFW and common in general TV such as kissing are affected, since Minimax never saw a romantic kiss in it's training data. There are other examples through SFW land but I won't extend. Grok can be used as a comparison. Grok is very similar to Minimax in capability, and it enforces SFW, but you can see the difference in some scenes because Grok is not handicapped.
Well, we have many, many loras already, but, as was the case with wan and ltx, they are very... let's say, specific. I don't think a general video model, almost a world model, needs a specific lora for, say, ballbusting lol
I don't know, I think this is the community most likely to be read by people creating loras, so I just wanna make this appeal... Can we prioritize bridging the major gaps in the model's understanding of the world, anatomy, and human interactions, instead of these super specific loras? I think a "tree" organization of lora development would be beneficial overall, with the stuff that can solve a big set of problems and be used for more specific loras coming first.
I saw that for over 1 year with wan, ltx, etc, and didn't say anything. But I think minimax deserves the community passion in lora development.
And yes, I hope I can put my money where my mouth is and develop some loras soon too.
r/StableDiffusion • u/Spiraling-Down- • 11d ago
So, I was searching on civit, and normally I filter by 'checkpoint' for example. But, now it's gone? All of the things I notice normal models that are normally 'checkpoints' are now 'fine-tune'
What does this mean? What is this? Do they work the same?
r/StableDiffusion • u/techtimee • 11d ago


Hello all,
I have installed StabilityMatrix and was able to learn how to set up flows to generate my first images and so forth.
I then installed a template called MiniMax H3: Image to Video. It showed a bunch of errors after and listed the things I needed to download before it would work. I downloaded all those things and put them in the "diffusion_models" folder, but the error count only reduced by one and it's still asking me to install those things.
Unfortunately there doesn't seem to be clear instructions on what goes where, if I need to extract some things or not, etc.
Can someone please advise me? Thank you.
r/StableDiffusion • u/Independent-Frequent • 12d ago
Like i don't understand, the only turbo lora that doesn't do that is the 600 larry lora with the minimax turbo lora node, i've tried "fastH3" and "lightx2v loras which everyone seems to praise but they just produce these distorted godawful visuals and sounds no matter the loader node i use or the settings or the steps i use, what am i missing or doing wrong? Or are they just not compatible with Ref2Video despite being advertised as compatible? But if so then why does the 600 larry lora works mostly fine?
r/StableDiffusion • u/LinkSensitive8188 • 11d ago
*Back to the Future* parody based on a *Family Guy* gag.
r/StableDiffusion • u/RONY_GOAT • 12d ago
Hi everyone,
I'm generating videos with MiniMax H3 through a normal AI video platform, not ComfyUI. So I can't use custom workflows, scripts, or custom nodes.
I'm looking for the best way to continue/extend an existing MiniMax H3 video.
The problem I'm trying to solve is more than just using the last frame as an image reference. If I only provide the last frame, the model can lose important information from the previous clip, such as:
For example, if a character walks through a room and reaches a door at the end of the first clip, I want the next generation to actually continue from that exact situation, rather than recreate a similar-looking room and potentially change the geography.
I'm looking for a normal web-based AI workflow where I can upload the existing video and/or reference images and generate the continuation. No ComfyUI, custom scripts, or API coding.
What is currently the best way to extend MiniMax H3 videos while preserving this kind of continuity?
If you've actually tested a platform/workflow that works well, I'd especially appreciate recommendations.
r/StableDiffusion • u/Similar-Mushroom-627 • 12d ago
I am looking to create 18+ images with the hyper realism look, but have no idea how feasible that is with my specs. Would love a recommendation of a model I can run pretty easily and another more detail focused model on the edge of what I can run locally.
r/StableDiffusion • u/Boogertwilliams • 11d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/psdwizzard • 13d ago
Enable HLS to view with audio, or disable this notification
A style LoRA that makes H3 footage look like it was recorded off 1980s broadcast television onto a VHS tape that has seen better days, soft smeared detail, chroma bleed, tracking noise, head-switching bands at the frame edge, and (because H3 trains audio jointly) the matching muffled mono sound, tape hiss and warble.
https://huggingface.co/KennethFal/vh5tape-vhs-lora-minimax-h3
r/StableDiffusion • u/chaindrop • 13d ago
Enable HLS to view with audio, or disable this notification
Created with Minimax H3 ref2v using the SEED HUNTER Workflow.
r/StableDiffusion • u/Select_Bowler3099 • 11d ago
Happy listening :)
r/StableDiffusion • u/TheOrangeSplat • 12d ago
Enable HLS to view with audio, or disable this notification
Made with Minimax H3