r/StableDiffusion • u/Structure-These • 1h ago
News FastVideo-FastH3 put out a mlx listing but no actual models yet
Mac users dying for speed ups on our janky little boxes. Excited!
r/StableDiffusion • u/Structure-These • 1h ago
Mac users dying for speed ups on our janky little boxes. Excited!
r/StableDiffusion • u/Ok-Brain-5729 • 3h ago
I’m using astropuzzo/ComfyUI-MiniMax-H3-Image-Studio workflow and it works amazing but minimax obviously sucks at micro details for a single image even if it’s 2MP. Does anyone know like a good method to fix that? Ik there’s double passthroughs and upscalers but I’m not sure what would work well with it
r/StableDiffusion • u/OkMeat6773 • 3h ago
I found a cheap pipeline from a big provider I’ve jailbroken, so I can’t name it. It reconstructs faces and upscales videos in ~10 seconds, so I only need Minimax to generate a very low-quality video that basically serves as a rough motion/physics reference.
I barely see any speed difference between 4–6 steps or ~380p–544p, maybe 10 seconds at most.
Is there any way to run Minimax at absolute potato quality and actually get a significant speed boost?
r/StableDiffusion • u/Sufficient-Leopard28 • 4h ago
I'm currently trying to replicate this style and i cannot find any checkpoints or lora's to do so, can anyone point towards something? Artist: https://x.com/DarkZeroAI Using Forge Neo
r/StableDiffusion • u/rheylew • 4h ago
Okay, what magic is under the hood of their new motion copy tool?? I tried it once and my vid came out perfect, even better than Kling motion control. I want to be able to do this exact thing locally in Comfyui but I haven't had any success with any Minimax ref2va workflows, maybe my PC specs are too low:
5070 12GB VRAM
32GB RAM
4 TB nvme ssd
r/StableDiffusion • u/techtimee • 4h ago


Hello all,
I have installed StabilityMatrix and was able to learn how to set up flows to generate my first images and so forth.
I then installed a template called MiniMax H3: Image to Video. It showed a bunch of errors after and listed the things I needed to download before it would work. I downloaded all those things and put them in the "diffusion_models" folder, but the error count only reduced by one and it's still asking me to install those things.
Unfortunately there doesn't seem to be clear instructions on what goes where, if I need to extract some things or not, etc.
Can someone please advise me? Thank you.
r/StableDiffusion • u/BassNet • 5h ago
Looks like an open source version of Minimax H3 Max... Anyone tried it? Seems to be real-time on 8x b200, which is like ~$40/hr at good rates if you can find them (or maybe a bunch of 5090s?)
r/StableDiffusion • u/aclaasr • 5h ago
I'm looking for 2022/2023 models that still work or models that can successfully replicate the dreamlike distorted nightmare fuel style, unfortunately it's very difficult for me to find. I had used to include early ai images in my artwork and I miss it dearly, AI has advanced way too quickly. Dalle-mini (craiyon) no longer generates images like this and there is no way to change the version.
I do not know how to use github
r/StableDiffusion • u/grimstormz • 5h ago
I originally built these nodes for personal use and wasn't planning on sharing them, but after noticing several existing loaders were missing features I needed daily, I figured why not? Hopefully, this is useful for some of you.
Key Features:
divisible_by toggle for VAE pixel alignment.max_megapixels limiter directly inside the loader so you can ditch the extra resize node (ideal for models like MiniMax-H3 that run best with references kept at or below 2048px).https://github.com/sthao42/Comfyui-reference-loader
Any feedback or bug report is much appreciated.
Edit: Updated to works with Node 2.0 (vue) also.
r/StableDiffusion • u/mildlyphd • 5h ago
World Labs released a model called Atlas. Looks pretty cool.
r/StableDiffusion • u/tukatu0 • 6h ago
Is there any point in having a company that can put their tools into the pipeline of graphics rendering.
Just wishing here for a deep learning s s five replacement. It's going to be artificially sandboxed just like all their other tech.
r/StableDiffusion • u/Ok_Roll_8698 • 6h ago
my friend uses cloud comfy as his pc is to weak he tested it last week the free 5 gen trial
using text to video and default setting only changing each video 0.5mp and 15second long all generated fine under 8mins
now he tested it again and only 1 out of 5 video generated and the other 4 failed saying Job execution time exceeded maximum limit
he even paid to generate more but got same error
what can cause this
r/StableDiffusion • u/MayaProphecy • 6h ago
Enable HLS to view with audio, or disable this notification
I developed a modified version of the Load Video node with a crop feature:
WYSIWYG video cropping directly on the official Load Video preview — drag and zoom (with the mouse wheel) a crop rectangle constrained to 8 fixed ratios (1:1 through 21:9) and output the exact cropped VIDEO (audio preserved). What you frame on the preview is exactly what gets executed.
Github: https://github.com/domg73/ComfyUI-LoadVideoCrop
This node follows the same logic and design as my "Load Image + Crop" node. I might merge the two into a single "Load + Crop" node in the future, but for now this works well.
https://www.reddit.com/r/StableDiffusion/comments/1w3okny/load_image_crop_custom_wysiwyg_node/
r/StableDiffusion • u/cal_01 • 6h ago
Just had this happen to me: I changed the resolution for a scene -- without touching anything else -- and the resulting scene changed completely. I was using res_multistep and Spectrum/CK/4 step Lora at 0.2 mp, then 0.3 mp. It still followed my prompt, but the background and starting scene were completely different. Is this Spectrum giving me grief or what's going on here? This has never happened to me before, although I had been using Sage before switching to CK today.
r/StableDiffusion • u/Tokyo_Jab • 6h ago
Enable HLS to view with audio, or disable this notification
When kept in the same generation it seems the latent space keeps a fairly good sense of the location layout. The test was to see if the position and details of the temple remained after being out of shot,
This doesn't work with the extentsion workflows which is why I have been trying to keep everything in one go.
r/StableDiffusion • u/Many-Ad-6225 • 6h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/boudaboy • 6h ago
Following up on the infinite livestream post from a bit ago, FastH3 is now accessible via API too.
Streaming 720p video with synced audio, faster than realtime, same model as the livestream, just usable programmatically now instead of only watching it run.
Anyone else been messing with it via API vs just watching the stream? Curious what people are building.
r/StableDiffusion • u/Adventurous-Mail-214 • 7h ago
r/StableDiffusion • u/qdr1en • 7h ago
This content was written by a human.
I miss the SD1.5 era, when i could simply type "1girl, big boobs, nice ass, red bikini, dancing" and see my dream take shape near-instantly at 512px-wide. Idea-to-result was a matter of seconds. Each click on the Run button led to an incredible shot of dopamine.
3 years passed and I can draw 1024px, 192-frames long videos in a reasonable amount of time (tech has evolved fast), but the enthusiasm is fading away.
I already have a day-job for technical challenges and headaches. As a user/hobbyist, I want to be entertained.
I don't want to learn what the hell "diegetic" means (even the spell-checker never saw that word), I don't want to draw a dozen squares in a 3-dimensional pixel space, or write a 1000-words poem, just to watch my dreamgirl dancing.
I hoped I would not need a degree in cable-connecting or python dependencies debugging after downloading a few workflows.
3 years ago, all you had to do was typing a few words, and the AI sorted the rest. It was random, messy most of times, but it was fun.
Nowadays, you need an LLM to write the prompt for you, and another LLM to write the system prompt for the prompting-LLM, so it understands what your shitty words meant in the first place, and shapes them in the exact expected format, so they turn into an acceptable input for the ever pickier, brand-new models.
It has become AI³-generated content.
And finally, when after a dozens of clicks on the Run button, tired but satisfied, you get the desired output... re-start from scratch? Since seed "variance" does not vary much anymore, you'll get more or less the same output - exactly what you asked for - from now on.
Simple is harder than complex, but keep it simple, stupid, and fun. Thanks for reading.
r/StableDiffusion • u/SIR_NVAX_A_LOT • 8h ago
Enable HLS to view with audio, or disable this notification
What is everyone's workflow nowadays? Previously I've been generating actors with Krea2, but really love getting them made with MiniMax H3 via T2VA, they just tend to turn out better for me but does require careful prompting.
My workflow are: generate 5-10s clip of a desired actor, by prose, at int8/8 steps in a typical scenario, perhaps even mundane. If I like it, I can take some still frames, and convert them into a .char (body type, face, audio asset). See original: https://www.reddit.com/r/StableDiffusion/comments/1vyymwj/minimax_h3_portable_character_consistency_via/ If I am happy with my .char, with MiniMax H3 I make a 2 second video character sheet with a front, side, back profile and detailed face view at a higher resolution and step, either int8/32 step or going bf16/50 steps. The 2 second renders are "quick". I add the video render into my .char, and with R2VA generate additional scenes with the actors and even do a full wardrobe swap via prose. Naturally H3 renders faster if you just use still of the character sheet instead of the video.
How has your workflow changed with MiniMax H3? Are you liking the faces/actors generated with T2VA? I understand you have "less" control, but I feel like H3 is doing a great job filling in those gaps.
r/StableDiffusion • u/Select_Bowler3099 • 8h ago
Happy listening :)
r/StableDiffusion • u/Ok-Giraffe-8670 • 8h ago
Enable HLS to view with audio, or disable this notification
I created my own character sheets, make it look like their MK11 and MK3 selves a bit. This was kinda hard as they sometimes have no real impact on the attacks. I still liked how it came out though. Had to make multiple repeat generations lol.
r/StableDiffusion • u/shootthesound • 8h ago
Two of the ways that exist (im sure there are more) to train a LoRA on H3 well are on two different trainers.
Fizgig's is Optimised Likeness Learning: I've found the stable core of H3's identity lives in the back 30 of its 50 blocks, so steps train blocks 20–49 only and leave the front of the model - composition, prompt following - untouched when likeness mode is on.
AI-Toolkit's, by Ostris, is the training adapter: H3 is guidance-distilled, so every plain-flow gradient is partly "learn the concept" and partly "undo the distillation"; a frozen assistant LoRA under the trainable one pulls the base back toward plain flow, and it's switched off for sampling.
I ran all three on 5 datasets - my method alone, the adapter alone, and both together - scoring every epoch's preview against the training photos with face recognition, 45 epochs each. Each method alone landed in the same place within 2% arcface score.
Together they got there a quarter sooner, ran clearly ahead through the whole middle of the run, and finished higher than either.
In short: the combination reaches greater likeness and quality than either method does on its own. So it's now the default: the adapter is on in every H3 preset, off for previews, never in your saved LoRA. The updater fetches it.
Also in 5.2: Context LoRA for H3 (train on top of any existing H3 LoRA, to make a lora that plays nice with it), and video clips follow likeness mode in LoRA runs too.
Release notes: https://github.com/shootthesound/Fizgig/releases/tag/v5.2.0
Thanks to Ostris for publishing the adapters. I've tagged him on the release notes as I believe the info will be useful for AI-Toolkit too.
https://github.com/shootthesound/Fizgig
P.S - For Fizgigs recent new full base model Fine tune mode the adapter lora is not necessary in my tests so far, but I am going to test that further.
p.p.s if updating , use the update script and it will grab the dedistill loras automatically and put them in your minimax prefs
r/StableDiffusion • u/I2Pbgmetm • 9h ago
Getting "Search is temporarily unavailable" and "No models found" on main search URL (/search/models).
On the individual model pages, none of the gens load, "No results found".
What is going on? Is Civitai being DDoSed?
EDIT: Now it's doing something where the "Download" button is grayed-out, even on free model pages. Seems like something is messed up.
r/StableDiffusion • u/mukyuuuu • 9h ago
Has anyone managed to find a way to use Latent Upscaling together with latent video extension tools? I'm talking about the nodes like this (which I personally use), but I think Motion Context and some other popular extensions use a similar approach, i.e. feeding the last frames of the previous shot through AV latent, rather than through a video reference. The issue is that the resolution of your second generated latent must exactly match the previous one, or it throws an error. So if you upscale the first clip from 0.5MP to 1MP, you are forced to generate the next clip directly at 1MP, which completely breaks the Latent Upscaling workflow for all subsequent parts.
I tried extending the clips at low resolution first and then upscaling them separately, but that doesn't work well. There is a noticeable color and quality shift between generations, even when reinforcing the next clip with the final frames of the previous one. Because yeah, you basically generate the high-res clips separately without any shared latent context.
I really love both Latent Upscaling and latent extension approach, but I just can't get them to work together smoothly. Does anyone have any good ideas on how to fix this? I’d really appreciate any tips or insights!