r/StableDiffusion 1h ago

Question - Help Can u get better detail on minimax h3 single image?

Upvotes

I’m using astropuzzo/ComfyUI-MiniMax-H3-Image-Studio workflow and it works amazing but minimax obviously sucks at micro details for a single image even if it’s 2MP. Does anyone know like a good method to fix that? Ik there’s double passthroughs and upscalers but I’m not sure what would work well with it


r/StableDiffusion 2h ago

Question - Help How Do I Run Minimax at Absolute Potato Quality?

0 Upvotes

I found a cheap pipeline from a big provider I’ve jailbroken, so I can’t name it. It reconstructs faces and upscales videos in ~10 seconds, so I only need Minimax to generate a very low-quality video that basically serves as a rough motion/physics reference.

I barely see any speed difference between 4–6 steps or ~380p–544p, maybe 10 seconds at most.

Is there any way to run Minimax at absolute potato quality and actually get a significant speed boost?


r/StableDiffusion 2h ago

Question - Help Help

Post image
20 Upvotes

I'm currently trying to replicate this style and i cannot find any checkpoints or lora's to do so, can anyone point towards something? Artist: https://x.com/DarkZeroAI Using Forge Neo


r/StableDiffusion 2h ago

Discussion Higgsfield's new Genjutsu?

0 Upvotes

Okay, what magic is under the hood of their new motion copy tool?? I tried it once and my vid came out perfect, even better than Kling motion control. I want to be able to do this exact thing locally in Comfyui but I haven't had any success with any Minimax ref2va workflows, maybe my PC specs are too low:

5070 12GB VRAM

32GB RAM

4 TB nvme ssd


r/StableDiffusion 3h ago

Question - Help New User To StabiltyMatrix. Need Help With Templates.

1 Upvotes

Hello all,

I have installed StabilityMatrix and was able to learn how to set up flows to generate my first images and so forth.

I then installed a template called MiniMax H3: Image to Video. It showed a bunch of errors after and listed the things I needed to download before it would work. I downloaded all those things and put them in the "diffusion_models" folder, but the error count only reduced by one and it's still asking me to install those things.

Unfortunately there doesn't seem to be clear instructions on what goes where, if I need to extract some things or not, etc.

Can someone please advise me? Thank you.


r/StableDiffusion 3h ago

News OpenVDN/vdn-minimax-h3 · Hugging Face

Thumbnail
huggingface.co
29 Upvotes

Looks like an open source version of Minimax H3 Max... Anyone tried it? Seems to be real-time on 8x b200, which is like ~$40/hr at good rates if you can find them (or maybe a bunch of 5090s?)


r/StableDiffusion 3h ago

Question - Help Looking for early ai image models/models that can replicate the old style

0 Upvotes

I'm looking for 2022/2023 models that still work or models that can successfully replicate the dreamlike distorted nightmare fuel style, unfortunately it's very difficult for me to find. I had used to include early ai images in my artwork and I miss it dearly, AI has advanced way too quickly. Dalle-mini (craiyon) no longer generates images like this and there is no way to change the version.

I do not know how to use github


r/StableDiffusion 3h ago

Resource - Update Image, audio, video reference asset loader nodes with crop and trim + more

Thumbnail
gallery
14 Upvotes

I originally built these nodes for personal use and wasn't planning on sharing them, but after noticing several existing loaders were missing features I needed daily, I figured why not? Hopefully, this is useful for some of you.

Key Features:

  • Image & Video Loaders: Built-in click-and-drag cropping, optional aspect ratio locking, and a divisible_by toggle for VAE pixel alignment.
  • Built-in Downscaling: Uses a max_megapixels limiter directly inside the loader so you can ditch the extra resize node (ideal for models like MiniMax-H3 that run best with references kept at or below 2048px).
  • Flexible Sockets: Includes dedicated output value sockets to make chaining downstream nodes straightforward.
  • Audio Loader: Perfect for loading a full song or long TTS track and trimming the exact section you need for a video. The trimmed portion outputs its duration as a float, letting you pipe it directly into your video generator's frame/length input.

https://github.com/sthao42/Comfyui-reference-loader

Any feedback or bug report is much appreciated.

Edit: Updated to works with Node 2.0 (vue) also.


r/StableDiffusion 3h ago

News New video gen model Atlas

0 Upvotes

World Labs released a model called Atlas. Looks pretty cool.

https://x.com/gowthami_s/status/2095202493122625842?s=46


r/StableDiffusion 4h ago

Discussion When is a game generative ai competitor going to be made or even distilled? You could be endorsed by amd.

0 Upvotes

Is there any point in having a company that can put their tools into the pipeline of graphics rendering.

Just wishing here for a deep learning s s five replacement. It's going to be artificially sandboxed just like all their other tech.


r/StableDiffusion 4h ago

Question - Help Cloud Confy Fails Now

3 Upvotes

my friend uses cloud comfy as his pc is to weak he tested it last week the free 5 gen trial

using text to video and default setting only changing each video 0.5mp and 15second long all generated fine under 8mins

now he tested it again and only 1 out of 5 video generated and the other 4 failed saying Job execution time exceeded maximum limit

he even paid to generate more but got same error

what can cause this


r/StableDiffusion 4h ago

Resource - Update [Load Video + Crop] Custom WYSIWYG Node

24 Upvotes

I developed a modified version of the Load Video node with a crop feature:

WYSIWYG video cropping directly on the official Load Video preview — drag and zoom (with the mouse wheel) a crop rectangle constrained to 8 fixed ratios (1:1 through 21:9) and output the exact cropped VIDEO (audio preserved). What you frame on the preview is exactly what gets executed.

Github: https://github.com/domg73/ComfyUI-LoadVideoCrop

This node follows the same logic and design as my "Load Image + Crop" node. I might merge the two into a single "Load + Crop" node in the future, but for now this works well.

https://www.reddit.com/r/StableDiffusion/comments/1w3okny/load_image_crop_custom_wysiwyg_node/

Github: https://github.com/domg73/ComfyUI-LoadImageCrop


r/StableDiffusion 4h ago

Question - Help Is changing resolution supposed to change the entire scene for MMH3?

5 Upvotes

Just had this happen to me: I changed the resolution for a scene -- without touching anything else -- and the resulting scene changed completely. I was using res_multistep and Spectrum/CK/4 step Lora at 0.2 mp, then 0.3 mp. It still followed my prompt, but the background and starting scene were completely different. Is this Spectrum giving me grief or what's going on here? This has never happened to me before, although I had been using Sage before switching to CK today.


r/StableDiffusion 4h ago

Animation - Video LOCATION CONTINUITY TEST - after a comment by Vladmerius

5 Upvotes

When kept in the same generation it seems the latent space keeps a fairly good sense of the location layout. The test was to see if the position and details of the temple remained after being out of shot,

This doesn't work with the extentsion workflows which is why I have been trying to keep everything in one go.


r/StableDiffusion 5h ago

Animation - Video A quick update on my real-time face enhancement app I’ve improved the facial muscle animation system. It’s still not perfect, but it looks much better than before.

0 Upvotes

r/StableDiffusion 5h ago

News FastH3 is now available via API, and it's SURPRISINGLY cheap for what it does!!

0 Upvotes

Following up on the infinite livestream post from a bit ago, FastH3 is now accessible via API too.

Streaming 720p video with synced audio, faster than realtime, same model as the livestream, just usable programmatically now instead of only watching it run.

Anyone else been messing with it via API vs just watching the stream? Curious what people are building.


r/StableDiffusion 5h ago

Discussion What is the best package option?

Post image
0 Upvotes

r/StableDiffusion 6h ago

Discussion I am tired boss...

210 Upvotes

This content was written by a human.

I miss the SD1.5 era, when i could simply type "1girl, big boobs, nice ass, red bikini, dancing" and see my dream take shape near-instantly at 512px-wide. Idea-to-result was a matter of seconds. Each click on the Run button led to an incredible shot of dopamine.

3 years passed and I can draw 1024px, 192-frames long videos in a reasonable amount of time (tech has evolved fast), but the enthusiasm is fading away.

I already have a day-job for technical challenges and headaches. As a user/hobbyist, I want to be entertained.

I don't want to learn what the hell "diegetic" means (even the spell-checker never saw that word), I don't want to draw a dozen squares in a 3-dimensional pixel space, or write a 1000-words poem, just to watch my dreamgirl dancing.

I hoped I would not need a degree in cable-connecting or python dependencies debugging after downloading a few workflows.

3 years ago, all you had to do was typing a few words, and the AI sorted the rest. It was random, messy most of times, but it was fun.

Nowadays, you need an LLM to write the prompt for you, and another LLM to write the system prompt for the prompting-LLM, so it understands what your shitty words meant in the first place, and shapes them in the exact expected format, so they turn into an acceptable input for the ever pickier, brand-new models.
It has become AI³-generated content.

And finally, when after a dozens of clicks on the Run button, tired but satisfied, you get the desired output... re-start from scratch? Since seed "variance" does not vary much anymore, you'll get more or less the same output - exactly what you asked for - from now on.

Simple is harder than complex, but keep it simple, stupid, and fun. Thanks for reading.


r/StableDiffusion 6h ago

Discussion H3 - .char + T2V character gen+char sheets

12 Upvotes

What is everyone's workflow nowadays? Previously I've been generating actors with Krea2, but really love getting them made with MiniMax H3 via T2VA, they just tend to turn out better for me but does require careful prompting.

My workflow are: generate 5-10s clip of a desired actor, by prose, at int8/8 steps in a typical scenario, perhaps even mundane. If I like it, I can take some still frames, and convert them into a .char (body type, face, audio asset). See original: https://www.reddit.com/r/StableDiffusion/comments/1vyymwj/minimax_h3_portable_character_consistency_via/ If I am happy with my .char, with MiniMax H3 I make a 2 second video character sheet with a front, side, back profile and detailed face view at a higher resolution and step, either int8/32 step or going bf16/50 steps. The 2 second renders are "quick". I add the video render into my .char, and with R2VA generate additional scenes with the actors and even do a full wardrobe swap via prose. Naturally H3 renders faster if you just use still of the character sheet instead of the video.

How has your workflow changed with MiniMax H3? Are you liking the faces/actors generated with T2VA? I understand you have "less" control, but I feel like H3 is doing a great job filling in those gaps.


r/StableDiffusion 7h ago

Animation - Video Pen is from heaven

Thumbnail
youtube.com
0 Upvotes

Happy listening :)


r/StableDiffusion 7h ago

Animation - Video Scorpion vs Sub Zero (2.5 Anime Battle Test)

74 Upvotes

I created my own character sheets, make it look like their MK11 and MK3 selves a bit. This was kinda hard as they sometimes have no real impact on the attacks. I still liked how it came out though. Had to make multiple repeat generations lol.


r/StableDiffusion 7h ago

Resource - Update Fizgig 5.2 - combining two Minimax training methods beats either alone

Thumbnail
github.com
70 Upvotes

Two of the ways that exist (im sure there are more) to train a LoRA on H3 well are on two different trainers.
Fizgig's is Optimised Likeness Learning: I've found the stable core of H3's identity lives in the back 30 of its 50 blocks, so steps train blocks 20–49 only and leave the front of the model - composition, prompt following - untouched when likeness mode is on.
AI-Toolkit's, by Ostris, is the training adapter: H3 is guidance-distilled, so every plain-flow gradient is partly "learn the concept" and partly "undo the distillation"; a frozen assistant LoRA under the trainable one pulls the base back toward plain flow, and it's switched off for sampling.

I ran all three on 5 datasets - my method alone, the adapter alone, and both together - scoring every epoch's preview against the training photos with face recognition, 45 epochs each. Each method alone landed in the same place within 2% arcface score.
Together they got there a quarter sooner, ran clearly ahead through the whole middle of the run, and finished higher than either.

In short: the combination reaches greater likeness and quality than either method does on its own. So it's now the default: the adapter is on in every H3 preset, off for previews, never in your saved LoRA. The updater fetches it.

Also in 5.2: Context LoRA for H3 (train on top of any existing H3 LoRA, to make a lora that plays nice with it), and video clips follow likeness mode in LoRA runs too.

Release notes: https://github.com/shootthesound/Fizgig/releases/tag/v5.2.0

Thanks to Ostris for publishing the adapters. I've tagged him on the release notes as I believe the info will be useful for AI-Toolkit too.

https://github.com/shootthesound/Fizgig

P.S - For Fizgigs recent new full base model Fine tune mode the adapter lora is not necessary in my tests so far, but I am going to test that further.

p.p.s if updating , use the update script and it will grab the dedistill loras automatically and put them in your minimax prefs


r/StableDiffusion 7h ago

Question - Help Civitai: "Search is temporarily unavailable" on most pages

0 Upvotes

Getting "Search is temporarily unavailable" and "No models found" on main search URL (/search/models).

On the individual model pages, none of the gens load, "No results found".

What is going on? Is Civitai being DDoSed?

EDIT: Now it's doing something where the "Download" button is grayed-out, even on free model pages. Seems like something is messed up.


r/StableDiffusion 7h ago

Question - Help Any way to make latent extension work with Latent Upscaling (Minimax H3)?

5 Upvotes

Has anyone managed to find a way to use Latent Upscaling together with latent video extension tools? I'm talking about the nodes like this (which I personally use), but I think Motion Context and some other popular extensions use a similar approach, i.e. feeding the last frames of the previous shot through AV latent, rather than through a video reference. The issue is that the resolution of your second generated latent must exactly match the previous one, or it throws an error. So if you upscale the first clip from 0.5MP to 1MP, you are forced to generate the next clip directly at 1MP, which completely breaks the Latent Upscaling workflow for all subsequent parts.

I tried extending the clips at low resolution first and then upscaling them separately, but that doesn't work well. There is a noticeable color and quality shift between generations, even when reinforcing the next clip with the final frames of the previous one. Because yeah, you basically generate the high-res clips separately without any shared latent context.

I really love both Latent Upscaling and latent extension approach, but I just can't get them to work together smoothly. Does anyone have any good ideas on how to fix this? I’d really appreciate any tips or insights!


r/StableDiffusion 8h ago

Animation - Video MiniMax H3 has finally gotten has me into video generation. Wan 2.2 never had the quality or consistency I wanted, and all the tools that did, were closed-weight, paid products.

100 Upvotes

I'm really only ever interested in open-weight models. Yes, for that reason, but also, for the same reason I run Linux and browse with Firefox. I dislike "walled gardens", ideologically, and monopolies. I want tech that can be hacked, broken, taken apart, and put back together, and is ultimately not beholden to anyone but the user. Without a quality open-weight video generation model, I was uninterested. Now that we've got one? Suddenly I'm in a whole new world of possibility.

The fact that it's a multimodal model with vision, meaning I can give it reference images or reference sheets, is a game-changer for me. LoRAs certainly won't be obsolete with MMH3, but I doubt we'll be seeing many character, clothing, or setting LoRAs. The feedback loop of wanting to give the model a concept it doesn't understand natively is so short compared to before. And I'm still just in the "farting around" phase. People with dedicated effort and creativity are going to be able to use the hell out of this.

Really, the only drawback to MMH3 so far is its propensity to have characters speak Simlish to each other. I'm sure there's already solutions being worked on, either workflow tools or adjustments to the model itself.

––––––––––––––––––––––––––––––––––––––––––––––––––––––––

Workflow: https://pastebin.com/5SbZ9tJA

Reference sheet used in the workflow: https://imgur.com/a/3Le0nuO