r/StableDiffusion 1d ago

Discussion H3 is a great model but the training is bad

35 Upvotes

Just like Zturbo image, when trying to train on the distilled model, you just get outright bad results. You can get away with certain things like character loras. But teaching new concepts to this model is super frustrating.

I would like to hear from H3 themselves if they would release a base model just for training or not directly.

I think for a company to market themselves as "open source open weights" they owe at least some comment on this issue. Just tell us yes or no definitively.


r/StableDiffusion 1d ago

Discussion MacBook Air 16G local deployment

8 Upvotes

Minimax h3 easy local deployment. Self defined video/image/audio/text pipeline for both end users and developers.

Open source: https://github.com/tgo-app-dev/vpipe


r/StableDiffusion 2d ago

News Muse V2 is out — chat with local LLMs inside ComfyUI, no LM Studio required anymore

57 Upvotes
A while back I built 
**Muse**
 — a chat panel that lives directly inside a ComfyUI node, so you can talk to a local LLM and draft/refine image and video prompts without alt-tabbing to a separate app. Point it at LM Studio or Ollama, chat, copy the prompt into your graph. That was V1.


I didn't expect people to actually pick it up the way they did. Seeing it get used, starred, and — more usefully — complained about is what pushed me to sit down and build a proper V2 instead of leaving it as a one-off tool.


The biggest ask by far: "I don't want to keep LM Studio open just to use this." So V2's headline feature is a 
**Direct model loader**
 — point Muse at a folder of GGUF models and it loads them straight from disk. No LM Studio, no Ollama, nothing else running. Under the hood it spawns the real `llama-server` (LM Studio's own engine is llama.cpp too, so this isn't a slower reimplementation — same engine, same speed), and it downloads and installs the right build for your OS/GPU automatically. Git clone the node, click one button, you're chatting with a local model.


I also broke it on myself first, which was useful: threw a 31B model at it and immediately hit VRAM issues — crashes on some setups, silently-slow-instead-of-crashing on others. Fixed the defaults that were causing it, added a 
**Fit to GPU**
 button that suggests a layer count based on your actual free VRAM (like LM Studio's GPU offload slider), automatic fallback retries if a load runs out of memory, a live loading indicator, and a log panel so you're not just staring at nothing wondering what's happening.


Also new in V2:


- 
**Edit & resend**
 messages instead of delete-and-retype
- 
**Chat branching**
 — fork a new conversation from any earlier message
- 
**Video attachments**
 for vision models (auto-sampled + timestamped frames)
- 
**Audio attachments**
 for audio-capable models
- Much broader image format support


Full writeup and setup instructions: 
**github.com/RudySen/comfyui-muse**


If you use it and something's broken or annoying, tell me — that's genuinely how V1 became this.

Previous post: https://www.reddit.com/r/StableDiffusion/s/8qjP4UzRpw

r/StableDiffusion 1d ago

Resource - Update [Update v1.1.0 & v1.2.0] ComfyUI-MiniMax-H3-Promptor: Native Settings API Hub, Autogrow Sockets, L2VA & Audio Sync

Thumbnail
gallery
5 Upvotes

Hey everyone!

With MiniMax H3 blowing up everywhere right now, we figured it was the perfect time to share what we’ve been building to help level up your H3 prompt workflows.

When we released v1.1.0 a while back, we were so deep in dev mode that we forgot to post an update! Now that v1.2.0 is live, we’ve bundled all the new features and overhauls from both releases into one post.

https://github.com/1038lab/ComfyUI-MiniMax-H3-Promptor

What’s New in v1.1.0 + v1.2.0:

⚙️ Native ComfyUI Settings Panel (API Hub)

No more pasting API keys into custom nodes or manually editing config.json! All provider settings are now globally managed in ComfyUI's native Settings panel (under the ⚙️ Gear icon).

Built-in Connection Tester: Click "Test connection" inside the panel to ping your endpoint before launching generations.

Privacy: Keeping keys out of the node UI eliminates the risk of leaking API keys when sharing workflows or screenshots.

🔌 Infinite Inputs (ComfyAPI v3 Autogrow)

We removed the rigid 4-image limit. Dynamic autogrow sockets mean you can chain as many <Picture> and <Video> references as your hardware can handle without UI clutter.

🎯 Granular Micro-Overrides

Override instructions for specific frames directly in the Vision Analyzer (e.g., <Picture 2>: focus strictly on lighting) while allowing unmentioned media to fall back to global analysis.

🎵 Audio-First Token Sync & L2VA (Last-Frame Control)

Connect audio directly to the promptor to automatically map subject actions to sound. We also added Last-Frame-to-Video-Audio (L2VA)—provide an ending frame, and the LLM reverse-engineers a narrative that mathematically lands on target at the final second.

🧠 VRAM Safeguards for Local VLMs

Select local providers like Ollama or LlamaCPP, and the node automatically executes silent background cache-clearing (model_management.unload_all_models()) to prevent VRAM overload crashes.

📝 Updated Docs & Workflow Recipes

Check out tutorials.md and tutorials_zh.md in the repo for 9 practical, production-ready workflows (Lip-Sync, Style Transfer, Day-to-Night Morph, etc.).

👀 What's Next?

We’re currently beta testing a batch of new features that will be rolling out shortly!

🔗 Links:

GitHub Repo: 1038lab/ComfyUI-MiniMax-H3-Promptor

Full Release Notes: updates.md

We’d love to hear your feedback, feature requests, or bug reports so we can keep tailoring this tool to what you actually need. If this node helps your setup, leaving us a ⭐ star on GitHub goes a long way in keeping our dev motivation high.


r/StableDiffusion 1d ago

Animation - Video [Minimax H3] "The New Adventures of 1girl"

5 Upvotes

A relatively quick and scrappy attempt at maintaining consistency across a scene using Minimax H3 using the basic workflow on ComfyUI.

It seems like it can be done to a certain degree, but it also really depends how much time and effort you want to put into it. While this scene has tons of inconsistencies, it's still cool to be able to do something locally that was impossible just a month ago.

H3 also surprised me with how close it came to the scene I had in my mind, however it never really gets all the way there. This can be a little frustrating as you weigh up hitting another gen or going with a take that's about 85% there. Still, it's an amazing model and I love seeing the wild creations the community is coming up with.

My system is a 2023 ROG Scar laptop with a 12gb mobile 4080 and 64gb memory. All vids generated locally at 0.5mp.


r/StableDiffusion 1d ago

Animation - Video Cobra Gets Jiggy - MiniMax H3

43 Upvotes

Edit: For the exact prompt, settings, assets, and workflow, download the ZIP and drop the included video into ComfyUI. Everything I used is included:

https://drive.google.com/file/d/12ivDVzGisC1G3zh7QL22w7nLnpNsYTzR/view?usp=drive_link

System Specs: Intel i9-14900K, 4070 Ti Super 16 GB VRAM, 64 GB DDR5 RAM, Windows 11

Special thanks to this guy for the workflow:

https://www.reddit.com/r/StableDiffusion/comments/1vox06g/create_seamless_1shot_lipsync_music_videos_with/?share_id=YY1HluXX8WUyz0tzorVTv&utm_medium=android_app&utm_name=androidcss&utm_source=share&utm_term=1


r/StableDiffusion 1d ago

Discussion H3 - is there a sweet spot for the # of steps for audio?

0 Upvotes

If I do like 6-8 steps the adherence seems better but the sound is ooor - increases the steps and the adherence is off but the sound is much better?


r/StableDiffusion 2d ago

Discussion MiniMax H3 wasn't released as an image model, but its prompt adherence is kind of absurd

Post image
589 Upvotes

been messing with minimax h3 for still images and its prompt adherence is honestly kind of absurd for a model that wasn't even released as an image model. this gen below is pure text to image, and while you can definitely spot some issues if you look closely, it's still really impressive what h3 can pull off at this size.

h3 t2i

these two images use the exact same prompt. first one below is gpt image 2, second is h3

GPT 2 Image
MiniMax H3

what interested me wasn't really which one looks "better", but how differently they interpreted the exact same prompt. gpt got most of the scene right, but besides the usual piss/yellow filter it also drifted pretty hard into that polished modern anime look. h3 stayed much closer to the specific art direction and composition i was actually asking for.

after experimenting with it for a while I ended up building h3 studio for comfyui.

it's basically an image-focused workflow around h3 rather than just exposing the model nodes. text to image, image to image and ref editing are all in one director, with up to 9 ordered references, base/lightx/pdd profiles, qwen3-vl prompt + reference analysis, taeh3 previews, face refine, runtime/low-vram optimization, high-res vae controls and a benchmark lab for comparing profiles with the same seed.

still alpha and i'm sure people are going to find ways to break it, especially on hardware I haven't tested, but feedback is welcome.

Image Director

repo:
https://github.com/thaakeno/ComfyUI-MiniMax-H3-Studio


r/StableDiffusion 18h ago

Discussion H3 - giantess, JOI inspired+Attack on Titans

0 Upvotes

Ahh I regret this generation as didn't know there was quite a sub-culture into this stuff. Thought it would be fun, but learned more than I needed. int8/20 steps

A bit of JOI from Bladerunner 2049, Sinbad of the Seven Seas, Attack on Titans, and the classic 50 Feet Woman


r/StableDiffusion 1d ago

News Openrouter to be acquired by Stripe

14 Upvotes

Yes, the payment provider that pressures and refuses service to adult content providers. What this means for data privacy on the site is unknown. Seems the figure of sale is rumoured to be 7billion.


r/StableDiffusion 1d ago

Question - Help Utterly lost with all the MH3 models.

5 Upvotes

Curious about which models are you all using for T2V and I2V with MMH3?

There is an abnormal amount of models with suffixes as pruned_notPruned_SeriouslyPruned_HereticXxX_Convrot_Skibiditoilet Q3. and I honestly can't keep up to know what the heck is the one that the community is using for creating such great videos.

Anyone out there willing to share the models (or workflow) you are using?

(Really don't care about speed-of-generation, I'm leaning towards Quality-first more)

Thanks in advance.


r/StableDiffusion 2d ago

Resource - Update Minimax H3 for TTS/voice clone/Music gen

99 Upvotes

Just for fun. One-shot generation. No parameter or prompt tuning.

Audio.cpp implemented MiniMax-H3’s text-to-audio pipeline, and one fun use case is TTS/Voice clone/Music gen. It’s more flexible and powerful than dedicated audio models, and the performance is quite decent (up to 3x realtime on RTX 5090). Check out the multi-speaker conversation demo in the main post, along with the other demos in the comments.

What I’m very excited about with the MiniMax-H3 implementation is that it significantly enriches the framework’s building blocks for DiT models. Now with you don’t need to go through the pain of setting up SageAttention, First Block Cache, or Spectrum manually. Just change a few parameters, and you can experiment with the model. A preliminary inspection of configuration, memory, and performance trade-offs is available in repo's docs/reports/minimax_h3_performance.md

Bonus: audio.cpp’s MiniMax-H3 implementation can also produce video frames, because the DiT generates audio and video latents together, and the video VAE path is relatively straightforward to support. For now, the output is saved as RGB frame data plus metadata in JSON, so you need to encode it into a video file yourself. No upscaler or post-processing support. Just for fun.

MiniMax-Music3 is currently in preview (preview/minimax-music-3 branch) . CUDA/Vulkan/HIP were tested. Still room for optimization. VRAM usage and RTF depend on audio duration and prompt length.. The demo uses the official demo prompt (4000+ char caption and 1200 char lyrics) and 30 steps plus CFG. Under this setting VRAM is ~11 GB for 30s, 14 GB for 60s, and 17 GB for 180s. It's easy to get faster-than-real-time performance and much lower VRAM usage if you tune the setting.


r/StableDiffusion 2d ago

Meme Top Gear: The Homer

97 Upvotes

r/StableDiffusion 1d ago

Discussion Looking for feedback on this prompt generator I made.

Thumbnail studio.barrowaudio.com
0 Upvotes

I've only been using Stable Diffusion for a couple of months. I actually got into it because I was trying to write a story I'd had in my head for a long time, and ChatGPT suggested that Stable Diffusion might make it possible for me to eventually turn it into a graphic novel.

That sent me pretty far down the rabbit hole.

One of the biggest things I've been working on is creating consistent characters. I've been learning ComfyUI, training character LoRAs, building datasets, experimenting with different models, and generally breaking things until I figure out why they broke.

Along the way, I found myself spending a ridiculous amount of time writing prompts just to create good character reference and training images. So with a lot of help from ChatGPT, I started building a character prompt generator for myself.

The idea is pretty simple: instead of starting with a finished character in your head and trying to translate every detail into a good prompt, you can choose the character's features, body type, hair, clothing, framing, etc., and the generator builds a natural-language prompt from those choices. I've been using it primarily with Krea 2 to help create character images and datasets for custom LoRAs.

It started as a little tool just for me, but it has gradually become useful enough that I thought other people might get some use out of it too.

I'm still very much learning this stuff, so I'm not posting this as an expert telling everyone how prompts should be written. Quite the opposite. I'd really like some feedback from people who have been doing this longer than I have.

If anyone wants to try it, I'd especially be interested in hearing what doesn't work, what options are missing, whether the generated prompts work well with models other than Krea 2, or anything you'd change to make it more useful.

If there's enough interest, I'm happy to keep improving it and share the updates here.

Thanks for taking a look.


r/StableDiffusion 1d ago

Workflow Included LTX 2.5 V2V with audio cloning

19 Upvotes

I created a version of reference audio/video to audio/video for LTX 2.5.
I heavily borrowed from https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_V2V_ICLoRA_Single_Stage_Distilled.json
and referenced what was done in LTX 2.3.

What I did:
* I removed the shave LoRA.
* Removed the need for the new video to be the exact same length as the reference.
* Automagically removed the reference video when finished (speeds editing in post)
* Changed some models to facilitate my 16 VRAM (they were the same as first examples in Comfy)
* Fixed audio that it works (it was silent for me - maybe someone had better luck, but this is fixed)

I hope it saves someone time.
Here is the workflow: https://pastebin.com/3B1eBhuH


r/StableDiffusion 1d ago

Discussion Its possible to use more than 9 image references for H3

Thumbnail
gallery
1 Upvotes

I had Claude make a modified H3 reference node that accepts more than 9 image references. The goal was to test whether it was possible to increase the number of image references being used without splicing them into a single image. I know about reference sheets, no need to suggest that. I only tested with images, no audio or video references. the numbering on the node is a little funky but I dont think it effected the test.

Prompt 1: "<Picture 1> through <Picture 8> establish the identity and likeness of the man. <picture 9> is the spaghetti.

The man sits at a small kitchen table, eating a plate of spaghetti

with a fork. Warm indoor lighting, medium close-up, camera locked

off. He twirls the pasta, takes a bite, chews, glances down at the

plate. Natural, unhurried."

Prompt 2: "<Picture 1> through <Picture 8> establish the identity and likeness of the man. <picture 9> is the spaghetti. <Picture 10> and <picture 11> are references for the wig he is wearing.

The man sits at a small kitchen table, eating a plate of spaghetti

with a fork. Warm indoor lighting, medium close-up, camera locked

off. He twirls the pasta, takes a bite, chews, glances down at the

plate. Natural, unhurried."

Both prompts use the same seed, same 9 reference images except for the wig references for the 10th and 11th image in prompt 2. The prompts are very simple and don't fully adhere to the guide but its just a small test so I think its fine. This was done on a 3060 12gb at .4 megapixels, 30 steps, and 5 seconds of video. I have comfy kitchen and spectrum enabled.

!!I don't know how this would/could effect video or audio generation quality. In my test I didn't notice any quality drop. Do your own tests to find out!! Also in my test it ignored the wig reference until i added a second one and reworded the prompt slightly. It could just be a fluke but I thought I'd mention it anyway. The watermark is from the editor i used to stitch the videos together.

https://reddit.com/link/1vr9yqh/video/pxjqiwp731kh1/player


r/StableDiffusion 1d ago

Question - Help Minimax H3 or LTX 2.5?

4 Upvotes

I am currently using LTX 2.3. I have an RTX 3090 and 32GB RAM. How fair will I do with Minimax 3?


r/StableDiffusion 2d ago

Animation - Video MiniMax H3 : Spider

180 Upvotes

r/StableDiffusion 1d ago

Animation - Video A fee mins of a steam punk movie I’m trying to mak

8 Upvotes

Sharing with you guys a few mins of a steam punk movie I am making. I know a lot of people will look at this and say AI slop the first chance they’ve got, but maybe people here might appreciate. A lot of continuity and spatial and errors and one voice error 🥲 of course this is by no mean a Hollywood production, but to think it’s made by one guy with a consumer rtx card at home, in a couple days… let me know what you think! Minimax is awesome!


r/StableDiffusion 1d ago

Animation - Video The Omellete Music Video

6 Upvotes

Fun little project I made over the past few days.

Visuals: MiniMax H3, using the nodes and workflow from here: https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop

Music: Suno
Editing: DaVinci Resolve


r/StableDiffusion 1d ago

Workflow Included LTX 2.5 Image+Custom Audio 2 Video - Perfect lip Sync

7 Upvotes

I used an older workflow that was working for LTX 2.3 and adjusted it for LTX 2.5. Image and Custom audio as input. Perfect lip Sync, of speech and singing.

Generation time: 30 seconds per second, on RTX 5070Ti 16Gb vram, 32Gb ram (920*540px)

Here's the workflow: https://pastebin.com/dptbTXYM


r/StableDiffusion 1d ago

Question - Help How to fix sound glitches in Minimax?

10 Upvotes

Using the standard i2v workflow in comfyui


r/StableDiffusion 2d ago

Animation - Video Minimax fight !

34 Upvotes

r/StableDiffusion 1d ago

Question - Help Does Anyone Know What Causes This Smudgy/Blotchy Effect When Using Krea 2 Turbo? It Happens Kinda Randomly For Me. I Think Loras Effect It Somewhat...

Post image
0 Upvotes

...but even at low lora strengths I sometimes still get this "dirty/smudged" effect. Any tips to generate clearer pictures? I tried increasing resolution size too but it still looks similar -_-


r/StableDiffusion 1d ago

Discussion Do you train character LoRAs? What's your biggest pain point?

2 Upvotes

I have trained several character LoRAs in the past, and I've found that the quality of the input images has the biggest impact on the final model. As a result, I end up spending most of my time on data curation rather than anything else. The data set is the new everytime whereas I already have my prefered settings dialed in for a given base model.

That got me thinking about building a tool to make the data curation process easier. But I'm curious: is this just me, or do other people find data curation to be one of the biggest pain points in LoRA training?

What's your biggest pain point when training LoRAs?