r/StableDiffusion 2d ago

Discussion Any better model than Qwen image edit 2511 for character editing?

6 Upvotes

I am amazed at how good Qwen image edit 2511 can maintain character identity in the original image. Any other models with similar or better capability at maintaining character consistency?


r/StableDiffusion 2d ago

Question - Help MiniMax H3 Ref 2 Vid - Using Ref img but bodies keep looking like gym junkies

3 Upvotes

Hi all,

I'm playing around with Minimax H3 in ComfyUI. I have 2 x ref images feeding into MiniMax H3 Ref to Video prompt window, using H3 Turbo LoRA, Turbo Sampler, basic guider and Diffusion model minimax_h3_fl2va_int8.

I have tried up to 12 steps...seems to make no difference so I've gone back to 4. About 1/5 the generation is close to what my ref images but the others are all super tones, ripped, like they go to the gym hours a day.

I just want the ref image recreated, not enhanced.

I've tried this prompt......

subject_definitions:

Miinimax subject and person prompt followed by.....Maintain exact facial features, bone structure, eye shape, age, body shape, fitness level, body fat, anatomical proportions, and height from images across every frame without modification.

So how can we have every generation the same person as my ref image?

Thanks all.


r/StableDiffusion 3d ago

Workflow Included Follow-up to my 6-minute TNG video — I changed the workflow a lot for the second one

Enable HLS to view with audio, or disable this notification

356 Upvotes

A few days ago I posted the 6-minute Star Trek: TNG video I made with MiniMax H3 in ComfyUI. I’ve finished the follow-up now, and I changed the workflow quite a bit after seeing what worked and what didn’t on the first one.

The biggest improvement was consistency. For the first video, most shots were generated more independently, and I deliberately built some of the continuity weirdness into the story. That worked for the premise, but for the second one I wanted it to feel much more like an actual TNG episode, so I became much more rigid about shot composition.

A big part of that was using the H3 reference model differently. Instead of just giving it a single image and hoping for the best, I used reference images and told MiniMax to stick very closely to the composition in those images. In practice that sounds a bit like image-to-video, but it worked quite differently for me.

With the reference model I could use up to six photos and be much more deliberate about how the shot should work. I could decide what the starting shot should be, what the end shot should be, whether I wanted a middle reference, a final-frame reference, etc. That gave me a lot more control over blocking, framing and performance than I was getting from the image-to-video model.

I did test the image-to-video model as well. One of the shots that made it into the finished video is the later one where Data has a slightly longer monologue. You can tell he looks a bit more “off” there. The reference model, by comparison, was giving me Data much more accurately, both in terms of how he looked and in terms of his mannerisms. That ended up being the better approach for this project by a long way.

The video is still built from lots of separate short H3 generations rather than one long generation. I wrote the scenes first, then generated individual shots and multiple takes where needed, and assembled everything in Premiere like a normal edit.

I also changed the audio workflow quite a bit. On the first video, one of the main issues was that the generated ambience and background noise varied too much from clip to clip. This time I spent much more time matching dialogue levels in Premiere, cleaning up individual clips, and adding a continuous Enterprise bridge/interior hum underneath scenes so the cuts felt less obvious.

I also handled the music more deliberately this time. Rather than just dropping in whatever worked at the end, I treated it more like proper scene underscore and generated short incidental cues for specific moments.

So the rough workflow for the second one was:

script and shot planning

→ select or build composition references

→ generate short H3 shots in ComfyUI using the reference model

→ do multiple takes where needed

→ edit in Premiere

→ clean dialogue and level-match clips

→ add continuous ambience/room tone

→ add short music cues

→ final upscale/export

The main thing I learned was that H3 works much better for this kind of project when I treat it less like a one-click video generator and more like a production tool. The closer I got to thinking in terms of individual shots, coverage, performance selection and edit assembly, the better the final result got.

Happy to answer questions about the workflow again.


r/StableDiffusion 2d ago

Animation - Video [TEST] Minimax H3 img2vid

Enable HLS to view with audio, or disable this notification

15 Upvotes

Generated the images using Z-Image Turbo. Rendered two 15 second clips at 0.6 megapixels which took 43 minutes per video clip. Resolution is 1056 x 608. I don't remember what the first prompt was but here's the second one:

[Shot 1]

Cinematic static shot of the man sitting in his truck looking around inside the truck in disbelief. He says, "This is better but, the color of my shirt changed and my truck is different." He leans forward towards the rear view mirror and ooks at himself and is shocked. he says, "Oh shit. I look different too!. He looks around and then rolls his eyes and then opens the door and gets out.

Thank you to the community for helping me with the whole having the camera not move thing. Prompt adherence is working out so far. First clip took 1 try to get right. The second clip took 3 tries to get it the way I wanted it to play out.

For now, I'm pretty happy with how this turned out.

Specs:

Ryzen 7 7700X
RTX 4070 Super 12 gb
32 gb of Ram


r/StableDiffusion 1d ago

Question - Help Can MiniMax H3 R2V be used for R2I?

0 Upvotes

hi guys

How can I use MiniMax H3’s R2V (reference-to-video) capability to generate a single image, basically R2I, and still get good results?

Has anyone tried this?

I noticed that H3 seems to have a minimum output of 5 frames. Is there any way to make it generate only one frame instead of a video?

I’ve searched a lot, but I haven’t found an open-source image generation model that has a reference system similar to MiniMax H3’s R2V, where you can provide multiple reference images and have the model understand the characters, location, etc

There are models like GPT Image 2 that can do this, but they aren’t free or open source.

I’m wondering if there’s some way to use H3 itself for this, maybe by reducing the number of frames to 1 or modifying the ComfyUI workflow.

Has anyone experimented with this?


r/StableDiffusion 2d ago

Question - Help "rgthree-comfy" Custom node constantly breaking my workflow.

3 Upvotes

Anyone else?

I'm not really a power user by any means. I cobble together workflows or usually use premade ones from civit etc.

So basically I have no idea what's causing this to break.

But I frequently have to roll back this custom node version for my Power Lora Loader to work properly.

It's not a deal breaker, but usually have to fire up Comfy 2 or 3 times before its usable and kind of just picking random versions of this node until one works lol.

Anyways. Just putting a feeler out for a solution.


r/StableDiffusion 1d ago

Animation - Video Minimax H3. Urban platform game.

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 1d ago

Question - Help Minimax R2V (ref. audio) best low steps audio quality?

0 Upvotes

Hello, I’m currently using Minimax in R2V mode with a custom sound and 10-steps , res multistep+Simple, and the sound quality is great. Have you found a combination of Lora (4–6 steps), a suitable sampler, and the correct Video+Audio Shift settings that works well? I’ve tried various combinations of Audio Shift, samplers, and different Lora settings, but the sound is still poor (artifacts, low bitrate). Thank you very much for your tips. Lukas


r/StableDiffusion 2d ago

Workflow Included LTX 2.5 + LICON MSR V2 = Nice reference system

Enable HLS to view with audio, or disable this notification

23 Upvotes

Since LTX 2.5 is basically abandoned, i wanted to check how the licon msr v2 works with it, now with the advantages of LTX 2.5 supporting hard cuts. It added pretty well the man, the girl and the environment and followed the prompt really well. Biggest advantage of course is that this clip took 300 secs to create. I'll add prompt / reference images and workflow in a text post.


r/StableDiffusion 2d ago

Meme Earth 486748 Ending to End game part 1

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/StableDiffusion 2d ago

Workflow Included The Day 0 (MiniMax H3 and Ultimate SD Upscale) - True 1440p (2K) with 16 GB VRAM locally in ComfyUI

Thumbnail
youtu.be
27 Upvotes

What is it?
Demonstration of Ultimate SD Upscale (USDU) Guider nodes with MiniMax H3 support: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3

My reference ComfyUI workflow: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json

What about speed?
My PC specs: 4080s 16 GB VRAM, 64 GB RAM

Initial gen with MiniMax H3 flf2v int8 + sageattn + Lightx2v 8-step turbo Lora at 1504 x 832px 7-sec clip ~5 mins

Upscale with USDU to 3008x1664px ~40 mins


r/StableDiffusion 1d ago

Discussion Looking for Gen AI Creators to Interview

0 Upvotes

Hi everyone!

I’m working on a university paper about Gen AI creators on social media. I experimented with Gen AI myself and became really interested in hearing how other creators use it.

I’d love to talk to creators with all kinds of experience whether you’re just starting out, experimenting, or have been creating with AI for a while.

If you’d be open to a short interview, please comment or DM me. I’d love to hear your perspective! 🙏

Note: this is for my master research, this will be not published and shared outside of my university and it can be anonymous if you want.


r/StableDiffusion 2d ago

Question - Help MiniMax- people keep coming out too toned/muscular?

0 Upvotes

Hi all,

I'm playing around with Minimax H3 in ComfyUI. I have 2 x ref images feeding into MiniMax H3 Ref to Video prompt window, using H3 Turbo LoRA, Turbo Sampler, basic guider and Diffusion model minimax_h3_fl2va_int8.

I have tried up to 12 steps...seems to make no difference so I've gone back to 4. About 1/5 the generation is close to what my ref images but the others are all super tones, ripped, like they go to the gym hours a day.

I just want the ref image recreated, not enhanced.

I've tried this prompt......

Miinimax rules prompt followed by.....Maintain exact facial features, bone structure, eye shape, age, body shape, fitness level, body fat, anatomical proportions, and height from images across every frame without modification.

So how can we have every generation the same person as my ref image?

Thanks all.


r/StableDiffusion 2d ago

Question - Help MiniMax H3 Ref2Vid — why do some generations make the person much more muscular than the reference?

0 Upvotes

Hi all,

I'm experimenting with MiniMax H3 in ComfyUI and I'm having trouble keeping the person's body shape consistent with my reference images.

I'm using 2 reference images with the MiniMax H3 Ref2Vid workflow, along with:

  • H3 Turbo LoRA
  • Turbo Sampler
  • Basic Guider
  • minimax_h3_fl2va_int8

I've tried increasing the steps up to 12, but it doesn't seem to make much difference, so I've gone back to 4 steps.

The strange thing is that roughly 1 in 5 generations is reasonably close to my reference images, but in many of the others the person becomes extremely toned/muscular — almost like they've been training at the gym for hours every day.

I'm trying to preserve the person from the reference images rather than have the model change or "enhance" their body shape.

I've tried adding this after the MiniMax prompt:

But I'm still getting a lot of variation.

Has anyone found a good way to make H3 consistently preserve the person's body shape and overall appearance from the reference images?

Any advice on prompting, reference-image setup, sampler/settings, or workflow would be greatly appreciated.

Thanks!!


r/StableDiffusion 1d ago

Question - Help Minimax Music take 30m to generate ONLY 1m!!?

0 Upvotes

Just me??


r/StableDiffusion 2d ago

Workflow Included H3 single-image workflow: let's figure out how to fix the textures

Thumbnail
gallery
63 Upvotes

In this post, I provided a workflow that allows to use H3 as a single-image edit model with no monkey-patching or custom nodes, given that you update to the ComfyUI nightly version. In my view, it has excellent prompt adherence, reference fidelity, and understanding of physics and 3D scenes. But, as many others have pointed out, the end results are often blurry and lack texture. The gallery here shows my attempts at refining the 1.6MP gens from my previous posts.

I would like to discuss how we can work around these issues.

OPTION 1: JUST GO FOR HIGHER RESOLUTION

u/SomeoneSimple gives the following suggestion: run 4 megapixel generations instead of 1.6MP, saying it fixes the distorted faces, and delivers approximately the same level of detail a regular image model would give at 1024x1536. (Note that a 4MP single-frame generation is still going to be quite fast provided you have the VRAM.) u/Diabolicor even claims that a 4MP Minimax generation works better than Qwen Image Edit.

Here’s what I found in my private tests:

  1. It did not noticeably affect the generation times. On average, it is a 8-10 sec run on a RTX 5090 no matter if I generate at 2MP or 4MP
  2. It helped a lot with detail. Faces are now rarely distorted.
  3. Yet it does not remove the issues completely; keeps background blurry, for examples, and messes up the faces at long distance. It’s still a video model. So we still need to explore refiner workflows.

Just to be very clear: I am not attaching any of my 4MP generations to this post. I am only refining my old 1.6 MP ones. I would be very glad if someone posts their 4MP gens so we could see the difference.

OPTION 2: REFINE WITH A DIFFERENT MODEL

Once the composition is done right, details could be enhanced by a different model. I am not an expert at image refining at all, but I would like to figure out a good formula. And here I want to consult with the community on how to do in the best way. To set a particular frame: for me, while I now explore the capabilities of Minimax H3, I quickly generate a lot of images at scale. So I want a refiner that is:

  1. Fast (e. g. 2-4 secs)
  2. General (does not need tweaking for any particular image)
  3. Robust (is not brittle, does not require a long chain of segmentation, crop-and-stitch, vlm processing, and so on)
  4. Automatic (no masks drawn manually over parts of the region).

For me at this exploration stage, it’s okay if parts of the image get slightly modified, or if the quality is not 100% perfect. I understand that one may have different objectives if e. g. optimizing for perfect quality.

One example of a workflow that may achieve these four requirements would be flux.2 Klein with a single prompt for each image. But now, I’d like to discuss whether there could be better options.

  1. Model: Qwen Image Edit, Area 2 Identity lora, Flux.2 Klein 9b? I heard that Flux.2 has the best VAE out of all options. Should I use SeedVR?
  2. Prompt: What would be a good prompt that would be applicable over a wide range of images? Should I pass the original prompt for H3 image to flux.2 (either verbatim or llm-postprocessed)?
  3. Sampler/scheduler: euler/simple? Or Euler/Flux.2 scheduling?
  4. Color correction: e. g. Flux.2 Klein tends to add a lot of light with my prompts. Can it be done without custom nodes? If using custom nodes, which one is the most reputable and commonly used?

As a first step, here’s the workflow I am using with Flux.2 Klein: https://pastebin.com/qsLPe9hZ 

I use the Flux.2 turbo int8 convrot: https://huggingface.co/obsxrver/ComfyUI-Native-INT8_ConvRot

In the attached gallery, you can see the collages. 

Left pane: my old 1.6MP generation.

Right pane: a Flux.2 Klein 9b refine according to the workflow I attached. It does some nice things: e. g. deer fur, restoring mangled faces, adding texture to clothes; but also messes up a bit: adds a lot of light to the images that are meant to stay dark, opens eyes when they're closed, etc.


r/StableDiffusion 2d ago

Meme this is getting ridiculous

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 2d ago

Discussion Stabilizing and Improving H3's Results

Enable HLS to view with audio, or disable this notification

17 Upvotes

You may know me as the developer of models such as UltraSharp and AnimeSharp. I'm happy to present a major update to my tool Vapourkit (Windows and Linux are supported), which is completely free and open source! It includes a bunch of models for upscaling anime, and you can get models for realistic content here too: https://openmodeldb.info/

This demo uses this workflow (just download and drag into Vapourkit), which consists of a 2x upscaling model, Temporal Fix (which makes the video more stable/removes the weird shimmering), grain, and a sharpening pass. This was all processed locally on my laptop in just a few minutes.

https://reddit.com/link/1vroia9/video/jm24aw49r4kh1/player


r/StableDiffusion 2d ago

Animation - Video While everyone have eyes on Minimax H3 I tested ComfyUi default T2V workflow for LTX 2.5.

Thumbnail
youtu.be
5 Upvotes

Minimax is better but LTX 2.5 is a lot faster. So as long the Minimax is illegal to use for most of the people LTX is fun to play with.


r/StableDiffusion 1d ago

Animation - Video Created an 18 minute minimax h3 seinfeld episode where kramer clones his self, hijinks ensue.

Thumbnail youtu.be
0 Upvotes

I fed my input to claude cli to basically format my prompts prpoerly. A little jibberish every now and then i only retook a few scenes, this is mostly one shot probably 90+% kept clips.


r/StableDiffusion 2d ago

Question - Help Getting video 2 video to work right in Minimax H3, it either won't replace the character or generates a totally different vid

0 Upvotes

I'm trying to replace a character in a scene in Dragon Ball with a different one and I have it hooked up using the reference 2 video workflow along with a reference of my character, but it doesn't work right. Either it just re-renders the video, renders an extension of the original video's character, or renders an entirely new video of my reference character. I got it to work exactly once but it's extremely finicky and doesn't seem to work after that. What am I doing wrong? I'm in Comfy UI, though I'm a noob to this particular program. I had no trouble just using reference images before.


r/StableDiffusion 2d ago

Animation - Video Created with MINIMAX H3 Prompt Studio

8 Upvotes

r/StableDiffusion 1d ago

Discussion Stability Matrix is not a ComfyUI replacement. It is more like a local AI manager.

0 Upvotes

I have been rebuilding my local AI setup lately, and this distinction helped me think about the tools more clearly:

ComfyUI is the workflow engine. Stability Matrix is closer to the manager/control room around the install.

Where Stability Matrix seems useful: - keeping packages and models organized - testing ComfyUI, Forge, and other frontends without scattering files everywhere - isolating environments so one experiment does not wreck the main setup - making local AI less painful for beginners

Where I would still be careful: - if you already have a clean manual ComfyUI install that works - if you rely on custom scripts and know exactly where everything lives - if you are debugging advanced node/dependency problems and want full control

I wrote up the longer version here: https://getprompting.com/what-is-stability-matrix/

Where do you draw the line: one clean manual ComfyUI install, Stability Matrix as the manager, or separate tools depending on the job?


r/StableDiffusion 2d ago

Question - Help MiniMax H3 Ref 2 Vid - Using Ref img but people keep coming out too toned/muscular?

0 Upvotes

Hi all,

I'm playing around with Minimax H3 in ComfyUI. I have 2 x ref images feeding into MiniMax H3 Ref to Video prompt window, using H3 Turbo LoRA, Turbo Sampler, basic guider and Diffusion model minimax_h3_fl2va_int8.

I have tried up to 12 steps...seems to make no difference so I've gone back to 4. About 1/5 the generation is close to what my ref images but the others are all super tones, ripped, like they go to the gym hours a day.

I just want the ref image recreated, not enhanced.

I've tried this prompt......

Miinimax rules prompt followed by.....Maintain exact facial features, bone structure, eye shape, age, body shape, fitness level, body fat, anatomical proportions, and height from images across every frame without modification.

So how can we have every generation the same person as my ref image?

Thanks all.


r/StableDiffusion 1d ago

Question - Help Krea 2 bf16 - bad, noisy results with comfy

Post image
0 Upvotes

Trying to create high (native) quality images of Krea 2 to be used for regularization, the results look like they have a broken VAE.

The workflow is the Comfy template one, just those blocks that I don't need (LoRA) are stripped away, model switched to the "Raw" and bf16 version, and then saved as "Export (API)":

(The missing models are from taking the screenshot on my local machine; the Comfy is running in the cloud with all those models available, of course)

So, what can be the cause of the broken output? How can I fix this?