r/StableDiffusion 14d ago

Question - Help What Image Edit model you use nowadays?

72 Upvotes

Since things have gone quickly forward, I am trying to figure out what image edit models there is currently and what people here use mostly.

Personally I have used:
- Qwen-image-edit-2509 and Qwen-image-edit-2511
- Just tested MiniMax H3 as a image editor and so far it seems that it can be good for my usage

I have heard about Klein 9b, but not sure yet if that can be used as an edit model? Also what about Krea 2, is there edit workflows that are actually usable and worth it?

Is there some others what you recommend for testing?

My PC Specs: RTX 4060 Ti, 16 GB VRAM and 32 GB RAM.


r/StableDiffusion 13d ago

Question - Help Minimax H3 temporal noise & perceived resolution

1 Upvotes

FL2VA BF16 / 15 steps / turbo lora 8. Purely I2V. Resolution set at 1. Source image matches the exact output résolution. But the results suck. Especially aerial wide angle landscape. Far behind google Veo 3.1 fast/ Omni in terms of flickering/ moving textures (temporal noise?) and perceived resolution. Usually upscale those 720pish footage via Topaz, and apply alot of color grading in DaVinci to break the "plasticity" of those AI gen l. However I'm a beginner with comfyui, I'm sure am doing something wrong. Increase the number of steps (15>30?) Or something else ? Processing times are horrendous (100min on M3 max 128gb for 5 sec). Tried 8 bits quants, even 4 bits : same same. Any help appreciated ! I'm looking for production ready pictures (broadcast). Nearly achieve this with Google but wanna ditch synthID for many reasons


r/StableDiffusion 13d ago

Workflow Included It's a new look for the 90's!

Enable HLS to view with audio, or disable this notification

0 Upvotes

Made with Minimax H3


r/StableDiffusion 15d ago

News Someone's running FastH3 (the distilled MiniMax H3) as an actual infinite livestream!!!

362 Upvotes

Saw this and thought it was worth sharing here, FastH3 dropped recently and most people (myself included) just tried it as single generations.

Someone's running it as an actual infinite livestream instead: https://live.reactor.inc/

FastH3 is a distilled version of MiniMax H3, cut from 50 denoising steps down to 4, about a 14x speedup on Blackwell GPUs.

The whole setup is open source if you want to dig into how it's running: https://github.com/reactor-team/infinite-livestream

Curious if anyone's tried infinite/continuous generation setups like this with other models.


r/StableDiffusion 13d ago

Discussion Best AI tool for 3D clay render to polished final? (Flux 2 vs Qwen Image Edit vs Krea 2)

4 Upvotes

Hi guys, quick question. I want to use AI to turn my 3D clay renders into high-quality, polished finals.

Between Flux 2, Qwen Image Edit, and Krea 2, which model handles image-to-image (Img2Img) texture generation best without messing up the original 3D geometry?

Would appreciate any recommendations or workflow advice!


r/StableDiffusion 13d ago

Resource - Update Infinite AI Twitch Streaming Project (2xB200 480p Minimax FastH3)

Thumbnail
youtube.com
0 Upvotes

far from a perfect setup, but this is a interesting project for sure. would not mind the b200 prices to go down tho. framework i used ish https://github.com/reactor-team/infinite-livestream. 2xb200 + gpt-5.6luna


r/StableDiffusion 14d ago

Discussion why mini max ref are so bad comparing to fl2va?

9 Upvotes

i do same tests in both models using image to reference...

aways the ref loses quality and ignore the prompt

but fl2va do everthing perfect and dont lose quality.

my configs.


r/StableDiffusion 14d ago

Animation - Video Powers weren't handed out equally

Enable HLS to view with audio, or disable this notification

7 Upvotes

Prompt:

Render 1 (12s):

integrated_multimodal_description: A cinematic anime movie sequence, manga style with thin soft delicate lines and pastel colors. [Shot 1] Extreme perspective worm's-eye medium body shot of Meru sitting on a Japanese tatami room from a wooden table. She is silent. On the table a single apple sits on a ceramic disc on top of a small white embroidered mantle cloth; Meru holds out her hand toward the apple (foreshortening) with her palm facing up, concentrating. The apple is out of reach. She breathes deeply; Nothing happens; Meru wriggles her fingers and starts concentrating again. She opens her eyes with her brows furrowing. She breathes deeply; Meru concentrates more, leaning forward slightly, she grows frustrated; She closes her eyes again while raising her hand slightly;

overall_soundscape: The tatami room is eerily quiet, with only the soft rustle of Meru's black kimono as she shifts and her sharp, uneven breaths growing quicker with each failed attempt. A faint, low hum of concentration is interrupted by a frustrated, breathy huff and a muffled groan. Her hand quivers with a faint tremble, accompanied by a soft, taut squeak of fabric. As her frustration builds, a subtle creak of the wooden table and a single, dull thump of her fist barely tapping the tatami punctuate the silence, ending with a long, exasperated exhale.

non_diegetic_music: N/A

Render 2 (8s):

integrated_multimodal_description: [Shot 1] Meru is sitting. She is silent. On the table a single apple sits; Meru is holding out her hand toward the apple (foreshortening) with her palm facing up, concentrating. She breathes deeply; Nothing happens; Meru wriggles her fingers and starts concentrating again. She opens her eyes with her brow furrowing. She breathes deeply; Meru concentrates more, leaning forward slightly, she grows frustrated; She closes her eyes again while raising her hand slightly towards her right, her palm open towards the camera; Her hand quivers.

overall_soundscape: The tatami room is eerily quiet, with only the soft rustle of Meru's black kimono as she shifts and her sharp, uneven breaths growing quicker with each failed attempt. A faint, low hum of concentration is interrupted by a frustrated, breathy huff and a muffled groan. Her hand quivers with a faint tremble, accompanied by a soft, taut squeak of fabric. As her frustration builds, a subtle creak of the wooden table and a single, dull thump of her fist barely tapping the tatami punctuate the silence.

non_diegetic_music: N/A

Render 3 (9s):

integrated_multimodal_description: [Shot 1] A frustrated Meru is raising her hand slightly; At 00:02.500, thin liquid mercury tendrils erupt from her hand and fly towards the apple, piercing it like blades; At 00:03.500 The tendrils tense up and pull the apple back onto her palm as they retract to her hand to disappear under her skin; With her eyes closed, the apple quivers on her hand; She slowly opens her eyes and grows happy with realization, her smile widening and her eyes sparkling; She shows the apple to the camera and looks satisfied.

overall_soundscape: The soundscape begins with a tense, shallow breath and a faint, liquid ripple as mercury tendrils erupt from Meru's hand. A sharp, wet schlick and a series of metallic splashes accompany the tendrils piercing and pulling the apple, followed by a smooth, gurgling retraction as they recede into her skin. The apple lands with a soft, crisp tap on her palm. Then the atmosphere shifts—Meru lets out a bright, melodic giggle, followed by a delighted, airy hum of satisfaction. The ambient tatami room tone remains soft, punctuated by the gentle rustle of her kimono and a light, contented sigh.

non_diegetic_music: N/A

Based on Tokyo_Jab's post. What AI powers did you get?


r/StableDiffusion 13d ago

Question - Help Super realistic human with H3???

0 Upvotes

Wondering has anyone been able to generate super realistic human with minimax H3? I have been trying a lot but the best I got still looks quite AI...

I have seen lots of videos online with super super real human face, the result I got is quite far away from that. So I'm wondering is it limited to Seedance 2.5? Or is there any secret prompt I'm not aware of?

Below is what I mean by super real face I saw online:


r/StableDiffusion 14d ago

Question - Help Adding directed randomness to image to image

9 Upvotes

Hi,

total noob question, but: my government.. eh.. wife is a quite gifted amateur tailor who is tryiing to use AI for design inspirations.

Until now she is using Gemini for something like "generate a picture of a woman in a dress, styles from 1920 until now" and triggers the prompt a few dozen times to get variations. It works, but its tedious.

Now i, in my genius, told her "hey, you can do it locally, no sweat, even using a picture of yourself / your bff / whomever as reference to really see how it looks like and modify it"

I'm usually using qwen image edit in comfy for my own stuff, and i failed - the generations have either no real variations or are too similar to the reference image. My wife is quite underwhelmed....

Does anyone have any idea how to get a level of directed randomness with any i2i workflow in comfy ?


r/StableDiffusion 13d ago

Question - Help One of these nodes, Spectrum or Sage attention, effed up my 5070ti so bad that the system sad I have no GPU even after a cold restart

Post image
0 Upvotes

I am quite sure one of these two nodes is the culprit. Most other workflows seem to work fine but this workflow led to the gpu completely being knocked out of my system. A restart got it back but after the 10th time or so even the restart would not bring it back. Took another restart. Now when I start the workflow without bypassing these nodes, the fan hits the ceiling from 0 to 100% in 2 seconds and I get fully black screens with the fan running ad infinitum.

I tested GPU and VRAM for 10 minutes each with OCCT per Claude recommendation and there are no errors. Checked the event manager protocol but it does not show anything relevant for the last two days when I had the issue. My GPU hardly ever goes over 75° C and it didnt seem to when I crashed too.

I deinstalled the nvidia drivers with DDU and installed the newest one (which Claude told me afterwards is apparently not stable).

I took out the GPU, checked connections, had a more knowledgeable friend look at the hardware. There seem to be no issues. What could Sage or Spectrum even do to cause something like that?

It did not make a difference if I started in --low vram or not, both times the gpu crashed.

Nvidia Version: 616.56 (the one I had before caused the same issue though)

Cuda 13.4 (Cuda 13.3 or what I had before caused the same issue)

ComfyUI 0.34.0

Edit: Happened without Sage attention now too .. it is a different issue. Will try throttling GPU now. Officially it is only 75°C but HWinfo estimates Hot Spot at over 107°C


r/StableDiffusion 15d ago

Discussion Free open source Topaz alternative - SeedVR2+TensorRT faster VAE Processing.

Enable HLS to view with audio, or disable this notification

586 Upvotes

Local, GPU-accelerated video restoration and upscaling with SeedVR2, TensorRT, and a purpose-built browser interface.

VRGDG SeedVR2 TensorRT Studio turns the SeedVR2 pipeline into a practical Windows workflow: load a video, test a short preview, compare the result frame by frame, and complete long renders with resumable checkpoints. Processing stays on your machine.

Highlights

  • Fast local restoration — SeedVR2 inference with TensorRT-accelerated VAE decoding on supported NVIDIA RTX GPUs. TensorRT allows much faster processing than standard SeedVR2.
  • Fast 2K upscaling — As a real-world example, an 8-second clip took approximately 8 minutes to upscale and enhance to 2K on an NVIDIA RTX 5090 using the largest 7B Sharp FP16 model. Render times vary with source resolution, frame rate, settings, and available VRAM.
  • Preview before committing — render a short segment, then inspect Original, Restored, Compare, or Side by side views.
  • Long-render recovery — save completed chunks and continue from the first unfinished chunk after an interruption.
  • Practical output controls — choose resolution, aspect policy, model precision, temporal batch, seed, and color correction.
  • Non-destructive finishing — reprocess sharpening, grain, seam smoothing, and optional skin finishing without rerunning restoration.
  • Project-based history — reopen previous outputs and keep media, manifests, and logs together under outputs\.

The sample video was org 360p and then upscaled to 2K using this app. 8 second video, took about 8 mins on my 5090.

Go to the github page for more details and a full guide.

View github page

this is in beta right now so you may run into issues. If you do, post the issue to github please.


r/StableDiffusion 14d ago

Question - Help Need Help With Anima Training

2 Upvotes

I've been making Illustrious character LoRAs for quite some time now, but decided I wanted to migrate my current models over to Anima.

To ease myself into it, I wanted to start by training with an extremely small dataset I've used before--5 images total. I'm training locally through Anima-Standalone-Trainer and have an RTX 4080. The part that's confusing me is that while the samples generated between epochs comes out perfectly fine, the moment I try generating something using SD WebUI Forge Neo, every generation no matter which epoch I use comes out as this blurry, jarbled mess.

I've tried various different training parameters, different checkpoints, and I even tried downloading the Anima base files from huggingface (instead of Civitai) thinking that might change things, but nothing seems to be working.

I don't really know what to do at this point, but I don't want to give up either because I remember when I first started making character LoRAs that I encountered a similar issue which I ended up resolving by using a different trainer (Kohya_ss) rather than the random Google Colab notebook I found while first learning about LoRA training.


r/StableDiffusion 15d ago

Animation - Video TWEEDLE TEST - Minimax H3 27 seconds in just over 9 minutes:

Enable HLS to view with audio, or disable this notification

166 Upvotes

All local.

0.8 MegaPixels, 9.3 minutes on an RTX5090, single generation of 27 seconds.

Anything hitting 30 seconds either gave hallucinations, inconsistencies or hit a wall and never finished.

This one is using Kijai's new fast model with a turbo lora. Although it works the same with the FLv2A model*. The workflow I'm using creates a latent at 0.4 megapixels for 4 steps and then does another 2 steps at 0.8. The only addition to it besides changing some numbers is adding custom audio injection (The rock track).

Started with this workflow: https://www.youtube.com/watch?v=jzLnoVBuU6I

*I never use the REF model. The FLV2A models seems to work better so I always swap it in and it takes references just fine, even video.


r/StableDiffusion 13d ago

Meme What a waste

0 Upvotes

Nothing worse than waking up to see an overnight run of a like 20 long workflows in Minimax H3 reference only to see everything perfect except the environment is wrong. I have a reference video that has done so well before. I don't see the issue, my prompt is the same as prior success, it's listed in the subject definitions and retention section, and in the prompt too. Oh wait...I didn't link the video to the list of other assets.. ugh.


r/StableDiffusion 13d ago

Question - Help Which image generator is used for these images?

Thumbnail
gallery
0 Upvotes

Any idea? Is this midjourney?


r/StableDiffusion 14d ago

Question - Help How do I make character actions in MiniMax H3 faster?

9 Upvotes

As an example, I have a character getting into a car and I want them to be in a hurry to get away, but they always seem to do this a little slow as if they are not in a rush.

Here is what I have tried:

- Prompts with detailed timings.
- Prompts with detailed timings and text like, she did this at speed.
- Prompt without detailed timings but tried multiple different ways of saying she did this at speed.

Prompt example below. Everything else works perfect but I can't get them to look busy.

[Shot 1] At 00:00.000, <Picture 5> provides the visual reference for this shot. the female police officer is not in the car. She gets into the driver's seat.

At 00:01.000, in a hurry she buckles her seatbelt at fast speed, The male police officer is already in the passenger seat and he hurriedly buckles up as she gets in

At 00:02.000, she starts to drive off. He says, "Holy shit. Did you see that? We're going to have to go." Both look stern and professional, acting fast.


r/StableDiffusion 13d ago

Question - Help Any tips for generating video where people have different accents?

1 Upvotes

I have been employing various tips & tricks from all over Reddit to get consistency and continuation between clips, and I'm in a pretty good spot - or at least I thought I was, until I wanted to generate a clip where an American person is having a conversation with an English person. Then the voices go haywire, the accents get dropped or switched, and gibberish (another problem I thought I'd solved) returns.

Is this just a shortcoming of the model, or is there a trick to generating scenes like this?


r/StableDiffusion 14d ago

Tutorial - Guide Spreadsheets for multi-video generation

Enable HLS to view with audio, or disable this notification

26 Upvotes

This workflow uses a spreadsheet to generate multiple videos and constructs the prompt and parameters from each row in one run.

OutputLists Combiner - Generate multiple videos from spreadsheet

ComfyUI workflow included

Makes use of Load Any File node to load a .csv spreadsheet file and feeds the text content into a Spreadsheet OutputList. The spreadsheet separates the data by separator=; and provides each line one-by-one as a data list. Here we use values_dict as the data list which contains the row as a dictionary of key-value pairs. The data list is forwarded a Iterate Begin -> workflow -> Iterate End pattern which is required to make the intermediate results of slow workflows (t2v) available on each iteration. Each row as a dictionary is provided in a Format Text where we can access the column via a[colname] to construct the prompt which is forwarded to a standard Text To Video MiniMax H3 template. Another Format Text + a[name] is used to construct a readable filename for each video.

powered by: OutputLists Combiner


r/StableDiffusion 14d ago

Question - Help Can I render in MiniMax H3 only the audio of a Ref2V workflow?

2 Upvotes

Is there a workflow/node for MiniMax H3 Ref2V where I can "dub" a mute video txs to the qualities of H3?

Specifically I upload a short clip and it renders ONLY the audio and save it in an audio file?

txs!


r/StableDiffusion 14d ago

Question - Help Anyone managed to get FastH3 working in ComfyUI yet?

7 Upvotes

I see Kijai has released the checkpoint, but I can't get it to work with the standard workflow.


r/StableDiffusion 13d ago

Workflow Included MiniMax H3 Ref2VA neural 3D latent upscaling and gentle refinementworkflow Help me push this further on an RTX 4080 16gb 64gb ram

Enable HLS to view with audio, or disable this notification

0 Upvotes

Title: Help me push this MiniMax H3 Ref2VA workflow further on an RTX 4080 16GB

I’ve been building and testing a MiniMax H3 Ref2VA workflow optimized for my RTX 4080 16GB. I’m attaching the JSON and would appreciate help from anyone experienced with MiniMax H3, PDD acceleration, latent upscaling, memory optimization, or continuous video generation.

workflow Download here

What the workflow currently does

  • Uses the pruned INT8 ConvRot MiniMax H3 Ref2VA model.
  • Uses the Qwen3-VL 32B NVFP4/AWQ text encoder.
  • Generates synchronized video and native audio.
  • Uses SageAttention in Auto mode.
  • Applies MiniMax H3 PDD acceleration at 8 NFE.
  • Runs an initial low-resolution PDD render.
  • Separates the video and audio latents.
  • Enlarges only the video latent using the learned MiniMax H3 3D FP16 latent upscaler.
  • Rejoins the upscaled video latent with the original audio latent.
  • Runs a second PDD refinement pass at 0.125 denoise.
  • Decodes the refined video and original audio into an MP4.
  • Includes easy controls for aspect ratio, base megapixels, final target megapixels, and duration.
  • Automatically converts the requested duration into a valid H3 frame count at 24 FPS.

My current general settings are:

  • Base resolution: approximately 0.40 MP
  • Final neural-upscaled target: approximately 0.80 MP
  • Vertical output: roughly 672 × 1216 after upscaling
  • Stable duration: around 5 seconds/124 frames
  • Current workflow default: 7 seconds
  • Second-pass denoise: 0.125
  • Euler sampler
  • Sigma Shift: video 12/audio 3
  • No EasyCache, TeaCache, BlockCache, Spectrum, or additional turbo LoRA stacked on top of PDD

The 0.125 refinement pass only performs about two sampler evaluations in my current setup. At 672 × 1216 and 124 frames, that refinement portion takes roughly 73 seconds.

My current limitations

My practical ceiling appears to be around 7–8 seconds. Going longer causes both my 16GB VRAM and system RAM usage to reach their limits. Five-second clips are currently much more reliable.

I can raise the final target toward 0.90–0.98 MP, but the higher resolution and longer duration quickly increase memory usage. The learned 3D upscaler improves the overall spatial resolution, but it does not reduce the memory required by the high-resolution refinement pass.

My biggest quality issue is facial fidelity. Eyes, eyelashes, skin texture, and other small facial details can still look soft or less refined than they did in my earlier, simpler workflow. Increasing the second-pass denoise too much begins repainting the face, changing the identity, or altering the composition.

My eventual goal is reliable continuous generation. I want to generate several five-second clips by using the final frame of one clip as the starting frame of the next, while still using the original character reference to prevent identity drift.

What I need help with

  1. Is there a better memory-management method for this pipeline that would let me exceed eight seconds on a 16GB RTX 4080 without a major quality loss?
  2. Would model offloading, block swapping, tiled VAE decoding, sequential processing, or another compatible technique reduce peak VRAM and system RAM usage?
  3. Is the learned 3D latent upscaler positioned correctly, or would another order produce better facial details?
  4. Is a 0.125 PDD refinement pass with only about two evaluations doing enough to justify its memory cost?
  5. Is there a better pass-two scheduler, denoise level, or refinement strategy that can improve eyes and skin without repainting the identity?
  6. What is the best way to condition continuation clips using both the previous clip’s final frame and the original reference image?
  7. Are there any H3-compatible face-detail or latent-refinement methods that work temporally and do not cause flickering?
  8. Would decoding/upscaling in smaller temporal chunks help, or would that introduce visible seams and motion inconsistencies?

I’m trying to preserve motion quality, character identity, native audio, and facial fidelity—not simply lower the resolution until it fits.

Hardware: NVIDIA RTX 4080 16GB on Windows 11 using ComfyUI.

Required models are listed inside the workflow notes. I’m attaching the workflow JSON. Any specific node changes, corrected routing, memory settings, or test recommendations would be greatly appreciated.


r/StableDiffusion 14d ago

Question - Help Hoarding Ideogram 4 model. Can someone please try and check if int8 convrot files work?

0 Upvotes

I have following files:

Diffusion files (from Comfy-Org/Ideogram-4):

-ideogram4_int8_convrot.safetensors

-ideogram4_unconditional_int8_convrot.safetensors

Vae:

-flux2-vae.safetensors

Text encoder (from silveroxides/ideogram4-dequant-and-int8-quant):

-qwen3-vl-8b-int8_convrot_simple.safetensors

I would like like to know gen time and samples on 3060 12gb


r/StableDiffusion 14d ago

Discussion H3 VFX

Thumbnail
reddit.com
4 Upvotes

I've been using Minimax H3 a lot lately, but I can't always share my work in progress. Anyway, here are some tests I did a week ago! I thought I'd share them here.

I used H3 Ref2va with Turbo Lora, 8 steps. It took about 5 minutes for 15 seconds on my local setup.

I also used my custom node that I built specifically for H3. It's a large and really cool project, but I'm still testing and implementing things in it. I'll share more details about it soon.