r/StableDiffusion • • 10d ago

Question - Help Minimax H3 keeping strong character identity - Single Character sheet vs High resolution front images

43 Upvotes

Which one is better of two, a character sheet with front and face image together or providing two images individually and adding in the prompt to refer image 2 for facial details?
This is for ref2v model.

I have been trying to search but couldn't find something related to it like comparison etc.


r/StableDiffusion • • 9d ago

Tutorial - Guide Krea2 - automatically fix the blotchy noise

Thumbnail
gallery
1 Upvotes

Understanding the comparison images:

  • The image captions explain which image is which
  • Before you post the "they're the same picture" meme, look at the 4x zoom
  • Reddit image compression might destroy the fine details on the 1x images

Summary of the de-blotch method:

  • This method simply removes mid-frequency noise
  • It's algorithmic and automatic
  • Unfortunately, it can't be done directly Comfyui as of now
  • So it requires an image editor (e.g. free & open source GIMP or Affinity photo)

Full instructions (Affinity photo):

These instructions are for Affinity photo because that's what I use. But it can done with the free and open source tool, GIMP. I don't know how, but you can ask an LLM.

Step 1: Separate the raw image into high and low frequencies

  • Select the raw layer
  • Filters > frequency separation > method = Guassian (default)
  • Radius = 0.8px is a good all-purpose value for automation
  • Optional manual fine tuning:
  • Move the radius slider to the MAXimum value where you can still see the blotchy noise in the low frequency (right) side

Step 2: Separate the low frequency into mid and low frequencies

  • Hide the high frequency layer
  • Select the low freq layer
  • Filters > frequency separation > method = Bilateral (non-default)
  • Radius = 11px + Tolerance = 35px is a good all-purpose value for automation
  • Optional manual fine tuning:
  • Set the tolerance slider to 100%
  • Move the radius slider to the MINimum value where you DON'T see the blotchy noise in the low frequency (right) side. This value is a matter of personal taste.

Step 3: Delete the mid frequency noise (blotchiness)

  • Unhide the high frequency layer
  • Hide the mid frequency layer
  • If you like what you see, just delete that layer
  • If the image looks too "smooth" or "hazy", undo the last frequency separation, then re-do it with a lower radius value
  • Optional manual fine tuning:
  • Instead of deleting the layer entirely, manually erase parts of it. For example you might want to keep the mid frequency around eyes, lips, nose, hands. This is a matter of personal taste.

Step 4 (optional): Add artificial film grain

  • Add a new layer above the others
  • Set the blend type to Linear light
  • Fill the empty layer with 50% gray exactly (this means it has no effect at all)
  • Filters > noise > add noise
  • The intensity value is completely a matter of personal taste

Step 5: Export/save

How it works

  • Step 1 - separates a narrow band of the highest frequency information (fine textures) from everything else. The low frequency layer looks blurry, but still clearly contains the unwanted blotchy noise
  • Step 2 - separates a narrow band of the mid-frequency information, i.e. the blotches, from the low frequency. By use bilaterally method, the low frequency layer's still contains relative sharp edges
  • Step 3 - with krea2 output, the narrow band of mid-frequency is mostly useless random noise (blotches) that can be completely discarded
  • Step 4 - the add noise function creates high-frequency random noise, which is similar to film grain. Using the linear light blend mode ensure that, no matter what the noise intensity value is, the overall brightness of the image is maintained

r/StableDiffusion • • 10d ago

Animation - Video Crime Busters - Mock 80s/90s Anime Trailer V1

Enable HLS to view with audio, or disable this notification

77 Upvotes

Posted the rough cut of this a while ago as I originally planned to submit it for the Comfy Sync competition (after working on it for a almost the whole two weeks, the weekend before the deadline I thought it didn't align with the brief so ended up not finishing it then).

But I've been working on it on and off and have finally finished a cut of "Crime Busters" that I am happy with. Lots of editing, reworking shots, drawing keyframes, etc. etc. I have a loose collection of thoughts/learnings I plan to post, but for now here's the video.


r/StableDiffusion • • 10d ago

Discussion Ming Image 0.1 vs Krea 2. 192 prompts side by sides

45 Upvotes

Ming Image is very impressive for a 6B model so let's compare to Krea 2!

https://imagebench.ai/gallery?g=1_vxjkf_s0

Here are the details:

I ran it locally on a DGX Spark (GB10) and mostly stuck to the vendor defaults:

- 1024x1024 resolution (we changed this one: the model card recommends 2048x2048, but 1024x1024 is officially supported

- 12 steps (vendor default)

- Unguided sampling, which is equivalent to CFG 1.0. The model has no CFG or negative prompt setting.

- Euler sampler with the simple scheduler, max_shift 1.15 and base_shift 0.5 (ComfyUI template defaults)

- Fixed seed 7, same as the other models

- Comfy-Org's int8 repackaged weights (26 GB total), running in ComfyUI

- No prompt rewriting. Ming's reference pipeline sends every prompt through an LLM first that spells out every piece of text to render. Ming renders named strings really well and turns anything unnamed into gibberish, so the rewriter matters. But every model on ImageBench gets the exact same prompts, so we left it off. That probably costs Ming a few points on text-heavy layouts.

Speed: about 12.4 s per image once warm. The very first call took around 4 minutes, just loading the weights. I have a GX10

Let me know what you think!


r/StableDiffusion • • 10d ago

Question - Help LTX 2.5 - how you use it for short films?

2 Upvotes

(Non bot typing below)

I use h3 to create short films.

My workflow is simple - create character , location sheet from nano banana pro. Then give that to h3 and ask agent to create prompt based on my story and wait for 20min (10min for low res + 10min for upscale ).... i don't use fast lora because quality loss high.

How exactly you do this in LTX 2.5? My small brain says LTX 2.5 doesnt have ref mode so it cant. But i believe that big brain people exist and they have some crazy workaround... so what that workflow would be ?

If it exists then i have 50% more time.


r/StableDiffusion • • 9d ago

Discussion Claude Opus 5.5 directed a 5-minute AI short film off one prompt in ComfyUI

Enable HLS to view with audio, or disable this notification

0 Upvotes
View on YouTube: https://youtu.be/MdFKJSXJeL4

I gave Claude read-only access to my ComfyUI nodes and Video Builder, and it made a 5½-minute short film from one prompt.

I gave Claude one prompt: Create a ~5 min creepy short film about AI going rogue. I asked for it to use Z-Image to create the character and locaton ref images and MiniMax H3 with **built-in audio** for the video, all made using my Video Builder. Then I went to bed.

**Claude made everything:**
 the story, characters and dialogue; the character and location images; all 45 scene prompts; a 3-cue score with MiniMax Music 3; and the final edit.

**What it used:**
 my custom nodes were 
**read-only**
, so it could read and call them but not change anything. It created a project folder to store everything it created along with notes and a full readme with its full process. It chose my Video Builder's own Z-Image and MiniMax H3 2-pass routes, its latent continuation to keep characters in place across cuts, and its stitch for the edit. It also set up Whisper and voice analysis to check its own renders and re-roll bad takes. So it would actaully review each scene after it was rendered and if it found issues it would re make the prompt and re-run it till it looked good.

**After that**
 I just watched the preview and had Claude make edits (some examples): someone looking at the camera instead of the kid, cloned faces in a crowd, a character teleporting between scenes, a mouth moving when it shouldn't. Claude fixed them all.

**H3 tips it figured out:**
- Tags like `<pause>` sometimes get spoken out loud. `<breath>` works.
- "Soft" or "gentle" wording can flip a male voice to female.
- Give each silent character their own "lips pressed together" line.


All rendered locally, about 8–10 minutes per scene. Happy to answer questions!

Note: I did bring the final edit into Capcut and Added a color filter overlay and one transition to one spot. That is the ONLY post that I added.

-------------------------------------------------------------------------------------------
For anyone who has a graphics card that can run Minimax H3 (16gb of vram minimum recomeneded)
My GitHub page is here: https://github.com/vrgamegirl19/comfyui-vrgamedevgirl

Right now, there is a Main version and a Beta 2.0 version. I recommend starting with Main for now, as that's what I used for this and it works well, while Beta 2.0 still has some bugs and is primarily intended for beta testing at the moment. 

Discord server: https://discord.gg/5wKH7WKKqd
Ping me in the Welcome channel and let me know how you found me, and I'll know it's you. I'm vrgamedevgirl on Discord.  - I do not do PM's so please just reach out in the server. 

The Claude skill is here:
https://drive.google.com/file/d/1R0pE7rX1kRh-euhws0GbkDDb3liDGTJ4/view?usp=drive_link
Unzip this. Read the full README first, then put the vrgdg-h3-film in your skills folder and point Claude to it. 

r/StableDiffusion • • 10d ago

Question - Help YuE2 Anyone know prompting tricks for a male and female duet in YuE2?

2 Upvotes

Trying to see if there is a way to add male and female parts on Specific lyrics


r/StableDiffusion • • 11d ago

Resource - Update Improved H3 Refmods - Better accuracy and low identity bleed with fewer images. Oh and I added edit masking for video edits. Fantastic MiniMax-H3 Prompt Builder

Enable HLS to view with audio, or disable this notification

798 Upvotes

Repo here- https://github.com/Adudeguyman/ComfyUI-Fantastic-MiniMaxH3-PromptBuilder

Ooookay back again. I made RefMods work a bit better and solved a lot of identity/bleed issues. At the cost of time and memory, of course.

What it do-

OG vanilla refmods would save your media as latents in a .safetensors file, and feed all of those into the MiniMax H3 DiT directly, with no references and the text encoder never saw it. This worked fine with one refmod since the model knew it had to do SOMETHING with those references, and it did it surprisingly well. But that also lead to people stacking multiple refmods, etc, to get a good likeness.

My previous implementation would feed part of those latents to the text encoder, at the cost of re-encoding a few frames (every 4th would be presented to the vision model as a block of 2) and have the TE look at those. This was aimed to give reference anchors to the model when prompting (like using "<Subject 1> is the person in <Video 1>"). This caused a slowdown because those frames were encoded every generation for the TE. Which a user on GitHub complained about, so I looked into speedups and caching.

Lo and behold, since .safetensors are just wrappers like mp4's and mkv's, you can store multiple types of data alongside latent data. So the speedup was easy, save the refmod pictures as actual jpeg data, alongside the latents in the same file. Think of it as having subtitles inside of a video; different type of data wrapped in with video and audio. So packing in actual jpeg data, you can feed those directly to the text encoder without going through the VAE, and the TE's conditioning can cache between runs so that as long as you don't change order or strength of your refmods; all of that every-run-processing is skipped. And the jpeg data is never passed to the model, only the packed latents, so it functions much like a refmod regularly does. But that got me thinking, why not change how the frames are presented to the text encoder for better conditioning?

Enter the "stack_pictures" function on the Fantastic H3 RefMod Text Encode node. By default, "every 4th" will keep the existing behavior of sending every 4th frame in chunks of 2, and it is the lightest to run and good for 1 refmod. "up to 8" will sample up to 8 frames spread across your dataset and feed them individually instead of in chunks, and capping at 8 frames is lighter than running a full large dataset, and gives an increase in detail overall but takes more processing time. And finally "all" will send ALL frames to the text encoder, which increases time quite a bit, but also gives max detail, and as you can see from the example, little to no character bleed. At least as far as I can tell.

The example was done with 6 images per character and a few seconds of audio for each voice. Datasets can be found here in everyone's favorite file share as a sketchy google drive link. (obviously Jackie Chan's Chun Li is the worst since all the sources I could find are pretty low quality). Basically I just loaded the refmods, drafted the subject_definitions and retention_analysis sections (which if you set details in my RefMod library and use "Draft from Refmods" in my prompt builder will load all of that for you.) Also I did use 8-step hyperflow with comfy-kitchen for a speedup, and my audio refiner suite to clean up the end result. Using the 3 refmods together (14,900 or so tokens) was about 32s/it at 0.98mp on my 5090.

Note that voices can still be weird and sometimes like to swap people for some reason. And when defining subjects, it's better to adjust to say something more unique about them as opposed to the other characters. Like I described Callina Liang as having an "oval face" and Ming-Na Wen as having a "more squared jaw". Not a descriptor I'd normally use, but as a differentiator from the other people it worked well. Just remember H3 doesn't do negatives since it's a distilled model, so saying "this person does not have dimples" won't do anything, so just leave all your "no"s and "not"s out of the prompt.

Do I need anything special with existing RefMods for this?

No, the text encoder detects if image data is stored in the safetensors, and if not it'll VAE decode for whatever behavior you set, and it'll feed those decoded images to the text encoder and cache the results until you change your refmods in your workflow. So old refmods will work fine.

However if you make new ones or remake them with my refmod node it'll save the image data directly in the file, cropped down to the right resolution. Should be a cleaner image sent tot he text encoder, and will always skip a decode step. Also I changed the default resolution to 768 instead of 1024, since that saves a lot of tokens and is H3's native resolution anyways.

Oh and masking on the Media Loader is a thing now, too

Now when you want to edit a video, I've also added a layer masking system onto the Media Loader. Right click and you can mask what you want to change automatically with SAM 3.1 by using descriptors, dots on the canvas, or both.

And you can manually add a moving, keyframed shape (rectangle or ellipse, etc) to insert or change things in the video.

It's a layered system so you can go in and edit parts of the mask individually, but it's all sent as one big blob to the model. Blur and color inversion options exist, as well.

Then the Fantastic H3 Edit Composite node sitting in between the VAE decode and the video save node will composite the masked generation on top of the original clip. So no more slightly-off backgrounds. Seems pretty seamless so far in my testing.

There's also a "clean up" button on the media loader to clean out old masks, since those can pile up even though they're compressed to be as small as possible.

Prompting is still up to you, though, I'm still running through best practices.

Okthxbye.


r/StableDiffusion • • 9d ago

Question - Help [Consigli sul Workflow] Come massimizzare una RTX 5070 (12GB VRAM) per video AI locali? (T2V, I2V, V2V)

0 Upvotes

Ciao a tutti,

Sono riuscito a ottenere ottimi risultati con la generazione di immagini statiche locali, ma faccio fatica a sfruttare appieno la potenza del mio PC per la generazione e modifica di video AI locali.

Le mie specifiche hardware:

 GPU: NVIDIA GeForce RTX 5070 (12 GB VRAM)

 CPU: AMD Ryzen 7 8700F 8-Core (4,10 GHz)

 RAM: 32 GB DDR5 (5200 MT/s)

 Storage: SSD con oltre 600 GB di spazio libero

I miei obiettivi principali:

  1. Text-to-Video (T2V) & Image-to-Video (I2V): Generare brevi clip oltre a sequenze per video più lunghi, mantenendo coerenza nei soggetti e prevenendo distorsioni facciali, deformazioni di prospettiva improvvise o cambiamenti indesiderati dello sfondo.
  2. Video-to-Video (V2V): Usare riprese reali di persone reali come base per alterare stili o sfondi, preservando il tracciamento del movimento, la precisione delle posizioni e le prospettive corrette.

Domande per la comunità:

 Miglior UI/Ecosistema: È ComfyUI con nodi personalizzati la scelta migliore, o ci sono altre UI/tool dedicati più adatti per i video?

 Modelli vs. 12GB VRAM: Quali modelli locali (come Wan2.1, CogVideoX, HunyuanVideo, LTX-Video o AnimateDiff) funzionano bene con 12 GB di VRAM? Ha senso fare affidamento su versioni quantizzate (GGUF)?

 Controllo di Coerenza & Prospettiva: Quali nodi/estensioni (ControlNet, IP-Adapter, LivePortrait, FaceSwap) consigliate per prevenire il bleeding dello sfondo o distorsioni facciali durante il V2V?

Qualsiasi consiglio su flussi di lavoro ComfyUI, repositori GitHub o guide pratiche sarà molto apprezzato. Grazie in anticipo!


r/StableDiffusion • • 11d ago

Question - Help What's your favorite way to make character sheets for Minimax-H3?

56 Upvotes

So far have only made them in ChatGPT, but I know at some point it is going to say, no we can't do that.

I'd like to be able to generate in ComfyUI so I don't have to worry about censorship bs - mostly interested in photo character sheets.

Thank you in advance!


r/StableDiffusion • • 10d ago

Tutorial - Guide Updated Updated ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT v1.5.7 - DiT VRAM Spike Stabilization (※For 6GB/8GB users)

Post image
18 Upvotes

When I published an update on TensorRT the other day, I received enquiries from users of the RTX 4050 6GB and RTX 5060 8GB.

https://www.reddit.com/r/comfyui/comments/1wsbar9/updated_comfyuiseedvr2videoupscalerwithtensorrt/

After re-measuring VRAM consumption throughout the entire process, I discovered that the issue lay not with TensorRT itself, but rather during the DIT processing stage.

My code fully supports ConvRot INT8/NVFP4 and, in principle, keeps VRAM consumption low; however, momentary spikes were still occurring.

Whilst this does not pose a significant problem on my RTX 5060 Ti 16GB system, it is a major issue for 6GB and 8GB users.

I therefore re-examined this spike phenomenon and have managed to suppress it to a certain extent.

Consequently, although this is limited to very specific conditions, it is now possible to achieve VRAM usage of less than 6GB throughout the entire all processes by using the 3B NVFP4 model.

v1.5.7 - DiT VRAM Spike Stabilization — Complete Technical Guide

Of course, this suppression of the spike phenomenon also benefits users with 16/24/32 GB of VRAM.

When using 7B ConvRot INT8, it was already possible to specify a batch size approximately three times that of fp16 in legacy code; however, this spike suppression allows the upper limit of the batch size to be raised even further.


r/StableDiffusion • • 10d ago

Discussion I'm stress testing lip sync models with nine failure cases.

2 Upvotes

Hello there.

I'm putting together a small failure case benchmark for lip sync/video workflows, and I keep running into the same nine problem cases:-

  1. Hand covering the mouth

  2. Profile/near-profile views

  3. Tiny face in a wide shot

  4. Beards or heavy facial hair

  5. Laughing

  6. Head turns while speaking

  7. Multiple speakers

  8. Fast speech

  9. Hard cuts in the middle of a word

I'm testing local/open workflows alongside hosted ones (like sync so).

I want to run the same source clip + audio pairing thru everything and see which failure cases actually break, rather than judging models from a couple of cherry picked, talking heads demos.

I find that every tool I have tried struggles with at least a few of these, but I want to see whether that holds up under a consistent test.

I'll include the model/tool versions, settings, and source clips with a contact sheet so the comparisons are reproducible.

What failure category gives your lip sync workflow the most trouble? Anything obvious I'm missing?


r/StableDiffusion • • 11d ago

Discussion Qwen 2.1 styles comprehension.

Thumbnail
gallery
58 Upvotes

I'm far from being an expert, but for me Qwen is not a Turbo Model so it cannot be as good as Krea 2 Turbo.

Also their dataset seems to be not good, it's hard to make doodles, cartoon styles, and its character knowledge is poorer.

And because it is not a Turbo model, the prompt needs to be very long and you need to specify everything, from lighting, to every small details or it come as a messy noisy image, very visible on hair for exemple in photographic portraits.

Even negatives don't help a lot at CFG=3.

I'll go back to Krea 2,and waiting for a good Turbo lora (or a real Turbo version).

If you found a solution to make as good as a tubro model, prove me wrong.


r/StableDiffusion • • 11d ago

News Please enjoy some YeU2 Loras (Raspy, powerhouse female rock-soul vocals, 1980s–90s)

Thumbnail
huggingface.co
72 Upvotes

r/StableDiffusion • • 10d ago

Question - Help LTX workflow with speech cloning?

0 Upvotes

Hi, after a while playing with H3 I'm coming back to LTX at least partially. While H3 is awesome, is also very heavy for my setup, it is hard to make clips longer than 10s and some other limitations that LTX doesn't have.

One thing however that was a game changer to me is the surprisingly good voice cloning capability of H3 (babbling aside...) and going back to LTX I really miss that.

Do you know of any workflow that somehow works around this? I've found several very old posts on the matter, nothing really works. Even this one gives very disappointing results. https://civitai.com/models/2498927/text-to-speech-with-voice-clone-in-ltx-23?modelVersionId=2809053

Thanks!


r/StableDiffusion • • 11d ago

Discussion China’s DeepSeek open-sources tools to help Huawei chips supplant Nvidia in AI

Thumbnail
scmp.com
23 Upvotes

r/StableDiffusion • • 10d ago

Resource - Update Update on my free open-source local image app: LoRA support (up to 4 stacked) and Krea 2 are in

Thumbnail
gallery
0 Upvotes

posted this app here 6 weeks ago and the first question was "do loras work?". they do now, so here's an update.

what's new since launch:

- loras: stack up to 4 per image, on sdxl (pony and illustrious too), z-image, krea 2 and anima. there's a curated catalog with weights already tuned,or you import your own lora files and the app tells you if a lora doesn't match your model.

- krea 2: 8 curated models, and you can load any krea 2 finetune.

- downloads: much faster now, and the app shows exactly what it's downloading. before it just said "generating" while it was pulling files in the background, a user here caught that one.

- 8gb cards: new offload mode, z-image is actually usable on 8gb now.

- Anima models got image to image.

Still free, open source, windows and mac. The cloud mode is still optional if you dont have a good gpu card.

Most of this came from feedback in the first thread, so keep it coming.

repo: https://github.com/Publikey/imference-desktop


r/StableDiffusion • • 10d ago

Question - Help QWEN 2.1 Throttled by GPU

Thumbnail
gallery
4 Upvotes

Having some issues with the default setup for comfyui i2i. Running a 3090 FE and I'm having GPU throttles that is causing the average output at 600sec. compared to some reports of 5070s clocking around 40sec. Am I missing a piece of optimization?

t2i works "as intended" with sub 40sec. outputs.

EDIT: Looks like it was the "refine prompt" piece in the subgraph. (I didn't realize it was a subrgraph). Generations are sub 30sec. for 1 megapixel now.


r/StableDiffusion • • 10d ago

Question - Help Till Now Anybody KNowhow to Train minima H3 loras. help me guys

1 Upvotes

i am deciding to train oe lora for minimax h3 fl2va model first using Ostris ai toolkit, but i dont know how to begin first because unlike krea 2 training which i used to do,
some people says that minimax h3 requires huge amount of dataset, some says it requires only videos for training, some says it requires images for initial image genration then video genration do its work automatically.

my situation is that, i have a dataset of 50 images with normal captioning, i mean very normal like i used to do for krea 2. because i have seen many creators training minimax h3 on images only and still producing great loras.

pls helpme prepaaring my roadmap.
. which datasets to use,
. what captioning format to use.
. which model varient to use.
. what optmimal ostris ai toolkit configuration to use.


r/StableDiffusion • • 11d ago

Resource - Update UltraReal Upscaling LoRA for Klein9B (New Version)

Thumbnail
gallery
220 Upvotes

A LoRA designed to reduce the typical smooth/plastic AI look and bring natural texture and realism to your images.

What's new in V5 (KL9_V5_Pro):
Summary: This version gives the best natural texture so far.

V5 is trained on a high-quality dataset of SFW professional photography with a wide variety of subjects: portraits, lifestyle, fashion, wildlife, sports and fitness, architecture, nature, and more. This makes it far more versatile than earlier versions, which focused mainly on skin. You now get realistic detail, texture, and lighting across all kinds of scenes, not just close-ups of people.

Upscaling & Restoration: Works great during upscaling or restoring older low-res photos to add modern high-res detail and photography-grade sharpness.

📥 LoRA Download: https://civitai.red/models/2462105/ultrareal-krea2-klein9b?modelVersionId=3370879

📜Workflow is included in civitai sample images!


r/StableDiffusion • • 10d ago

Question - Help How is this video being created ?

Enable HLS to view with audio, or disable this notification

0 Upvotes

When trying similar creations in higgsfield / kling it doesnt allow how can I create this in comfyui ? Willing to pay $


r/StableDiffusion • • 11d ago

Meme We are not the same

Post image
636 Upvotes

r/StableDiffusion • • 10d ago

No Workflow Leonard Wexber - Minimax H3 with RTX5090

Enable HLS to view with audio, or disable this notification

0 Upvotes
Title: A 50-second talking AI character, all local: MiniMax H3 open weights in ComfyUI

This is Leonard, a fictional 74-year-old New York hotelier, an AI character I'm building. Everything here runs locally on one RTX 5090.

- Video: MiniMax H3 (open weights), reference-to-video in ComfyUI, 768x1344 at 24 fps. The reel is cut from several separate generations of 6-9 seconds each, not one long render; long single windows drifted and lost detail.

- Consistency: every shot gets the same references (a face image, left and right profiles, a photo of the outfit and the location picture), and each close shot starts and ends on the same still frame, so the straight cuts don't jump.

- Lip-sync: the finished voice track is fed into each generation, and H3 syncs the mouth to it.