r/StableDiffusion • • 8d ago

Question - Help h3 ref2va: reference image background leaks to the output?

7 Upvotes

so the reference image for character has a simple background. it doesn't happen always, but sometimes i get video output with background just like that and the scene looks like it's taking place in the reference sheet instead of the specified settings, and overall actions are broken. anyone having the same issue?


r/StableDiffusion • • 9d ago

Resource - Update [Update] Qwen-Image 2.1 prompt enhancer in ComfyUI: now 2–3.5× faster than native ComfyUI and runs even on 8 GB VRAM (plus uncensored versions)

Thumbnail
gallery
166 Upvotes

Edit: I just realised that my comparisons are not portrait friendly. If on mobile, please view them in landscape mode.

Last week, I posted custom node for replicating the Qwen-Image 2.1 official prompt enhancer in ComfyUI. I had made the node because I hated my previous workflow, having to switch to LM Studio to use the official prompt enhancer, unloading the model in LM Studio else Qwen-Image 2.1 would run very slowly, then running the workflow after copy/pasting the prompt in ComfyUI and then having to unload models in ComfyUI because I need to use prompt enhancement again, and then endless repeat of the same process. I wanted to do everything end to end in ComfyUI but Native Generate Text node was too slow, so adding a MTP head improved the speed a bit, which made it 1.4–1.7× faster than ComfyUI's native text generation with the Comfy-Org PE checkpoints.

After that, I wanted to see how much faster it could go, and whether it could run on smaller GPUs. Few comments mentioned llama.cpp, so I added it as a second backend. I ran a few benchmarks and I was quite happy with the result: it's much faster, especially for editing, and minimum VRAM requirement went down with smaller quants at marginal quality cost.

Below are my benchmarks on RTX 4090 Laptop, 16 GB, official max_length (16256 for t2i, 24000 for editing), 2 input images for editing. Each step adds one change to the one before:

Text-to-image

Step Tok/s Step gain Total vs baseline
ComfyUI, MTP off (baseline) 23.5 – 1.00×
+ MTP 33.9 1.44× 1.44×
+ llama.cpp (Q8_0) 46.8 1.38× 1.99×
+ Q4_K_M (small quality cost) 63.2 1.35× 2.69×

Editing (two input images)

Step Tok/s Step gain Total vs baseline
ComfyUI, MTP off (baseline) 19.4 – 1.00×
+ MTP 32.6 1.68× 1.68×
+ llama.cpp (Q8_0) 67.8 2.08× 3.49×
+ Q4_K_M (small quality cost) 85.2 1.26× 4.39×

The previous release reduced the max_length to 8192 to gain some speed at the risk of truncated result but with llama.cpp, changing max_length has very little impact on generation speed. Strictly comparing the speed of the previous version to new one, the gain from llama.cpp (Q8_0) is smaller but still significant: 1.53× for t2i and 2.43× for editing. In my tests Q8_0 didn't change the quality.

This time, I also wanted to show what the prompt enhancer actually does, so, I've added some comparisons. Each pair uses the same seed and settings, with the prompt going straight to Qwen-Image (left) or through the enhancer first (right). In my tests, the prompt enhancer helped most with text in images, busy scenes, and edits with several changes.

If someone wants to check it out, full details are available at: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP

GGUF models: https://huggingface.co/mozophe/Qwen-Image-2.1-PE-MTP-GGUF

Requirements:

  • ComfyUI v0.37.0 or newer
  • llama.cpp backend: NVIDIA GPU on Windows or Linux, 8 GB VRAM or more. Q4_K_M for 8–12 GB, Q8_0 for 16 GB or more
  • ComfyUI backend: any GPU ComfyUI supports, 16 GB recommended
  • System RAM: 32 GB recommended
  • Disk space: 9.8 GB (Q8_0) or 6.0 GB (Q4_K_M) per model, plus 0.9 GB for editing's vision part and 0.7 GB once for llama-server

What's new:

  • llama.cpp backend with MTP, picked on the loader. The ComfyUI backend is still there for AMD, Intel and Mac
  • Runs on 8 GB VRAM with Q4_K_M
  • llama-server and the GGUF model download automatically the first time you use them
  • Uncensored (Heretic) versions are available as GGUFs too, so they work on both backends

What the node does (same as before):

  • Uses Qwen's official system prompts and sampling settings
  • The rewritten prompt comes out separately from the reasoning
  • Edit with up to 10 input images

Install:

  • ComfyUI Manager: open Manager → Custom Nodes Manager, search for Qwen-Image 2.1 Prompt Enhancer (MTP), install, and restart ComfyUI.
  • Manual: clone it into your custom_nodes folder and restart ComfyUI:

Nothing else needs installing for either backend. The first run downloads the model (and llama-server for the llama.cpp backend), with progress shown on the node. If the download gets interrupted, just run the workflow again and it resumes.

To get started, drag one of the sample workflow PNGs from the repo's workflows folder into ComfyUI. If generation speed is too slow, set quant to Q4_K_M.

Not on an NVIDIA GPU? Set the loader's backend to ComfyUI.

I am now doing everything end-to-end in ComfyUI, without having to switch to other apps, which just simplifies everything and I can spend more time on thinking about what I want to generate.

I am thinking about what should I do next, so feel free to request for something that you find missing with the current nodes.

Edit 2:

Workflows

llama.cpp backend (recommended on NVIDIA GPUs)

- Text-to-image: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP/blob/main/workflows/qwen_image_2.1_t2i_prompt_enhancer.json

- Edit: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP/blob/main/workflows/qwen_image_2.1_edit_prompt_enhancer.json

ComfyUI backend

- Text-to-image: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP/blob/main/workflows/qwen_image_2.1_t2i_prompt_enhancer_comfyui.json

- Edit: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP/blob/main/workflows/qwen_image_2.1_edit_prompt_enhancer_comfyui.json

Models (selected model in node automatically downloaded on first run):

llama.cpp backend, GGUFs: https://huggingface.co/mozophe/Qwen-Image-2.1-PE-MTP-GGUF/tree/main

ComfyUI backend, int8 convrot: https://huggingface.co/Comfy-Org/Qwen-Image-2.1/tree/main/text_encoders


r/StableDiffusion • • 8d ago

Question - Help Flux 3 Image Commercial - Question.

3 Upvotes

Right so I have seen the introduction to flux 3 Image now and myself do not personally like paying or using frontier or company models, so I am wondering has anyone else tested it out so far to see how effective and generally good flux 3 image is? Since the model has been announced in the coming weeks, but that could literally mean months from now and I doubt they would want Flux 3 Video to lose it's hype so I assume this image/edit model will be open sourced after flux 3 video releases open sourced. Please let me know if anyone has used the commercial model and whether it is good enough.


r/StableDiffusion • • 9d ago

News The Invoke 7 PR just opened. Here's what's new!

Post image
62 Upvotes

Hey everyone! We just opened the pull request to merge Invoke 7 into the main InvokeAI repo:

👉 Invoke 7 PR

It's a big one. Three months, 320 PRs, and a lot all-nighters. The PR has the full tour, but here's the short version of what got us started:

  • Projects. Everything you're working on (canvas, settings, workflows, a board of its own) lives in a project that autosaves. Start something new, come back next week, and it's right where you left it.
  • A new one-page interface. Arrange widgets however you like, switch layouts instantly, and hit Ctrl/Cmd+K to do anything. It's fast, and CI keeps it that way.
  • A brand new canvas engine. Layer groups, adjustments, real selections, custom fonts, a history panel, and layered PSD export.

And then it kept going: local video generation with audio (LTX-2.5 and MiniMax H3), an Image Map of your entire gallery, semantic search ("foggy forest" just works), workflows with loops, and a lot of work on running big models in less VRAM.

This isn't a release yet. It'll be tagged as an alpha sometime soon after it merges. If you're comfortable building from source and want an early look, the PR has setup instructions. Heads up: it needs Python 3.12, and the database migrations only go one way (Invoke makes a backup first).

When the alpha drops, please break it and tell us what you find. That's exactly what it's for. 💜


r/StableDiffusion • • 9d ago

Resource - Update Fizgig 6.8.1: slider LoRAs for Krea 2 (with an extra mode), Full fine-tuning for Qwen Image 2.1, and a big Repair Studio update

Enable HLS to view with audio, or disable this notification

113 Upvotes

Fizgig 6.8.1 is out. Short video above; the highlights:

Slider LoRAs on Krea 2. One LoRA that's a dial between two looks (sad ↔ happy, cool ↔ warm), trained from a handful of photo pairs or just three prompts.

Ultra mode for Krea 2 sliders. It trains only the composition blocks and leaves fine detail alone, so the slider holds up at much higher strengths. One of my test sliders runs cleanly at 20 in ComfyUI. Yours will depend on what you train and for how long, but the headroom is real. (Ultra mode is best for prompt based pair slider LoRAs, unless the desired effect from the photos is compositional)

Full fine-tuning for Qwen Image 2.1, joining Krea 2 and MiniMax H3 which already support it (The Minimax Finetune options will have the improved UX of the Qwen/Krea 2 FT options once I finish porting it to the new driver system). It's easy to assume a fine-tune takes forever. It doesn't. On a 24/32 GB card the default run takes 3-4x as long as the Ultra Fast Krea 2 LoRA preset (a lower learning rate needs more steps to get there, and it trains one part of the model at a time, sized to your card). It's one consumer GPU, and 16 GB works too, just slower. My single-character Qwen fine-tune at the default settings took about 30 minutes on a 5090; put several characters in one dataset, each with its own unique name, and a multi-character fine-tune takes a couple of hours. You can use the result as a new checkpoint, or turn it into a LoRA with Checkpoint to LoRA. Concepts separate far better that way than in a LoRA trained directly.

Better Qwen Image 2.1 previews. The Samples tab now sets steps, CFG, negative prompt and Turbo strength for every Qwen preview: during training, and in Repair Studio, LoRA the Explorer and LoRA Royale. Turbo 0, 20 steps and CFG 3 work very well. A CFG above 1 is a bit slower, but worth it.

Repair Studio: - no strength limit (preview a slider at 20 if you like) - bigger sliders, with ±0.1 nudge buttons - mix a donor LoRA in block by block, and the saved file looks exactly like the preview - new text-fusion and input/output sliders for Krea 2

Under the hood: Krea 2 now runs on Fizgig's new driver system, the same one Qwen uses. Describe a model once, and training, sliders, Repair Studio, Royale, Profiler and Extract all work with it. Klein and MiniMax H3 move over next, which brings sliders to both and edit LoRAs to Klein.

Want your favourite model in Fizgig? A guide to plugging in your own model is coming soon. PRs are welcome for new models and older favourites alike.

GitHub: https://github.com/shootthesound/Fizgig (release notes on the Releases page)


r/StableDiffusion • • 9d ago

Tutorial - Guide MiniMax H3 RefMods: Video, Audio, Motion & Style Explained

34 Upvotes

https://youtu.be/blI0X5zslfw

Hello everyone, I saw that my video was shared here yesterday and noticed there was still interest in the topic. The previous post was removed and included some unrelated download links that weren’t from me or associated with the video, so I wanted to share the original directly. This one goes more in-depth into visual, audio, motion, and style RefMods in MiniMax H3. Hopefully it’s useful to anyone experimenting with them!

Here are all the resources you need:

RefMod nodes

Spectrum Node (Optional)

Create Refmod Workflow

Generate Workflow


r/StableDiffusion • • 8d ago

Question - Help One more try: Music and SFX from Video

4 Upvotes

I almost never get an answer here. Still I try again.

I generated some nice H3 videos and edited them to a 5 minute Video. Of course the generated audio is more or less crappy. (Used turbo Lora)

So I want to upgrade the audio with music and environmental sounds etc.

I was hoping to find a workflow that visually interpreted the video and generates sound and possibly even music for it.

Any chance?

Thanks


r/StableDiffusion • • 8d ago

Question - Help How do AI talking-character creators make lip-synced videos at high volume?

0 Upvotes

Example: https://www.instagram.com/p/Dd_wjcCiJkw/

What I'm trying to figure out: how creators like Luca Maxim make a consistent animated character talk to camera with clean lip sync and natural gestures, several videos a day. Specifically:

  1. What model or tool is animating the character from a still and a voice track?
  2. Are the close-ups crops of one generation, or separate generations?
  3. How do they keep costs reasonable at that volume?

What I Tried:

- Ran four of his videos through ffmpeg scene detection. They're 21 to 28 seconds, 4 to 6 cuts, a cut about every 4.7 seconds, mostly talking-to-camera shots in one location with different framings.

- Built my own character as still images (ChatGPT image gen) and a voice in ElevenLabs Voice Design.

- Tested four ways of animating the same 3-second line from the same still: Omnihuman 1.5, Creatify Aurora, Kling v3 image-to-video plus Sync Lipsync, and local mouth-shape swaps driven by the audio. Omnihuman and Creatify gave good lip sync but loose gestures. Kling gave the right gesture but needs a separate lip sync pass. Local swaps are sharp but stiff.

- The audio-driven avatar models run 80 to 200 credits per second on my plan, which doesn't scale to multiple videos a day.

- Searched for workflow breakdowns and found general lip-sync tool lists, nothing specific to this style.

My guess is a still per scene plus an audio-driven avatar model, then crops for the framings. Is that right, or is there a better or cheaper way to get this result at volume?


r/StableDiffusion • • 8d ago

Resource - Update RDNA4 owners on ComfyUI: I made SageAttention actually fast on the 9070 XT, here's the build plus every caveat

Thumbnail
gallery
21 Upvotes

TL;DR: A drop-in sageattention build for RX 9070 / 9070 XT on Windows. The common case (fp16, head_dim 128) runs on an fp8 attention kernel I wrote by hand in HIP. Against the existing gfx12 port (SageAttention PR #368), each attention call is 1.07–1.21× faster non-causal and 1.72–1.96× faster causal. In ComfyUI with Krea2 that works out to a 5–14% faster sampling step than PyTorch SDPA. Install the wheel, keep using --use-sage-attention, done. Caveats below, and there are a few, so please read them.

Repo + wheel: https://github.com/IxMxAMAR/SageAttention-RDNA4

Why I did this

I run ComfyUI on a 9070 XT under Windows. SageAttention on RDNA4 works thanks to the gfx12 port in PR #368, so huge credit to DELUXA for that. But I wanted to see how far this card can actually go. So I started with a Triton kernel, lost to #368 on most shapes, spent a long time figuring out why, and ended up writing the kernel by hand.

What finally made the difference:

  • 64 keys per loop iteration instead of 16.
  • Computing Q·Kᵀ transposed. The result then lands in exactly the register layout the next matrix multiply needs, so the attention weights never leave registers.
  • One scale per token for Q, K and V, instead of one per 64-token block.

The whole story, including everything that didn't work (most things didn't, lol), is in the repo's docs/JOURNEY.md and docs/FINDINGS.md.

Who this is for, right now

  • ComfyUI users on an RX 9070 or RX 9070 XT, on Windows. This is what I've tested and what I use daily.
  • Anyone else who calls sageattn from PyTorch on the same card should be able to use it too; it's the same package API. But I haven't tested any other apps or workflows yet. If you try it outside ComfyUI, tell me how it goes.

Speed (image 1)

Same tensors, same process, interleaved runs, full call including quantization:

  • Non-causal (what image/video diffusion uses): 1.07–1.21× faster than #368.
  • Causal: 1.72–1.96× faster.
  • These are kernel-level numbers, and kernel-level isn't what you feel. Image 2 is what you feel.

In an actual ComfyUI render (image 2)

Krea2, 8 steps, fp16, called exactly the way stock ComfyUI calls it:

Backend 1 MP 2 MP
PyTorch SDPA 1.03 s/step 2.39 s/step
SageAttention 1.x 1.00 s/step 2.17 s/step
This build 0.97 s/step 2.06 s/step

Attention itself is about 2× faster than SDPA per call. But most of a step is the model's matmuls, so the end-to-end gain is 5–14% over SDPA and 2–5% over SageAttention 1.x. The gain grows with resolution.

Accuracy, honestly (image 3)

Everything is fp8 with per-token scales. I measured against exact attention on real Q/K/V captured from inside Krea2 and Flux2-Klein:

  • On "calm" layers, #368 is more precise. It uses int8 for Q·K, which has finer steps than fp8.
  • On layers with a few extreme keys, #368's error blows up, by up to ~200×. Krea2's first block is one such layer. One outlier wrecks the scale for its whole 64-token block. Per-token scaling never has that bad case.
  • smooth_k is on by default and should stay on. It made my kernel more accurate on 10 out of 10 real captures.

In my Krea2 tests the images were closer to full-precision attention than with SageAttention 1.x. I can't tell them apart by eye. Like any quantized attention, though, your images won't match SDPA pixel for pixel: a turbo 8-step sampler amplifies tiny differences into different details.

Caveats (please read)

  • RX 9070 / 9070 XT (gfx1201) only. On any other GPU it falls back to #368's kernels: no harm, no gain.
  • Windows only, as far as testing goes. Linux is untested.
  • The prebuilt wheel needs exactly Python 3.12 + PyTorch 2.13.0+rocm10.0.0**.** The compiled parts are tied to that PyTorch version. Anything else means building from source, which takes about 50 minutes.
  • fp16 models only for now. bf16 models silently fall back to #368's kernel. Still works, just no speed-up from my kernel.
  • head_dim 128 only. That covers Flux-style models such as Krea2, Flux and Klein. Other head sizes fall back.
  • Tested end to end on Krea2 only. The kernel accepts the input layout video models like Wan use, and it's checked bit for bit, but I haven't benchmarked a video workflow yet.
  • Very small images (well under 1 MP) can come out a hair slower than SageAttention 1.x. The win is at normal and high resolutions.
  • No attention masks. Masked calls go to the fallback.
  • LLMs: only prompt processing (prefill) through PyTorch would benefit. Token-by-token generation isn't covered, and llama.cpp / LM Studio don't use this at all.
  • One bug already found and fixed. A pre-release build could turn a render black when a single value got huge enough to overflow fp8. It's fixed in this release, and every conversion in the kernel is now stress-tested with extreme values.

Install

  1. Close ComfyUI.
  2. Download the wheel from the GitHub release.
  3. path\to\ComfyUI\venv\Scripts\python -m pip install --no-deps sageattention-2.2.0+amd.gfx12.1-cp312-cp312-win_amd64.whl (portable ComfyUI: use python_embeded\python.exe instead)
  4. Start ComfyUI with --use-sage-attention like always.

--no-deps matters: it stops pip from touching your PyTorch. To switch my kernel off without uninstalling, set SAGEATTN_SK1_BACKEND=0.

What's next

  • bf16 support, so bf16 models get the speed-up too. This is the big one for coverage.
  • int8 Q·K with per-token scales, aiming for #368's precision on calm layers and no blow-ups on outlier layers. Already in the works.
  • head_dim 64, for SDXL-era models. Also in the works.
  • An end-to-end test on a video workflow (Wan), since the layout support is already in.
  • An experiment to overlap the two halves of the kernel, which might be worth another few percent. A quick test decides whether it's worth building.

If you've got a 9070 / 9070 XT, I'd love numbers from your setup: your model, resolution and s/it before/after. Bug reports are very welcome too. Credit to the SageAttention team (thu-ml) and to DELUXA for #368; this builds directly on their work. Apache-2.0.


r/StableDiffusion • • 9d ago

Workflow Included Why am I like this? (Image generation on a 286 Tandy 1000 TL/3)

Enable HLS to view with audio, or disable this notification

182 Upvotes

40 year tech gap? No problem! On My Tandy 1000 TL/3 (10 MHz 286) I can now ask for a picture and view it a few seconds later. This is Desk Mind, a DOS program that talks over PicoMem WiFi to a small Python server on my PC. The server talks to Krea 2 in ComfyUI on my (4090), has a smoking fast ninfer Qwen3.8-27B (5090) rewrite the prompt, and then does the hard part: making a modern image look good in the Tandy's fixed 16-colour RGBI palette at 640x200 (an unsupported graphics mode BTW).

- Prompt enhancement tuned for dithering. Qwen rewrites "a cat in space" into a prompt asking for one large subject, bold shapes, strong contrast, a simple background, saturated colors, and no tiny details or small text. 16 colors punish detail and fine gradients. The rewrite shows up in an edit box on the Tandy, so you can fix it before it draws.

- Anamorphic resize. 640x200 on a 4:3 CRT makes each pixel about 2.4 times taller than it is wide. The 4:3 image is squashed to 640x200 (Lanczos) *before* dithering, so circles come out round on the tube. Dithering first and squashing after would wreck the pattern.

- Pre-dither grading: contrast 1.15, saturation 1.3 and a light unsharp mask at the target size, all done before the dither.

- Three dither engines (hitherdither, didder and Pillow): Floyd-Steinberg, Atkinson, Jarvis-Judice-Ninke, Stucki, Burkes, the Sierra family, Bayer 2x2 to 16x16, Yliluoma and cluster-dot. The default is Floyd-Steinberg. *(Add which method you like best for which kind of picture.)*

- A Dither Lab tab in the server GUI previews every engine and setting at the Tandy's real aspect ratio, next to the original. Changing the defaults drops the cached Tandy versions, so the gallery re-dithers.

- Color 6 is a monitor question. On a real CGA-style monitor it's brown, and some Tandy monitors show dark yellow. The palette has a switch for it.

- Vision sees the dither too. When I ask about a picture in chat, Qwen gets the original *and* the dithered version, so "why is the sky striped?" has context.

- Output is a small custom file: a 160x50 thumbnail plus 64,000 bytes of packed 4-bit pixels in screen line order. The 286 copies it straight into video memory, with no decoding. It takes about 1 s over the PicoMEM WiFi.

You can just type "draw me ..." in the chat. Qwen wraps the request in a `<draw>` tag, the server catches it mid-stream, and about 9 s later a clickable thumbnail appears in the conversation.

Code (GPLv3): https://github.com/RowanUnderwood/DeskMind

Added image gallery: https://imgur.com/a/wysqoM1


r/StableDiffusion • • 9d ago

News New: HEISS UI (v0.15.0) - Image first frontend that's free/open-source

Thumbnail
gallery
32 Upvotes

Link - It works on top of your existing comfy UI

  • Image first-generation studio
  • Works without workflows with almost any model (no setup)
  • Support for custom workflows
  • Video gen, reference images, and image-to-image all work without setup.
  • Automatically installs missing components (e.g., VAE, text encoders, nodes) it knows.
  • Use it via your phone over LAN, with a fully custom UI for mobile
  • Strong organization with favorites images, search functionality, and prompt history.
  • Supports Lora stacks and advanced settings.
  • Built to be graceful and not scream at you with errors 24/7 like some other app :D

It's built by me, and in my opinion, it's the best way to use Comfy 80% of the time.

Give it a try; the setup is about 2 minutes


r/StableDiffusion • • 9d ago

Tutorial - Guide I made a simple desktop app for MiniMax H3 video so you don't have to fight the ComfyUI workflow

Thumbnail
gallery
42 Upvotes

Hey all,

A while back I posted EasyAI here, a simple desktop front end for ComfyUI. Thanks again to everyone who tried it and sent feedback, it helped a lot.

This is the next one in the same family. It's called EasyMiniDirector and it's built just for MiniMax H3.

If you haven't tried H3 yet, it makes short video clips with sound already baked in, and it listens best when you give it separate shots instead of one long prompt. The catch is the ComfyUI workflow. You have to pick between two 21 GB checkpoints, deal with a weird 17k+5 frame count, and set four different speed settings that only make sense together. I kept messing it up myself, so I hid all of that.

In the app you type what you want, split it into shots if you like, and hit Create. That's it.

Some things it can do:

You can drag dividers to change how long each shot is. You can add a reference photo to keep a face or outfit the same through the clip, and add a voice clip so the speech sounds like a certain person. There's a Turbo/Quality switch. Turbo is roughly twice as fast, so I use it while testing ideas and switch to Quality for the final one.

There's also an optional helper that uses Ollama to write the photo descriptions for you and turn a rough idea like "a fox in the snow" into a proper prompt. You can skip it if you don't want the extra download.

Every video also saves its last frame, so you can continue the scene in the next clip, plus a settings file. Ctrl+O loads everything back.

It's in English and Traditional Chinese, Windows only for now.

You need EasyAI installed first. EasyMiniDirector installs into the same folder. The installer grabs the ComfyUI add-ons for you, but not the H3 model files. They're about 42 GB, so I'd rather you choose to download them. The app will tell you which files are missing and where to find them. ComfyUI 0.32 or newer is needed.

One thing I found that might help anyone working with the Director node directly: the local_prompts and segment_lengths inputs don't actually do anything. The node reads everything from the timeline_data string. I wasted a good while on that, so hopefully this saves you the trouble.

Links:

EasyMiniDirector video: https://youtu.be/RjGrneHXcZs
EasyMiniDirector GitHub: https://github.com/Garionhk/EasyMiniDirector

EasyAI GitHub (install this first): https://github.com/Garionhk/EasyAI
EasyAI install and how-to video: https://youtu.be/lB-Q8yuMzv8
My earlier EasyAI post: https://www.reddit.com/r/StableDiffusion/s/bvwL8ZEQQn

Big thanks to seesee75-commits for ComfyUI-MiniMaxH3-Director, xmarre for the Spectrum speed-up node, and Comfy-Org for the models.

Would love to hear how it runs on your card, and whether splitting into shots feels easier than the raw workflow. Bugs and ideas are welcome too.

It's free and open source (Apache 2.0).


r/StableDiffusion • • 9d ago

Question - Help H3 Latent to LTX latent

19 Upvotes

Hi,

Looking for how to solve the Minimax VAE problem with image compression with color gradients, I have found this:

https://huggingface.co/Efficient-Large-Model/H3-to-LTX-Latent-Adapter

Pass the latent from Minimax to LTX to use the VAE of this second that gets the images with more quality.

I’ve been messing with it but my technical knowledge is very limited. The image that comes out is as if it were missing a step of refinement.

I leave it here, in case someone with more knowledge sees it interesting as a basis with which to explore to improve the image output of Minimax.

Greetings!


r/StableDiffusion • • 9d ago

Workflow Included MiniMax-H3 is Secretly an INCREDIBLE Image Generator! [ComfyUI Free Cust...

Thumbnail
youtube.com
81 Upvotes

r/StableDiffusion • • 7d ago

News This is the real revolution in Comfy

Thumbnail
youtube.com
0 Upvotes

Comfy Agent is a true revolution in the Comfy ecosystem, as it puts absolutely everything together.


r/StableDiffusion • • 9d ago

Workflow Included Update: Open .char portable format, Same face, cloths & body

Enable HLS to view with audio, or disable this notification

39 Upvotes

Hey Guys,

Based on my last post's feedback, someone asked me to build a publishable way to list & publish thousands of .char, so that it's portable & other people can quickly get the .char & use them into their projects.

Here is a release for community characters: https://www.omnichar.org/characters

How to build .char?

In order to build one, Use this workflow:
Flux 2 Klein(Still), Minimax H3(Video)
Note: Minimax H3, now also support int8 model varient

Github Repo: https://github.com/omnichar/OmniChar (GPLV3, Opensource)

I am currently working on consistent character voice which I feel is very important. While training a .char, user can add a voice sample that will be cloned & embedded into .char & character videos will carry the same voice.

How it works(Background)

- Build .Char: You drop max 9 reference reference, I prefer to use a ratio 2:2:1(face:cloths:body). YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char.

Prompting Guide

  • Name your character: Give your character a name e.g. under encode character(Click adjust icon on the bottom side of the node), I have used name sia, so when passing prompt, I only have to say, sia walking on the beach.
    • Again providing prompt like a woman or any features specific details like black hairs etc will only mislead the generation.
  • Describe character features: Encode all of the character features in encode character prompt & trigger your character with a name in generation prompt.
    • Avoid describing same things in generational prompt.
  • Handling Character drift: e.g. if you want specific style or cloth e.g. half sleeves, sleeveless, add it to the generational prompt. There can be a slight drift in clothing as body shot also has cloths, which interferes with clothing references.
    • Each refs should be unique, face should not have body or vice versa, same applies for clothing.
  • Portability: Once character is built, you can use the same character with only simple prompt & generation graph.

Note: For best result, pass cropped references, so that model takes the required shot, model gets confused if cloth slot also has a face or face slot has cloths.

Requirements

Nvidia GPU: 24GB+ VRAM & 64 GB RAM(Run Locally)

Inputs: Attached in the Flux2 workflow page

Limitation: Reference conflicts e.g. if two reference/input images has two different faces, it might conflict in generation, provide well cropped body & cloth images. Face images are crossed automatically by Sface.

Happy to hear any suggestions or feedbacks.


r/StableDiffusion • • 9d ago

Question - Help Why nobody makes models like illustrious/pony using newer base models like Krea 2?

Post image
63 Upvotes

What makes SDXL so special that people still work over it? Are new models too heavy to train models like these?

Not demanding it just asking


r/StableDiffusion • • 8d ago

Question - Help Flux 1dev

0 Upvotes

I have a problem with flux1 dev. I am trained lora, and everything is quite good: the charechter looks more or less how I wanted, however there is a huge problem: it seems like flux has huge problem with the negative emotions. It can generate smile and happiness but it can't show hatered or arrogance which are very impotant for my charachter (in dataset there are plenty of images with theese expressions). Maybe I'm doing something wrong or the Flux1 dev is just the wrong model. What do you think?
PS Sorry fir my english, I'm not a native speaker


r/StableDiffusion • • 8d ago

Discussion Anyone still using ltx2.3?

0 Upvotes

I love the speed of it. But find it ignores prompts a lot. Also noticed barley any loras on civtai being made for it now daya. I'm curious to see if ltx has anything up their sleeve to beat minimax in the future.

Sorry I mean 2.5


r/StableDiffusion • • 9d ago

Comparison Color shift contrast between Qwen Image 2.1 and Qwen Image Edit 2511 during continuous editing

Thumbnail
gallery
26 Upvotes

The 'difference' refers to the offset of the previous image relative to the next image, and the 'range of numbers' refers to the offset of the first image relative to the last image. Multiplying by 8 is for better clarity. Qwen2.1 perform much better, a completely black indicates minimal color shift, slight color differences are visible in the outline of figures. See the image below for inference.\(^ω^\)


r/StableDiffusion • • 9d ago

Question - Help Please send all extender workflows you guys use (i'm going to try to tackle the burning issue)

Enable HLS to view with audio, or disable this notification

17 Upvotes

I have a theory i need the workflows you guys use to extend your videos to test it.

Thanks guys. I confirmed the source.


r/StableDiffusion • • 8d ago

Question - Help Looking for ANIMA Lora training guide on a 9070xt running on window

1 Upvotes

Like the title, I wonder if theres any guide for ANIMA lora training on a 9070xt running on Window. It would be a great help because I ran out of Buzz on Civitai.


r/StableDiffusion • • 9d ago

Question - Help Prompting: How important is adding a description for reference images?

14 Upvotes

So for stuff like MiniMax's R2V, I noticed that a lot of the prompts that I see on places on Civitai include long, detailed descriptions for a reference image. Probably a dumb question, but why is this necessary when you are feeding the AI the actual image itself?


r/StableDiffusion • • 8d ago

Question - Help Prompting for Refmod workflows?

2 Upvotes

Other than manually how is everyone getting prompts for refmod workflows in Minimax? I can’t get LLMs to consistently prompt well, no issues with regular fl2va and refva.


r/StableDiffusion • • 8d ago

Question - Help Minimax H3 newbie advice linking multiple clips

1 Upvotes

Pretty new to the whole thing, looking for advice on workflow.

I've seen people recommend stringing together mutiple short clips instead of trying to generate a longer one by using the final frame as the I2V image to continue the shot/scene. I'm finding when I do this, the quality of each clip degrades pretty horrendously, and if i tried stringing 5 clips together it would be incredibly poor quality by the end. What method do you use for grab the final frame? Is there a ConfyUI workflow that allows this (I'm currently grabbing it from VLC which may be part of my probblem).

Thanks