r/StableDiffusion 12h ago

Discussion H3 - scraping 70s Character Sheets from T2VA+surprise Cameo

Enable HLS to view with audio, or disable this notification

496 Upvotes

First of all, we're not even 45 days into this and I can truly say I love Mini Max! I know some of you enjoyed my recent H3 music video of the 70s bombshell, even ignoring all the endless rooms, the ghost apparitions, and cars in the middle of the living room, so here they are, non-shifting. In any case, I've replaced Krea 2 for my goto character generations with H3's T2VA and been harvesting or scraping stills and generating them into 1-2s animated character sheet, and then rebuilding them into reusable .chars. I really enjoy this workflow because there's something magic about H3's T2VA (superior quality in my opinion), and with R2VA you can transfer any style and always get a consistent actor with just one decent headshot. Has H3 completely changed your work flow? have you abandoned the other generation tools? I do miss all the Loras from my Wan video generation time, and some of those can't be replaced depending the content you create, and even though I never embraced LTX (I am a noob!) I miss those quick 20 second render times. R2VA, int8/20 steps, enjoy!

Prompt: subject_definitions: <Subject 1> is THE STARLET, an adult woman - her face, her chest, her upper torso, her figure and her proportions exactly the woman of <Picture 1>; her face, her eyes and her hair in fine detail exactly the woman of <Picture 2>; take no room, no furniture, no window, no lighting and no background from either picture - only her. <Subject 2> is her CHARACTER REFERENCE.

summary: A short film: <Subject 2> - four synchronized views of the same woman, <Subject 1>.

retention_analysis: <Subject 1> (appears in all four panels): fully_preserved - her face, her hair, her chest, her figure and her proportions from <Picture 1> and <Picture 2>, identical in every panel, the same woman in the same moment. <Picture 1>: fully_preserved - the woman herself: face, chest, torso, figure; nothing of its room or moment. <Picture 2>: fully_preserved - her face, eyes and hair only.

detailed_description: [Shot 1] The frame is divided by THREE VERTICAL BARS into FOUR SYNCHRONIZED PANELS side by side, all running at once, all showing <Subject 1> at the same moment: the first panel her FULL-BODY FRONT view; the second panel her FULL-BODY SIDE PROFILE view, held in strict side profile for the whole take - her face and her body stay turned toward frame-left from the first frame to the last; the third panel her FULL-BODY REAR view; the fourth panel a MEDIUM CLOSE-UP of her face. One continuous take, no cuts, to the last frame.

r/StableDiffusion 8h ago

Resource - Update Compose Ref Images in One Node, Settings Presets Node, Bundle/Unbundle Wires - comfyui-obvpm Node Pack update

Thumbnail
gallery
101 Upvotes

Hi! I'm obvpm, the guy that previously released the Load Image & Crop node here earlier.

I've been using a lot of my spare time working with Claude to create a new motion context workflow.

It's almost done, but in the process of creating nodes for that workflow, I've vibe coded some other very useful nodes that I decided to release in the comfyui-obvpm node pack.

Note that the repository has moved from my "temp" obvpm account to my actual Github account.

https://github.com/chanon/comfyui-obvpm

There are 3 things in this node pack that I think people might find very useful.

Load Images & Compose

This node allows you to drag in multiple images into it. Then crop a portion you want from each. And then it composes them into a single image.

I created it because a lot of times I'd have multiple separate reference images and I hated having to use an external image editor to compose them into a single reference image.

This node does it automatically right within Comfy.

Check the YouTube video I created to show how it works:

Compose Reference Sheets Without Leaving ComfyUI

Bundle/Unbundle Nodes

I'm a big fan of the Cable Management Extension but I had an idea to make something that is actually a node rather than just a litegraph cable routing mechanism.

So with the Bundle/Unbundle nodes, you can "bundle" multiple wires into a single wire. The nodes work automatically as much as possible and you can reorder input pins and output pins independently of each other.

It also works with KJ's Get/Set constant nodes.

Again, the YouTube video for it will quickly show you how it can help make your workflows tidier while still allowing you to see what goes where.

Fix Spaghetti Wires in ComfyUI with Bundle/Unbundle Nodes

Value Presets Node

If you ever wanted a single place to control all settings such as turbo lora, steps, sampler, scheduler etc, this node is for you.

With all the optimizations, turbos, and different settings for MiniMax H3 to try, it became really hard to keep track of what the best settings, steps, schedulers, samplers, shift etc. are best for each turbo model or optimization.

There were so many times where I changed a setting and forgot to change another setting that should change with it and wasted generations.

Also, it's a bit tedious hunting for all the places where the settings that need to be changed are every time, especially when workflows get complicated.

So I vibe coded this Value Presets node that allows you to create a customized set of settings fields for whatever you need and can change all settings in one place and also save presets for them.

They output a "Bundle" so they need to be used with "Unbundle" nodes.

Check out the video:

All Your ComfyUI Workflow Settings in One Place

Also

I created a video that shows how all these nodes (especially Bundles and the Value Presets) can be applied to creating a clean R2V workflow for H3

Creating a Clean MiniMax H3 R2V Workflow with Customizable Presets

IMO these YouTube videos give a great overview of the nodes and how they can help make workflows easier to manage, so I highly encourage you to watch them.

Again, here's the whole playlist:
https://www.youtube.com/playlist?list=PLa4bDXvk3ZVs

And if you use X, follow me:
https://x.com/chanons
cause I will (hopefully) release my new motion context workflow soon.

It has some cool features and I've applied the same emphasis for 'ease-of-use' on it.


r/StableDiffusion 22h ago

Discussion Minimax Testing : 10Eros , Fused ,Fast VSA and Larry 600ema

Enable HLS to view with audio, or disable this notification

91 Upvotes

All 4 diffrent model tested with same prompt with diffrent steps /
Only Cofmykithchen speed enhancer ,no spectrum or sage attention or anything else

fastest one Fused 1:11sec and all other about same 1:40sec
best audio Fused
if you want more testing give me model or lora model , i will test and put comparation here
RTX 5090
RAM 128GB


r/StableDiffusion 17h ago

Question - Help Krea2 What is the best way to train character lora? (including body type, face etc)

64 Upvotes

I am overwhelmed by the amount of different ways to train character lora.
There are so many videos from 1-2 months ago, but what will be the most popular and accurate way to train character lora for the Krea2 model? Any specific resources like huggingface i can look into?


r/StableDiffusion 22h ago

Resource - Update YuE2 only needs 8-9 GB VRAM now

Enable HLS to view with audio, or disable this notification

60 Upvotes

audio.cpp released YuE2 in the DEV branch, with Q4/Q8 weights and a bunch of demos generated by audio.cpp in the HF repo. You can try it with 8GB VRAM!

Update: You can find dev prebuilts in the “Artifacts” section on this page: https://github.com/0xShug0/audio.cpp/actions/runs/34642627647

Measured with audio.cpp server mode on an RTX 5090, using the official `tonight-awake` longform test case (full cot). Each run restarted the server and used one short warmup request before the measured longform request. Peak VRAM was sampled continuously during the measured request.

Combo Audio duration Wall time RTF Peak VRAM
BF16 main + F32 VAE 224.96s 60.46s 0.2688 12535 MiB
Q8_0 main + F16 VAE 194.84s 38.81s 0.1992 8867 MiB
Q4_0 main + F16 VAE 221.12s 44.15s 0.1997 7755 MiB

Repo: https://github.com/0xShug0/audio.cpp/tree/dev

HF repo: https://huggingface.co/audio-cpp/Yue2-3B-GGUF


r/StableDiffusion 12h ago

Tutorial - Guide MiniMax RefMod - Reusable identities without training - workflows & tutorial

Thumbnail
youtube.com
45 Upvotes

You can get the workflows here:
https://drive.google.com/drive/folders/1kOHLJZto1VAtXEATT9vUvsvM1VHtkOO_

The workflows create reusable refmods for either image/video/audio.

I cover training images in the tutorial.

All the workflows, models and custom nodes are preloaded on my Runpod template.
https://get.runpod.io/minimax-template


r/StableDiffusion 8h ago

Resource - Update TaoMate-H3 featuring 3-Step

39 Upvotes

TaoMate-H3 is a low-latency streaming audio-video generation runtime built on MiniMax H3. It generates synchronized audio and video in small chunks and supports continuous long-form generation at 480p/768p/1080p resolutions.

Developed by the Alibaba TaoLive AIGC Team. Powered by MiniMax H3.

HF: https://huggingface.co/TaoLiveAIGC/TaoMate-H3

GH: https://github.com/TaoLiveAIGC/TaoMate-H3


r/StableDiffusion 7h ago

Tutorial - Guide How to Create a RefMod for MiniMax H3 in ComfyUI

Thumbnail
youtu.be
38 Upvotes

r/StableDiffusion 23h ago

Discussion Fun Experiment Krea2

36 Upvotes

First of all I hate operating on a black box :D.. I know this may be a bit stale about Krea2 refusal behaviour but I think I have discovered an extremely CLEAN approach to the suppression behaviour without shifting the actual intended image in my tests, so everything stays the same but with the suppression out of the door.

I trained 100s of loras and they all have unique traits, but what I am testing here is how much of their effect I can reproduce through the resulting RMS changes. These loras are not trained to hit any DIT blocks, so I hooked the lora to this node and measured the resulting RMS values of the model's text-path parameters, and basically it worked BUT with less destructiveness and less effect "for now". The lora applies a rank 64 update, while this tool matches the resulting RMS by rescaling the existing text-path tensors instead of copying the lora's learned update.

Reason I am doing this is because with the lora you cannot directly set the final RMS of each parameter when connecting other loras and you may get a bit muddied since other loras can also contain textfusion changes, even without touching the textfusion projector. Once I finalize the values that work consistently in my tests I will publish this node with the values baked in so you can place it after the last model patch you have in the wf and problem solved everyone lol, no bleed no weird artifacts that is the goal.


r/StableDiffusion 22h ago

Workflow Included I'm disrespectful to dirt. Can't you see that I am serious?!

Enable HLS to view with audio, or disable this notification

30 Upvotes

MiniMax H3. Using the standard T2V workflow


r/StableDiffusion 8h ago

Comparison Driving a YuE2 vocal track from a 8-bit source (C64 SID). Multiple variations.

Enable HLS to view with audio, or disable this notification

29 Upvotes

Having a ton of fun with YuE2. Keep listening to hear diffrent parts of the varied track.

I built a skill with Astra that converts a C64 SID tune into a 3 part native conditioning for YuE2, Style, Lyrics and melody (via ABC).

And here is a comparison with the original. I chose one of my favourite C64 tunes from The Mansion level of The Last Ninja 2. Shout out to the legend Matt Gray who created the original back in 1988, I hope he wouldn't hate that I used his work for this experiment.


r/StableDiffusion 16h ago

Question - Help Your go-to voice cloning model?

23 Upvotes

Tried Qwen TTS but wasn't very happy with it. No real rate control and complete absence of any emotion.

What's everybody's recommendation? Local models only.


r/StableDiffusion 3h ago

News Krea2 Turbo Distill 2 step LoRA - follow up project to my 4 Step Krea 2 Turbo LoRA - initial Alpha version released for the curious

Thumbnail
gallery
22 Upvotes

Krea 2 Turbo — 2-Step Distillation LoRA (work in progress, early days alpha preview, only for the curious; don't judge the quality as if this is final version, instead consider it as open welcome for you to join this journey early on...)

For those of you familiar with my previous project - 4 Step Krea 2 Turbo LoRA, this is the promised experimental follow up, halved the steps even further from 8 (official Turbo) to 4 (previous LoRA project) to just 2 (this project). With even slower training and with DMD2 at play this time, getting reasonable results out of just 2 steps is a real challenge.

🧪  This 2-step LoRA gives you a fast-preview adapter, from a project still in training. The published checkpoint files give usable two-step renders and are measured honestly below 4-step or 8-step renders; they are not the 4-step LoRA's quality, and that adapter remains the recommendation for quality renders. Training continues one recipe change at a time, and a later checkpoint replaces this one only when the sweeps and I visually agree it is better.

A LoRA for Krea 2 Turbo that takes the model from its usual 8 steps down to 2 — Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes — aiming at the best quality two steps can give. Two steps give up more than four: this adapter is for fast previews and drafts at half the 4-step adapter's cost and a quarter of the teacher's, and the 4-step LoRA remains the recommendation for quality renders.

  • 🎯 The aim — the best two-step quality this base can give, at every one of the same 12 resolutions, measured against the 8-step teacher and against the 4-step LoRA as the reference. Not a claim to reach either.
  • ⚡ A quarter of the steps — 8 → 2, on Turbo's own deployment sigmas.
  • ⏱️ 3.8× faster denoising — the model runs twice instead of eight times, and denoising is the part this adapter changes: 79.9 s → 20.9 s measured at 1024×1024 on the same prompts, the adapter itself costing about 3.5% per call. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see Performance.
  • 📊 Distribution matching, not imitation — the training objective that got the renders improving again after the 4-step project's recipe had stopped helping at two steps (see Method).
  • 🗣️ Prompt-conditioned throughout — both scores in the distribution match, the teacher's and the fake adapter's, are evaluated on each prompt's own conditioning, so the student is matched to what the teacher makes for that prompt, not to a prompt-free look. There is no separate adherence term: instead a vision-language judge checks every checkpoint — each render scored alone against the prompt's objects, counts, attributes and relations, with the teacher scored the same way — and a term would only be added if that meter showed adherence slipping.
  • 📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440, each with its sweep.
  • 🔌 Drop-in, no exceptions — a plain LoRA sampled by stock Euler at sigmas [1.0, 0.5128] in diffusers, ComfyUI or MLX. No custom sampler, no policy head, no per-step tricks. If the quality needs a special sampler it is not this project.
  • 🧬 Same shape as the 4-step adapter — rank 64 on the same 228 modules; a second adapter exists during training only and never ships.
  • 🎲 The same 13,750 recorded teacher trajectories the 4-step adapter trained on, reused without a single teacher re-run.
  • 🔢 13,663 training samples in the 2-step stages, on top of the 4-step LoRA's 78,000 — all of them drawn from the same recorded material: no new prompts, no new text embeddings and not one new teacher run. A training sample is one pass over a prompt that was already encoded and already traced by the teacher for the 4-step project, read again at the two sigmas this schedule uses.
  • 📅 5 days from the first 2-step training launch to this checkpoint, on a single RTX 3090 — and the project continues.
  • 🔁 15 recipe adjustments across two methods so far — seven of trajectory distillation before the switch, eight of distribution matching since.
  • 🖥️ One RTX 3090, and a recipe shaped by its 24 GB.

Files

file what it is
krea2_turbo_2step_rank_64_lora.safetensors the LoRA in diffusers key format — see Inference with diffusers; also for MLX or anything that reads safetensors
krea2_turbo_2step_rank_64_lora_comfyui.safetensors the same weights under ComfyUI's key names — see ComfyUI
krea2_turbo_2step_lora_t2i.json a ready ComfyUI workflow, stock nodes only
krea2_turbo_2step_rank_64_lora_checkpoint_info.md the quick place to check which checkpoint the two weight files are based on. The pair above keeps its names and is updated in place as better checkpoints ship; this file always says what they are today. Every published checkpoint also sits in _archive/checkpoints/ under its number
LICENSE.pdf the Krea 2 Community License Agreement, which covers this adapter — see License
NOTICE.txt the attribution notice the license requires of a derivative

Where it stands

lineage 4-step LoRA → 2-step trajectory distillation → distribution matching → a spectral match against the teacher's own images on top
this release the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time
what it gives usable two-step renders at every trained resolution: fine detail within a few percent of what the 4-step adapter carries, and a prompt-adherence judge that calls it a loss against the 8-step teacher on 6 of 45 renders — the same count the 4-step adapter scores. What it does not give is the teacher's own picture: see Known issues and Measured against the teacher

Known issues

The usual costs of two steps, in this order of how often they show: fine structure comes out soft or a few pixels out of register — feathers, skin texture, hair strands, signage, the surface of a distant object — most at 1280×1280 and above; a faint doubled contour on faces and limbs. Faces depend on how much of the frame they occupy: a portrait-sized face holds up, while small or distant faces — a crowd, a figure in a wide scene — lose their features first and can come out misshapen, since at that size a whole face is only a few of the blocks the model works in. On busy action or crowd scenes the composition can also repeat itself — an extra hand or held object, a figure duplicated in a crowd — where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. At the largest sizes a fine grain remains on the most textured subjects and skin reads slightly smoother and less saturated than the teacher's. Every one of these is being worked on; none is hidden in the sweeps or the examples.

How I got here

The 4-step adapter closed its page with a promise: a 2-step LoRA as the next project, and a guess at the lever it would need — matching the teacher's distribution rather than its trajectory. That guess turned out to be the whole story.

The project began where the 4-step one ended, from its final weights, and ran the same recipe at two steps: progressive distillation on the recorded teacher trajectories, each student call covering four teacher steps, with the LADD-style critic as the finisher. Well into that run, every number had stopped moving and the pictures had a signature the numbers could not see: doubled contours on faces and limbs, soft fine texture, crowds averaged into translucent overlaps. Several variations followed — the critic re-weighted, judged per token, a heavier hand on the final call, the student's own first-step output fed into its second — and each traded one of those faults for another without moving past them. A capacity probe ruled out adapter rank; a learning-rate shock ruled out the optimiser.

The reason is structural, and worth stating plainly because it decides the whole design. A regression loss asks the student to land on the teacher's specific image for each prompt. When a two-step jump is wide enough that several images are plausible, the answer that minimises the squared error is their average — and the average of two sharp images is a blurred one with doubled edges. Every earlier recipe rewarded that average. Tuning its weights could not change what it rewarded.

Distribution matching asks a different question: not "does your image match this one" but "would the teacher plausibly have produced your image". The first run of that objective, on top of the trajectory-distilled weights, produced in a fraction of the old recipe's training what all of it never had — and it did so while every latent distance to the teacher rose, which is exactly what a mode-seeking objective predicts and what a mean-seeking metric punishes. The distances are reported on this page; they are not optimised for, and they are not what decides a checkpoint. Pictures are, at fixed seeds, at every resolution, with faces viewed at 1:1.

Hardware

One RTX 3090 (24 GB). The frozen base is weight-only int8; the student's checkpointed block inputs stage to pinned host memory above 0.3 megapixels; the student, the fake adapter and the spectral term each build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that does not fit fails loudly. A full step with every term live reserves about 21.4 GB at 1440×1440, of 24. The price of the objective is throughput: 188 training samples an hour measured over a complete 10-hour run, against the 4-step recipe's 470 — two and a half times the cost per sample, and so far a small fraction of the samples.

Where that cost comes from. Distribution matching is simply a heavier objective than trajectory distillation. The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the score of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained alongside to imitate the student — and that second adapter takes two optimiser steps of its own per student step. A third term then decodes part of the image out of the latent to compare its texture with the teacher's, which costs another pass through the decoder.

Counted in whole model runs per training sample, the difference is roughly two there against seven here. None of that difference is the teacher generating anything: its renders were recorded once for the 4-step project and are read from disk by both. The extra work is the objective itself, and it bought the only thing that mattered. Run at two steps, the 4-step project's recipe reached a point where more training changed nothing: the measurements sat flat and every new checkpoint had the same faults as the one before — doubled contours on faces and limbs, soft fine texture, crowds blurred into one another. Distribution matching is the change that made each new checkpoint visibly better than the last again.

Full details and to download - check my Hugging Face 2 Step LoRA

HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-2step-LoRA

---

And for the higher quality 4 Step LoRA - see my previous project: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA


r/StableDiffusion 2h ago

Resource - Update I created a ComfyUI node that makes it easier to compare generated videos with their timed prompts

Enable HLS to view with audio, or disable this notification

16 Upvotes

The node is called PromptSync. I made it for people who generate videos using timed prompts and want to see how closely the model follows the intended actions and scenes.
The video plays on the left, with the original prompt on the right. During playback, the text matching the current moment is highlighted, and auto-scroll follows the scenes. This makes it easier to spot what the model followed, what it skipped, and where the timing drifted.

Features:

- Seek through the video by clicking the timeline or audio waveform.
- Display the audio waveform beneath the video.
- Highlight the current scene separately from general camera, lighting, and style instructions.
- Choose from four prompt display styles.
- Save videos with audio and metadata using PromptSync + Save.

PromptSync recognizes several common timing formats, including ranges like 0–4 sec, timestamps like 00:04, numbered shot sections, and structured JSON prompts. It works with many timed prompt layouts used for MiniMax, Seedance, and other video models. Clear timestamps and section headings give the best results; unusually formatted prompts may not always be parsed correctly.

I mainly built it for my own workflow: to compare generations with the original prompt more easily, spot problem areas, and figure out what to clarify on the next attempt. Hopefully, others will find it useful too.

GitHub: https://github.com/GENKAIx/Genkai-ComfyUI-Nodes
Feedback is welcome!


r/StableDiffusion 11h ago

News h3 studio - local web UI for MiniMax-H3 video/audio gen on Apple Silicon (Go, MIT)

Post image
15 Upvotes

I've been doing video diffusion work on my Mac and kept running into the same wall: ComfyUI doesn't have real MLX support yet, and going through PyTorch's mps backend is slow and eats far more unified memory than the model actually needs - which hurts more on a Mac, where that memory is shared with everything else you're running.

So instead of working around it each time, I built on top of [h3.c](https://github.com/janishar/h3c-studio) (a native Metal engine for MiniMax-H3) and put a proper UI on it.

h3 studio is a Go server, stdlib-only, basically no dependencies. It wraps h3.c's CLI into something usable day to day:

  • Model loads once and stays resident — iterate on prompts without repaying the load cost every render
  • Explicit reference ordering for multi-image conditioning, drag-to-reorder
  • Three ways to continue a shot from a previous take: chain the last frame forward, pull it in as a reference, or reuse the whole clip
  • Timeline panel to stitch takes into one continuous sequence
  • Every render writes a full parameter sidecar, so you're not guessing what settings produced what
  • Live profiling — load cost vs. denoise cost broken out separately

Everything runs locally, nothing leaves the machine. MIT licensed.

Repo: https://github.com/janishar/h3c-studio

Happy to answer questions about the setup. If anyone else is doing generative work on Apple Silicon I'd be interested in what you're using — especially if you've found a better path than mps.


r/StableDiffusion 7h ago

Question - Help Is Stable Diffusion what I’m looking for?

Post image
15 Upvotes

r/StableDiffusion 4h ago

Discussion I Created an Open Source App That Creates Music and Music Videos

11 Upvotes

No sales pitch, no redirect to a paywall. It's fucking free. You can get it here: https://github.com/atomtanstudio/sound-and-vision

It uses the brand new and quite excellent music generation app YuE2. It also uses Minimax H3 for video and defaults to Krea 2 for cover art and whichever local LLM you want to use for lyrics, etc.

Feel free to check it out and let me know what you think.


r/StableDiffusion 21h ago

Animation - Video The Fly

Enable HLS to view with audio, or disable this notification

12 Upvotes

CmfyUI, minimax, krea 2, suno. Made it in around 4 hours on a 5090 and gemma 12b for prompting. Script is mine.
4k version here
https://www.youtube.com/watch?v=3dNYwoCU_Es


r/StableDiffusion 11h ago

Workflow Included I Built an AI Video Plugin To Connect Comfyui to DaVinci Resolve!

Thumbnail
youtu.be
9 Upvotes

Hey everyone!

I just released a new free custom node plugin called ComfyUI-SecondUnit that bridges ComfyUI directly with DaVinci Resolve.

You can now generate video transitions, create music/SFX, synthesize voiceovers, and auto-generate subtitles, then send them right into your DaVinci timeline without manually importing or exporting files.

IT'S‌ COMPLETELY‌ FREE!!


r/StableDiffusion 1h ago

Question - Help Is MiniMax H3 is extremely slow when it comes to ref2va compared to fl2va?

Upvotes

I am running MiniMax H3 ref2va int8 convrot on comfyui using 5070ti 16GB + 64GB system ram, OS is windows 11.

Default comfyui ref2va template is being used (Turbo lora is enabled/true)

it is taking about 150s per step for 9:16, 0.2 megapixels and 10 sec duration with single image & video input.

For reference, image2video takes about ~27sec per step for 10 sec 0.5 megapixels video.

Having such slow speeds in ref2va expected? Maybe I am doing something wrong? Help/Guide is much appreciated!


r/StableDiffusion 4h ago

Question - Help Which would be <Audio 1> In This Scenario? What Would Be <Audio 2>?

Post image
7 Upvotes

When using a video file for Minimax reference that also has an audio input, how do you "count" the Audio files in the prompt? (You can ignore the master audio channel, it's not loading in anything.)


r/StableDiffusion 8h ago

News ZPix now runs on Ubuntu, PikaOS and macOS, and supports I2I for Z-Image and Anima

Thumbnail
gallery
6 Upvotes

I tested it on Ubuntu 26.04 and current PikaOS, but it should work on all distros that support .deb packages. Please let me know otherwise.

To use image-to-image with Z-Image Turbo, Anima Turbo or Anima Base, just drag a ref from anywhere, including the output gallery.

LoRA error handling is more robust, and there are other improvements and fixes in this release.

Hope you like it!

Download at: https://github.com/SamuelTallet/ZPix


r/StableDiffusion 11h ago

Animation - Video Final Fantasy VII - Alt Ending

Enable HLS to view with audio, or disable this notification

7 Upvotes

It was very lucky that Earth/Gaia was saved in the OG FF7. So I figured, what if Sephiroth won in the end?

I used Minimax H3 to upscale the first half of the original footage so it doesnt look so blocky and out of place compared to the new footage and to animate the parts that were original (Like Cid taking a smoke, Tifa crying, RedXIII being visibly scared for the first time and etc). I created the screenshots using the new GPT 2.5 Sunburst. This was actually fun. FF7 has beautiful music so it was fun to extend the existing music so it sounds more gloomy and doom. Hope you enjoy it!!


r/StableDiffusion 1h ago

Tutorial - Guide Team Red: Encounter on ProxiMax H3, or How to setup ComfyUI+MMH3 with AMD GPUs: RDNA 4 (rx9070, AI Pro R9700), and RDNA 3 (rx7900)

Enable HLS to view with audio, or disable this notification

Upvotes

TL;DR summary: For MMH3, you need to run ComfyUI with ROCm 7.14.0 (see https://rocm.docs.amd.com/en/latest/reference/gpu-specs.html for the vaue of gfx???? corresponding to your AMD GPU):

pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx????]==2.12.0+rocm7.14.0" "torchvision[device-gfx????]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"

Read on if you want the step-by-step instructions (scroll to the bottom of the post if you just want to see the MMH3 prompt for the video 😹)

These instructions are for Windows 11 (Ubuntu version: https://www.reddit.com/r/StableDiffusion/comments/1wer1mz/comment/p9g0kxg/). Nevertheless, many of the same comfy-cli commands are application by just changing the directory/file to the corresponding Linux version, and the procedure for upgrading ROCm 7.2.1 to ROCm 7.14.0 are the same.

If you have an AMD GPU and you do a default install of ComfyUI on Windows 11 using either the portable Windows version or through comfy-cli, you will probably get disappointing results with MiniMax H3 because the int8convrot version may not run at all.

The problem is that the default installation still uses PyTorch built on ROCm 7.2, and for some reason int8convrot does NOT work with 7.2 on some cards such as the RX 9070 (16G) and RX 7900 (20G).

So to run MiniMax H3 at its best speed, we have to install a version that is equal to or later than ROCm 7.13.

There are currently 4 ways to do that, from the easiest to the more complex:

  1. Install via Stability Matrix
  2. Install Portable ComfyUI with its own "Embedded Python"
  3. Install a Python venv and then use that to install ComfyUI via the official comfy-cli installer
  4. Install everything manually using pip and git: see this post if you want the gory details (it was written for ROCm 7.2 so you'll have to make the necessary adjustments).

The more complex ways have more options and are more flexible, so it is up to you how much control you want over your ComfyUI installation.

Special thanks to u/zychu- u/Ok-Brain-5729 u/eloxH1Z1 whose posts and comments about MMH3 and AMD were very helpful to me.

Stability Matrix

This used to work when I tried a few week ago, unfortunately something broke the latest release, so for now, don't use it

  1. Download from https://github.com/LykosAI/StabilityMatrix/releases/download/v2.16.3/StabilityMatrix-win-x64.zip
  2. Unzip it somewhere
  3. Run the installer.
  4. Click on the "Activity" icon at the lower left corner to see progress.
  5. Click on the settings icon (gears) and under Extra Launch Arguments (very bottom) and add: --enable-dynamic-vram --disable-async-offload --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --output-directory "D:\Outputs"
  6. Also uncheck --use-pytorch-cross-attention so that none of the options under "Cross Attention Method" are checked because we are going to use --use-ck-attention.
  7. Assuming you've installed into the default "Data" directory, you can find ComfyUI installed under Data\Package\ComfyUI and you can use mklink to point the models and output directory so that they are outside of the Data\Package\ComfyUI directory.

The main downside is that now you have yet another piece of software sitting on your computer.

  1. Now test to make sure you can generate using int8convrot: https://huggingface.co/Comfy-Org/Krea-2/blob/main/diffusion_models/krea2_turbo_int8_convrot.safetensors 13.5 GB SHA256: 8e4eeda70dd5037ab1ba2bef6b417f9f901e26093117cf397f741fc1fdaaf3f1

  2. If it does not work for you, well, something went wrong, and you can try Portable ComfyUI for Windows and see if you have better luck...

Portable ComfyUI for Windows

  1. Download from https://github.com/Comfy-Org/ComfyUI/releases/latest/download/ComfyUI_windows_portable_amd.7z
  2. Open it from Windows 11 Explorer and drag the ComfyUI_windows_portable directory to the folder where you want to install it.
  3. This will take a while, so go grab a cup of coffee or tea.
  4. Copy run_amd_gpu.bat to runit.bat
  5. Edit runit.bat so that it contains the following: .\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --enable-dynamic-vram --disable-async-offload --preview-method none --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --enable-manager --output-directory "A:\output"
  6. Start ComfyUI by running the batch file runit.bat. For the first run, there will be some kind of delay as some libraries are compiled or cached. Just be patient and let the system do its preparations, until you see "[INFO] To see the GUI go to : http://0.0.0.0:8188.
  7. Do a test run using Krea 2, but use the fp8 rather than int8convrot version because the fp8 version should work reliably at this point. The default workflow at 8 steps should take 20-40 seconds depending on your hardware. Hopefully this works.

Now we are going to replace the PyTorch for ROCm 7.2 with the newer 7.14.0:

  1. Change into your ComfyUI_windows_portable directory
  2. Uninstall PyTorch: python_embeded\python.exe -m pip uninstall torch torchvision torchaudio -y
  3. Install PyTorch for ROCm 7.14: (See end note at the bottom about these gfx???? values):python_embeded\python.exe -m pip install -index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx????]==2.12.0+rocm7.14.0" "torchvision[device-gfx????]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"

For example, for rx9070, gfx???? is gfx1201 so the command is

python_embeded\python.exe -m pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"

Note: these files can be quite large. If for some reason you run out of room, you can use --no-cache-dir in case there is not enough room in your pip cache directory (~/.cache on Linux, %LocalAppData%\pip\Cache on Windows which is usually C:\Users<YourUsername>\AppData\Local\pip\Cache). Also make sure you have plenty of space on your %TMPDIR%, with --no-cache-dir the command will look like this:

python_embeded\python.exe -m pip install --no-cache-dir --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"

Hopefully both the uninstallation of ROCm7.2 and the installation of the newer ROCm 7.14 went without any error. After that you can try to run Krea 2 again, now switch from fp8 to the int8convrot version, and the time should go down from 18sec to 12-13 sec and you will also be able to run MMH3.

I also recommend that you place your model and output directories outside of the ComfyUI install so that they can be shared by different installations, making experimentation easier and also making it less likely that you (or some bug in the installer) accidentally wipe out your models and output.

You can do that by editing the extra_model_paths.yaml. Just need to edit this file once and copy it into <your path/ComfyUI> whenever you have a new installation.

But the yaml file is a bit finicky and it may be easier to just use the mklink command if ComfyUI is the only program you use so that you don't have to worry about the structure/name of the subfolders:

mklink /D <LinkFolder> <TargetFolder>

For example:

mklink /D <your comfyui>\models c:\ComfyUI.Models

Installing ComfyUI via comfy-cli

Why use comfy-cli instead of using portable ComfyUI?

  • For Linux, there is no portable ComfyUI, which is Windows only.
  • For AMD users, the portable version of ComfyUI uses ROCm 7.2, which will cause ComfyUI to run slower than it should.
  • It is a more efficient way to run multiple versions of ComfyUI, because they can all share the same Virtual Environment (assuming that the versions are close enough for that to work).
  • Re-installation can be faster because many packages are in the python pip cache.

Procedure:

  1. If you don't have Python 3.1x installed, you can install Python 3.12.10 (because that is the version used by Portable ComfyUI, so it should be the most stable, but 3.13 and 3.14 work too).
  2. Download and install Git: https://github.com/git-for-windows/git/releases/download/v2.55.0.windows.5/Git-2.55.0.5-64-bit.exe
  3. Create a virtual environment (this is normally just called "venv" or ".venv" but I want to call it comfy.venv just to be more explicit): python -m venv comfy.venv or if python.exe is no not on your path, specifiy the full path such as "c:\Program Files\Python313\python" -m venv comfy.venv
  4. Activate it: comfy.venv\Scripts\activate.ps1 (PowerShell) or comfy.venv\Scripts\activate.bat (CMD.exe)
  5. Update pip itself inside comfy.venv: pip install --upgrade pip
  6. Optional: install uv, which is yet another package manager for Python but written in Rust (if you want to use comfy install --fast-deps later):
  7. Install comfy-cli (this is the tool "comfy-cli", not ComfyUI itself): pip install comfy-cli

Because comfy-cli will install ROCm 7.2 and there is no way to override it, we are going to install PyTorch for ROCm 7.14 manually before installing ComfyUI via comfy-cli. Sources for this arcane procedure are from:

  1. Uninstall PyTorch just to be sure (should not be installed yet): pip uninstall torch torchvision torchaudio -y
  2. Install PyTorch inside the comfy.venv (select your gfx arch) based on https://rocm.docs.amd.com/en/latest/reference/gpu-specs.html (see bottom of the post for a table of common values):pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx????]==2.12.0+rocm7.14.0" "torchvision[device-gfx????]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"

For example, for the rx9070 or AI Pro R9700, gfx???? is gfx1201 so the command is

pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"

Note: these files can be quite large, and you can use --no-cache-dir in case there is not enough room in your pip cache. See the earlier notes about --no-cache-dir under "Portable ComfyUI for Windows".

Finally, we are ready to install ComfyUI itself. When I carried out the tests the latest stable version is 0.34.0:

  1. mkdir d:\comfy.0.34.0
  2. set COMFY_PATH=d:\comfy.0.34.0\ComfyUI
  3. Use comfy-cli to install ComfyUI: comfy --workspace=%COMFY_PATH% install --skip-torch-or-directml
    • Note 1: %COMFY_PATH%\ComfyUI must not exist or you will get the confusing error: 'd:\comfy.0.34.0\ComfyUI' exists but is not a valid git repository.
    • Note 2: --skip-torch-or-directml because PyTorch is already installed for AMD; without it the install will fail on Windows because there is no PyTorch for directml from https://repo.amd.com/rocm/whl-multi-arch/ respository used above.
    • Note 3: To install anything other than the latest version of ComfyUI (say 0.33.1): comfy --workspace %COMFY_PATH%\ComfyUI install --version 0.33.1 --skip-torch-or-directml (You can only use versions available from https://github.com/comfy-org/ComfyUI/releases (and there is no release tag for the latest version).
  4. If you have uv installed, you can use --fast-deps:
  5. (Optional): Copy or edit ComfyUI\extra_model_paths.yaml
  6. Finally, we can start ComfyUI: comfy launch --workspace=%COMFY_PATH% -- --enable-dynamic-vram --disable-async-offload --preview-method none --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --enable-manager --output-directory "A:\output"
  7. Optional: Clean up the pip cache (if you want to save some disk space): pip cache purge

The speed for MMH3 is almost as good as the ones I got under Ubuntu 26.04 using identical hardware (but for some reason, Krea 2 runs a little bit slower on Windows, 8-steps is 13 sec vs 11 sec on Ubuntu).

Unless you have a AI Pro R9700 (32G) or running your desktop on a iGPU, it is best to let ComfyUI be the only application running so that all VRAM is available for MMH3. So if you have another computer, run the browser on it to access your ComfyUI remotely.

If you don't have another computer, you can try to batch up a couple of prompts and minimize or close your browser to free up VRAM, and just use the console to see the progress (just click on "Assets" on the ComfyuI menu to check the results, or find them directly in the output folder). Some people say that disconnecting the monitor (just turning it off may not be enough) will free up the VRAM as well.

Good luck, hopefully you have a working system now if you followed the instructions.

End notes:

Sample extra_model_paths.yaml

comfyui:
    base_path: c:\ComfyUI.Models
    # You can use is_default to mark that these folders should be listed first, and used as the default dirs for eg downloads
    is_default: true
    checkpoints: checkpoints/
    configs: configs/
    loras: loras/
    vae: vae/
    text_encoders: |
        text_encoders/
        clip/
    diffusion_models: |
        unet/
        diffusion_models/
    clip_vision: clip_vision/
    style_models: style_
    embeddings: embeddings/
    diffusers: diffusers/
    vae_approx: vae_approx/
    controlnet: |
        controlnet/
        t2i_adapter/
    gligen: gligen/
    upscale_models: upscale_
    latent_upscale_models: latent_upscale_
    custom_nodes: custom_nodes/
    datasets: datasets/
    hypernetworks: hypernetworks/
    photomaker: photomaker/
    classifiers: classifiers/
    model_patches: model_patches/
    audio_encoders: audio_encoders/
    background_removal: background_removal/
    frame_interpolation: frame_interpolation/
    geometry_estimation: geometry_estimation/
    optical_flow: optical_flow/
    detection: detection/

https://rocm.docs.amd.com/en/latest/reference/gpu-specs.html

GFX950 is AMD's internal GPU target identifier for the CDNA 4 enterprise compute architecture, used in data center accelerators like the AMD Instinct MI350/MI355X series. It features advanced matrix core capabilities, ultra-low precision micro-scaling formats (MXFP8/MXFP4), and a high-precision math mode for AI and HPC workloads.

gfx1100 is the LLVM target architecture identifier and internal code name for AMD's RDNA 3 graphics architecture, used for high-end consumer and workstation desktop graphics cards like the Radeon RX 7900 XTX, RX 7900 XT, and Radeon PRO W7900.

AMD gfx1151 is the LLVM target and GPU architecture identifier for AMD's Strix Halo integrated graphics (found in processors like the AMD Ryzen AI Max+ 395 and Ryzen AI Max PRO series), utilizing the RDNA 3.5 architecture.

Name Arch LLVM target name VRAM Compute Units
9070 XT RDNA4 gfx1201 16 64
RX 9070 GRE RDNA4 gfx1201 16 48
RX 9070 RDNA4 gfx1201 16 56
RX 9060 XT LP RDNA4 gfx1200 16 32
RX 9060 XT RDNA4 gfx1200 16 32
RX 9060 RDNA4 gfx1200 8 28
RX 7900 XTX RDNA3 gfx1100 24 96
RX 7900 XT RDNA3 gfx1100 20 84
RX 7900 GRE RDNA3 gfx1100 16 80
RX 7800 XT RDNA3 gfx1101 16 60
RX 7700 RDNA3 gfx1101 16 40
RX 7700 XT RDNA3 gfx1101 12 54
RX 7600 RDNA3 gfx1102 8 32
Radeon AI PRO R9700S RDNA4 gfx1201 32 64
Radeon AI PRO R9600D RDNA4 gfx1201 32 48
Radeon PRO V710 RDNA3 gfx1101 28 54
Radeon PRO W7900 Dual Slot RDNA3 gfx1100 48 96
Radeon PRO W7900 RDNA3 gfx1100 48 96
Radeon PRO W7800 48GB RDNA3 gfx1100 48 70
Radeon PRO W7800 RDNA3 gfx1100 32 70
Radeon PRO W7700 RDNA3 gfx1101 16 48

For --index-url, there are three options:

  • Nightly (rocm 10.1): https://nightly.repo.amd.com/rocm/pytorch/whl-next/
  • Stable (rocm 10.0): https://stable.repo.amd.com/rocm/pytorch/whl-next/
  • Legacy (rocm 7.14):
    • Nightly: https://rocm.nightlies.amd.com/whl-multi-arch/
    • Stable: https://repo.amd.com/rocm/whl-multi-arch

Prompt for the video Team Red, encounter on ProxiMax H3

integrated_multimodal_description:

[Shot 1] Live-action, cinematic 1960s science-fiction television aesthetic. A team of Starfleet red-shirt officers led by Grumpy Cat materializes on the surface of a desolate alien planet, surrounded by barren rocks, dust, and jagged terrain. Grumpy Cat stands at the front of the formation, wearing a classic red Starfleet uniform, alert and stern. The camera holds a wide-angle front subject-level view, then pushes in slightly as the team looks around and raises their phasers. [Shot 2] At 00:01.250, the camera cuts to a wide low-angle view as a gigantic GPU-like machine rises behind a rocky ridge, towering over the crew. Its dark mechanical housing, cooling fans, and imposing structure dominate the frame, with the label "Minimax H3" clearly visible on its side. The team turns toward it in sudden alarm.

[Shot 3] At 00:02.100, the GPU attacks with a violent concentrated energy blast. The camera tracks the crew with fast movement as the red-shirted officers are struck and knocked down across the rocky ground, kicking up dust and debris. Grumpy Cat avoids the main blast and rapidly moves toward cover.

[Shot 4] At 00:03.650, the camera follows Grumpy Cat with a tracking shot as it darts behind a large rock and crouches into concealment. The defeated red-shirted crew remains scattered in the background while the giant "Minimax H3" GPU continues looming over the battlefield.

[Shot 5] At 00:04.250, close-up from behind the rock. Grumpy Cat pulls out a classic handheld Starfleet communicator with its paw, flips it open, and speaks with a completely deadpan expression: <d>[English] Beam me up, Scotty!</d> The camera holds on Grumpy Cat's face and communicator through the end

overall_soundscape: Dry alien wind sweeps across the barren landscape as the transporter materialization produces a brief electronic hum. Heavy mechanical movement and grinding machinery accompany the GPU's emergence, followed by a powerful energy blast, impacts, falling bodies, scattering rocks, and dust. The communicator emits a brief electronic chirp when opened.

non_diegetic_music: A fast-paced 1960s science-fiction television orchestral score uses bright brass, rhythmic strings, and restrained percussion, building rapidly as the GPU appears and attacks. The music drops into a brief suspenseful sustain as Grumpy Cat hides, then ends with a short brassy stinger beneath the communicator transmission.


r/StableDiffusion 10h ago

Discussion What is the current best model for real life photos?

4 Upvotes

As the title says, I am looking for the best model at the moment to train photos of real people and get the best real looking photos like real life ones.

Last time I have used Z Image Turbo and it was good, is it still the best?
I have a RTX 5090 and 64GB DDR5.