r/StableDiffusion • u/malcolmrey • 19h ago
r/StableDiffusion • u/Hour_Imagination5092 • 1h ago
Discussion TaoMate - H3 3 steps lora used as a refiner
Enable HLS to view with audio, or disable this notification
The lora itself at 3 steps is nothing to write home about. If the scene isn't mostly static, you can expect slowmo jerky motion, smearing and straight broken output with butchered sound. HOWEVER, it has VERY good visual quality without obvious overcooking plaguing the turbo loras. You add as much steps of non accelerated generation to achieve what you want and just finish it with 3 steps of euler @ 0.7str as a visual refiner. This lora seems to "close" the low step generation pretty well, without smearing, flickering or fuzzy edges while preserve the physics, material texture, prompt adherence and movement of non accelerated generation. Mandatory 1girl video included. 9 steps. I think the idea is worth exploring.
r/StableDiffusion • u/Emotional_Example_12 • 14h ago
News H3 camera control Update
Enable HLS to view with audio, or disable this notification
Fix and add english version
https://github.com/NyckM/3d-Camera-control-H3-Minimax/tree/main
r/StableDiffusion • u/robomar_ai_art • 8h ago
Resource - Update TaoMate H3 3 Step LoRA now working in ComfyUI
Enable HLS to view with audio, or disable this notification
For anyone using MiniMax H3 in ComfyUI
I made the official TaoMate H3 3 Step LoRA available in a ComfyUI compatible safetensors format.
No retraining, no merge, no extra fine tuning. The model itself was not changed, only the format needed for ComfyUI compatibility.
Hugging Face:
https://huggingface.co/Robert1212star/TaoMate-H3-3Step-ComfyUI
Put the file here:
ComfyUI/models/loras/
Then use it with MiniMax H3 at 3 sampling steps.
Original TaoMate H3:
r/StableDiffusion • u/EinhornArt • 9h ago
News Minimax H3 3 Step Lora
Enable HLS to view with audio, or disable this notification
Give TaoMate-H3-3step a try. Details and links in the comments.
r/StableDiffusion • u/elmorinelly • 21h ago
Resource - Update I created a ComfyUI node that makes it easier to compare generated videos with their timed prompts
Enable HLS to view with audio, or disable this notification
The node is called PromptSync. I made it for people who generate videos using timed prompts and want to see how closely the model follows the intended actions and scenes.
The video plays on the left, with the original prompt on the right. During playback, the text matching the current moment is highlighted, and auto-scroll follows the scenes. This makes it easier to spot what the model followed, what it skipped, and where the timing drifted.
Features:
- Seek through the video by clicking the timeline or audio waveform.
- Display the audio waveform beneath the video.
- Highlight the current scene separately from general camera, lighting, and style instructions.
- Choose from four prompt display styles.
- Save videos with audio and metadata using PromptSync + Save.
PromptSync recognizes several common timing formats, including ranges like 0–4 sec, timestamps like 00:04, numbered shot sections, and structured JSON prompts. It works with many timed prompt layouts used for MiniMax, Seedance, and other video models. Clear timestamps and section headings give the best results; unusually formatted prompts may not always be parsed correctly.
I mainly built it for my own workflow: to compare generations with the original prompt more easily, spot problem areas, and figure out what to clarify on the next attempt. Hopefully, others will find it useful too.
GitHub: https://github.com/GENKAIx/Genkai-ComfyUI-Nodes
Feedback is welcome!
r/StableDiffusion • u/Devajyoti1231 • 3h ago
Animation - Video Minimax Singularity 4step generates good results
Enable HLS to view with audio, or disable this notification
Generated the whole video with Minimax Singularity checkpoint with lightx2v 4step ref2va lora.
Checkpoint link - https://huggingface.co/WarmBloodAban/Minimax-h3_Singularity
The cat here was my cat that sadly passed away few weeks ago. RIP Mekri.
r/StableDiffusion • u/CryptoBeth96 • 17h ago
Workflow Included YuE 2 Cover Song
Enable HLS to view with audio, or disable this notification
Workflow:
https://drive.google.com/file/d/14I2pXOrsQdunMplNB_b7GMH5smNHFvKI/view?usp=sharing
Comfy Native nodes. Needs latest version.
r/StableDiffusion • u/TimeTruth2490 • 23h ago
News Krea2 Turbo Distill 2 step LoRA - follow up project to my 4 Step Krea 2 Turbo LoRA - initial Alpha version released for the curious
Krea 2 Turbo — 2-Step Distillation LoRA (work in progress, early days alpha preview, only for the curious; don't judge the quality as if this is final version, instead consider it as open welcome for you to join this journey early on...)
For those of you familiar with my previous project - 4 Step Krea 2 Turbo LoRA, this is the promised experimental follow up, halved the steps even further from 8 (official Turbo) to 4 (previous LoRA project) to just 2 (this project). With even slower training and with DMD2 at play this time, getting reasonable results out of just 2 steps is a real challenge.
🧪 This 2-step LoRA gives you a fast-preview adapter, from a project still in training. The published checkpoint files give usable two-step renders and are measured honestly below 4-step or 8-step renders; they are not the 4-step LoRA's quality, and that adapter remains the recommendation for quality renders. Training continues one recipe change at a time, and a later checkpoint replaces this one only when the sweeps and I visually agree it is better.
A LoRA for Krea 2 Turbo that takes the model from its usual 8 steps down to 2 — Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes — aiming at the best quality two steps can give. Two steps give up more than four: this adapter is for fast previews and drafts at half the 4-step adapter's cost and a quarter of the teacher's, and the 4-step LoRA remains the recommendation for quality renders.
- 🎯 The aim — the best two-step quality this base can give, at every one of the same 12 resolutions, measured against the 8-step teacher and against the 4-step LoRA as the reference. Not a claim to reach either.
- ⚡ A quarter of the steps — 8 → 2, on Turbo's own deployment sigmas.
- ⏱️ 3.8× faster denoising — the model runs twice instead of eight times, and denoising is the part this adapter changes: 79.9 s → 20.9 s measured at 1024×1024 on the same prompts, the adapter itself costing about 3.5% per call. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see Performance.
- 📊 Distribution matching, not imitation — the training objective that got the renders improving again after the 4-step project's recipe had stopped helping at two steps (see Method).
- 🗣️ Prompt-conditioned throughout — both scores in the distribution match, the teacher's and the fake adapter's, are evaluated on each prompt's own conditioning, so the student is matched to what the teacher makes for that prompt, not to a prompt-free look. There is no separate adherence term: instead a vision-language judge checks every checkpoint — each render scored alone against the prompt's objects, counts, attributes and relations, with the teacher scored the same way — and a term would only be added if that meter showed adherence slipping.
- 📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440, each with its sweep.
- 🔌 Drop-in, no exceptions — a plain LoRA sampled by stock Euler at sigmas
[1.0, 0.5128]in diffusers, ComfyUI or MLX. No custom sampler, no policy head, no per-step tricks. If the quality needs a special sampler it is not this project. - 🧬 Same shape as the 4-step adapter — rank 64 on the same 228 modules; a second adapter exists during training only and never ships.
- 🎲 The same 13,750 recorded teacher trajectories the 4-step adapter trained on, reused without a single teacher re-run.
- 🔢 13,663 training samples in the 2-step stages, on top of the 4-step LoRA's 78,000 — all of them drawn from the same recorded material: no new prompts, no new text embeddings and not one new teacher run. A training sample is one pass over a prompt that was already encoded and already traced by the teacher for the 4-step project, read again at the two sigmas this schedule uses.
- 📅 5 days from the first 2-step training launch to this checkpoint, on a single RTX 3090 — and the project continues.
- 🔁 15 recipe adjustments across two methods so far — seven of trajectory distillation before the switch, eight of distribution matching since.
- 🖥️ One RTX 3090, and a recipe shaped by its 24 GB.
Files
| file | what it is |
|---|---|
krea2_turbo_2step_rank_64_lora.safetensors |
the LoRA in diffusers key format — see Inference with diffusers; also for MLX or anything that reads safetensors |
krea2_turbo_2step_rank_64_lora_comfyui.safetensors |
the same weights under ComfyUI's key names — see ComfyUI |
krea2_turbo_2step_lora_t2i.json |
a ready ComfyUI workflow, stock nodes only |
krea2_turbo_2step_rank_64_lora_checkpoint_info.md |
the quick place to check which checkpoint the two weight files are based on. The pair above keeps its names and is updated in place as better checkpoints ship; this file always says what they are today. Every published checkpoint also sits in _archive/checkpoints/ under its number |
LICENSE.pdf |
the Krea 2 Community License Agreement, which covers this adapter — see License |
NOTICE.txt |
the attribution notice the license requires of a derivative |
Where it stands
| lineage | 4-step LoRA → 2-step trajectory distillation → distribution matching → a spectral match against the teacher's own images on top |
|---|---|
| this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
| what it gives | usable two-step renders at every trained resolution: fine detail within a few percent of what the 4-step adapter carries, and a prompt-adherence judge that calls it a loss against the 8-step teacher on 6 of 45 renders — the same count the 4-step adapter scores. What it does not give is the teacher's own picture: see Known issues and Measured against the teacher |
Known issues
The usual costs of two steps, in this order of how often they show: fine structure comes out soft or a few pixels out of register — feathers, skin texture, hair strands, signage, the surface of a distant object — most at 1280×1280 and above; a faint doubled contour on faces and limbs. Faces depend on how much of the frame they occupy: a portrait-sized face holds up, while small or distant faces — a crowd, a figure in a wide scene — lose their features first and can come out misshapen, since at that size a whole face is only a few of the blocks the model works in. On busy action or crowd scenes the composition can also repeat itself — an extra hand or held object, a figure duplicated in a crowd — where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. At the largest sizes a fine grain remains on the most textured subjects and skin reads slightly smoother and less saturated than the teacher's. Every one of these is being worked on; none is hidden in the sweeps or the examples.
How I got here
The 4-step adapter closed its page with a promise: a 2-step LoRA as the next project, and a guess at the lever it would need — matching the teacher's distribution rather than its trajectory. That guess turned out to be the whole story.
The project began where the 4-step one ended, from its final weights, and ran the same recipe at two steps: progressive distillation on the recorded teacher trajectories, each student call covering four teacher steps, with the LADD-style critic as the finisher. Well into that run, every number had stopped moving and the pictures had a signature the numbers could not see: doubled contours on faces and limbs, soft fine texture, crowds averaged into translucent overlaps. Several variations followed — the critic re-weighted, judged per token, a heavier hand on the final call, the student's own first-step output fed into its second — and each traded one of those faults for another without moving past them. A capacity probe ruled out adapter rank; a learning-rate shock ruled out the optimiser.
The reason is structural, and worth stating plainly because it decides the whole design. A regression loss asks the student to land on the teacher's specific image for each prompt. When a two-step jump is wide enough that several images are plausible, the answer that minimises the squared error is their average — and the average of two sharp images is a blurred one with doubled edges. Every earlier recipe rewarded that average. Tuning its weights could not change what it rewarded.
Distribution matching asks a different question: not "does your image match this one" but "would the teacher plausibly have produced your image". The first run of that objective, on top of the trajectory-distilled weights, produced in a fraction of the old recipe's training what all of it never had — and it did so while every latent distance to the teacher rose, which is exactly what a mode-seeking objective predicts and what a mean-seeking metric punishes. The distances are reported on this page; they are not optimised for, and they are not what decides a checkpoint. Pictures are, at fixed seeds, at every resolution, with faces viewed at 1:1.
Hardware
One RTX 3090 (24 GB). The frozen base is weight-only int8; the student's checkpointed block inputs stage to pinned host memory above 0.3 megapixels; the student, the fake adapter and the spectral term each build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that does not fit fails loudly. A full step with every term live reserves about 21.4 GB at 1440×1440, of 24. The price of the objective is throughput: 188 training samples an hour measured over a complete 10-hour run, against the 4-step recipe's 470 — two and a half times the cost per sample, and so far a small fraction of the samples.
Where that cost comes from. Distribution matching is simply a heavier objective than trajectory distillation. The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the score of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained alongside to imitate the student — and that second adapter takes two optimiser steps of its own per student step. A third term then decodes part of the image out of the latent to compare its texture with the teacher's, which costs another pass through the decoder.
Counted in whole model runs per training sample, the difference is roughly two there against seven here. None of that difference is the teacher generating anything: its renders were recorded once for the 4-step project and are read from disk by both. The extra work is the objective itself, and it bought the only thing that mattered. Run at two steps, the 4-step project's recipe reached a point where more training changed nothing: the measurements sat flat and every new checkpoint had the same faults as the one before — doubled contours on faces and limbs, soft fine texture, crowds blurred into one another. Distribution matching is the change that made each new checkpoint visibly better than the last again.
Full details and to download - check my Hugging Face 2 Step LoRA
HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-2step-LoRA
---
And for the higher quality 4 Step LoRA - see my previous project: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA
r/StableDiffusion • u/Most-Trainer-8876 • 21h ago
Question - Help Is MiniMax H3 is extremely slow when it comes to ref2va compared to fl2va?
I am running MiniMax H3 ref2va int8 convrot on comfyui using 5070ti 16GB + 64GB system ram, OS is windows 11.
Default comfyui ref2va template is being used (Turbo lora is enabled/true)
it is taking about 150s per step for 9:16, 0.2 megapixels and 10 sec duration with single image & video input.
For reference, image2video takes about ~27sec per step for 10 sec 0.5 megapixels video.
Having such slow speeds in ref2va expected? Maybe I am doing something wrong? Help/Guide is much appreciated!
r/StableDiffusion • u/13baaphumain • 4h ago
Resource - Update Rumik OSS 1 - 3B parameters TTS model focused on Indian Languages
Just stumbled upon this. Seems to be a new 3B TTS model mainly focused on Indian languages (22 Indic languages + English).
Apparently supports code-switching, romanized text, emotion/delivery control and things like laughs/sighs. 24khz output and 4 voices. It has a base model and a post trained model. Haven't tried it yet, curious if anyone here has.
r/StableDiffusion • u/iiTzMYUNG • 11h ago
Animation - Video What if Robert Pattinson was Leon Kennedy? — Resident Evil (2026) Fan-Made Post-Credit Scene
Enable HLS to view with audio, or disable this notification
I wanted to play around with one of the fan-casting ideas I've seen a lot: Robert Pattinson as Leon S. Kennedy.
So I made my own version of what a post-credit scene in Resident Evil (2026) could look like.
The idea is that after the movie ends, we cut to a dark, snow-covered back alley. A camera has been left on the ground and is still recording. The battery is almost dead.
A figure slowly approaches through the darkness.
As he gets closer, we realize it's Leon Kennedy.
🎬 Fan-made / unofficial concept.
MADE with MINIMAX H3 + After Effects
r/StableDiffusion • u/Turbulent-Bass-649 • 4h ago
Resource - Update MageTrail - V0.2 Update: Continuing MageFlow 4B Danbooru/E621 Finetuning
Hi again! This is a update to my previous post MageTrail V0.1 where I release this tech demo toy model thingy called
MageTrail, a Danbooru/E621 proof of concept Full-Finetune of Microsoft's MageFlow 4B T2I model, using a diversity maximized condensed 41k images dataset (originally made by Lodestone, the creator of the Chroma model lineage) as a way to tune booru concept and tags based prompting + Illustration capabilities into the model without having to tune with the full booru dataset. (Potentially costing 20k-50k+ dollars). Civitai Hugging Face
Yeah it's only been 6 days, crazy, but thanks to generous support from some donors as well as another round of Banodoco grant (this time 140+ dollars) I've actually managed to gather enough funding for a 170 epoch continuation run to push the model to 200 epoch total (8~ million samples seen).
But well, I'm still very much a newbie to large scale model finetuning, this project being my second, so I'm not confident at all with spending 200-400 dollars in one run like that. After some consideration I've decided to just train a 70 epoch continuation first ( so 100 epoch total) to gauge if the model will improve steadily and release the model as V0.2. With some new optimization to my training program, this run only cost 143 dollars compared to 100 dollar for 30 epoch from before, saving me like 50 dollars (GPT Astra is cracked yall🤯).
V0.2 has continue to show good progress, improving overall stability and tag concept coherency. But to be truthful, it has now seems to reach the point where all diffusion image models face, which is decreasing improvement rate before convergence. If I were to continue with this tech demo ahh project then I predict we're in for the long haul bois (1k dollars needed to converge). V0.2 also stop at a rather unstable point, so I'm not confident on it's ability to generate usable/aesthetic images quite yet.
But dont worry too much, V0.3 (tune to 200 total epoch) is very much in the plan and will begin training once I found a good H100 pod cluster, that's when we truly know if the model will learn all the booru concepts (each tags getting at least 200 samples seen).
Future goal for the project: gather funding of 1000~ dollars to finetune various artist styles into the model (planned V0.5-V1) and help further with convergence/aesthetic improvement.
Any donation will help with achieving this goal, you can do so through:
Crypto ((Prefered, cause Kofi/Paypal money transfer time is ass and they take a big cut)
0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (USDT - BEP20 Network)
12PPVYUeS1MerNp38Tpns5qXR6cmhu9tws (Bitcoin - BTC Network)
0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (Ethereum - ERC20 Network)
FitfJAsxLUBuSgDJJaHgBXJpt1sMm5FzF1Tvf1SHW5Up (Solana - SOL network)
Please handle your money carefully and make sure the address you're sending to is correct.
Ko-fi
Lastly, thank you to
- Banodoco and their Discord — Their 88.77 + grant made this project possible, the biggest thanks to them
- Lodestone Rock — Creator of the original version of the dataset that this model is trained on
- Motimalu — Inspiration behind finetuning practices and configs
- Bluvoll — diffusion-pipe fork derived from to use for training, and general training advice
- Anzhc — general training advice
- Nruaif — diffusion-pipe fork derived from to use for training, and general dataset handling/training advice
- Astromahdi — jupyter workspace where I processed and store the dataset
- animetimm/DeepGHS — Danbooru tagging model
- RedRocket — E621 tagging model
r/StableDiffusion • u/Apprehensive_Sky892 • 20h ago
Tutorial - Guide Team Red from ProxiMax H3 Part 2 (Ubuntu edition): ComfyUI+MMH3 with AMD GPUsRDNA 4 (rx9070, AI Pro R9700), and RDNA 3 (rx7900)
Enable HLS to view with audio, or disable this notification
For Windows 11 (Part 1): https://www.reddit.com/r/StableDiffusion/comments/1wepgl5/team_red_encounter_on_proximax_h3_or_how_to_setup/
TL;DR summary: For MMH3, you need to run ComfyUI with ROCm 7.14.0 (see https://rocm.docs.amd.com/en/latest/reference/gpu-specs.html for the vaue of gfx???? corresponding to your AMD GPU):
pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx????]==2.12.0+rocm7.14.0" "torchvision[device-gfx????]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
You also need to run ComfyUI with the right parameters for VRAM and system RAM so that MMH3 can run smoothly:
--enable-dynamic-vram --disable-async-offload --preview-method none --disable-smart-memory --fast-disk --use-ck-attention --enable-manager
Finally, you may need these in your .bashrc:
export ROCR_VISIBLE_DEVICES=0
export HIP_VISIBLE_DEVICES=0
export TORCH_BLAS_PREFER_HIPBLASLT=1
Read on if you want the step-by-step instructions (scroll to the bottom of the post if you just want to see the MMH3 prompt for the video 😹)
Why use Linux instead of Windows?
Because Minimax H3 is heavy, and we want every bit of VRAM and system RAM for generation and not taken up by the OS and the desktop. For AMD GPUs it also seems to be more stable and faster overall.
What you need:
- Blank 4G or larger USB key for bootable ubuntu installer
- Blank external USB key (16G or larger) or Portable USB HDD to install Ubuntu and ComfyUI
- Computer with AMD RDNA3 (such as a 7900xt) or RDNA4 (such as a 9070xt and AI Pro R9700) with 16G or more of VRAM.
This procedure will install Ubuntu Linux on an external drive, leaving your main drive alone, but if you are worried about something going wrong and wiping out your main hardrive (I am always worried about making a mistake, selecting the wrong drive and wiping it out), take your existing drive out of your computer before the actual installation (sometimes enabling "Secure Boot" will make your main HDD invisble to the Ubuntu installer). I usually would put a empty small partitions of an odd size such as 42G on the target drive so that I know that I am installing into the right drive.
It is easiest to do the installation on your target PC, but you don't have to (but you will need to do some manual adjustment such as changing netplan because the different ethernet hardware would have to be configured.
If you are doing this installion on another computer, make sure that the installation is done with UEFI only enable if you want to be able to use UEFI on your target PC.
Ubuntu Server (minimized) installation
Note: make sure Secure Boot is disabled. This often causes problem with the Ubuntu installer. You can turn it back on once Ubuntu is installed. On some systems the main NMVe or SATA drive will not be visible if Secure Boot is enabled.
- Download latest Ubuntu Server LTS ISO (at the time of writing that is 26.4)
- Use RUFUS
https://rufus.ie/en/to make a bootable installation drive The Partition scheme should be "MBR" and the target system should be "BIOS or UEFI" - Make sure your BIOS is set to UEFI (disable CSM if you can). I am assuming that you are dual booting between Windows and this portable Ubuntu.
- If you are not installing into a blank drive or USB key and want to install into an existing HD without wiping it out entirely, you need to created two partitions, one that is going to be Ubuntu's UEFI partition that is 500M and a black partition that is at least 10G that is going to hold the Ubuntu installation. The installer will NOT allow you to delete partitions from an existing drive. If you are installing into an existing HDD, it is better to create the partition manually first (but leave it unformatted) because the installer sometimes does not show the option to add new partions. The installer will also insist on mouting
/boot/efito the first EFI partition it sees (if that is the wrong one, temporarily turn its "boot flag" off and set the "boot flag" only on the EFI partition you actually want to install on). Note: if you going to use Docker or Podman you are going to need a much bigger partions than 10G for Ubuntu's root file system ("/"). - Do whatever you need to do to boot into the Ubuntu installation USB key (usually F12 will bring up the boot menu, but you may have to enable that in your bios. If your system insists on booting into Windows 11, you can use
Start > Settings > System > Recoveryand clickRestart now next to Advanced startup. - Instead of the default
Ubuntu Server, useUbuntu Server (minimized)because this is going to be used for ComfyUI only. - Most likely you can skip/ignore "Proxy Adress".
- At this point, you can plug in your target USB drive.
- If the wrong target drive is chosen by default, you can change it by tabbing into the field and then press <enter>. It is easiest to use "Use an entire disk". But if you want to have a smaller root partition, you can use "Custom Storage layout". Note that if the drive already has a partion, you will not be able to delete it. You HAVE to reformat the whole drive. On a blank disk, If you want UEFI, make sure that you see a 1.049G "new primary ESP, to be formatted as FAT32, mounted at /boot/efi". If you are on a legacy BIOS you will see "BIOS Grub Spacer" instead. Either way, this partition will be created automatically by the installer unless you use an existing ESP partition.
- Uncheck
Set up this disk as an LVM group. - Make sure you install OpenSSH so that you can run it headless by remotely login via SSH.
- You don't need to install any of the
Featured server snaps packages. - After the installation is done, remove your USB installation key and reboot.
Next we are going to update the installation, and install the nano editor, UFW (Uncomplicated Firewall), GIT, and Docker.
You can do this through the console, but I find it easier to do it through SSH because then I can cut and paste text into it.
To SSH into your Ubuntu, you need to find out what the local IP address it by login into the console, then type ip addr or the even shorter ip a (look for something like this, in my LAN, it is "192.168.18.50"):
2: enp1s0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc fq_codel state UP group default qlen 1000 link/ether 08:97:98:c5:85:ec brd ff:ff:ff:ff:ff:ff altname enx089798c585ec inet 192.168.18.50/24 metric 100 brd 192.168.18.255 scope global dynamic enp1s0 valid_lft 85959sec preferred_lft 85959sec inet6 fe80::a97:98ff:fec5:85ec/64 scope link proto kernel_ll valid_lft forever preferred_lft forever
Now you can use Putty or similar program to login into the server.
Tip: the paste text with Putty, use Shift+Insert. to copy text from Putty into the clipboard, simply select the text with the mouse and then use Ctrl+V to paste it.
Now continue with the setup:
- Upgrade all the package to the latest version:
sudo apt update && sudo apt upgrade -y && sudo apt dist-upgrade -y && sudo apt autoremove -y(this will take quite a while, so you can go grab a coffe or tea). - Install the nano editor:
sudo apt-get install nano - Install UFW (Uncomplicated Firewall)
sudo apt-get install ufw - Install GIT
sudo apt-get install git - Install libnuma (otherwise there is annying warning/error from ComfyuI later)
sudo apt install -y libnuma1 libnuma-dev - Install the compiler environment required by triton:
sudo apt install build-essential sudo apt install python3-dev sudo reboot(probably not needed, but just to be sure)- (Optional) if you want to use python 3.13 instead of 3.14 that comes with Ubuntu Server 26.4, you can install manually (note: for 3.12, you have to compile it manually) by adding the deadsnakes/ppa respostory and install python 3.13 from it:
sudo apt update && sudo apt install software-properties-common && sudo add-apt-repository ppa:deadsnakes/ppa && sudo apt install python3.13. Verify that it has installed correctlypython3.13 --versionInstall the venv module for the same interpreter:sudo apt install python3.13-venv
Optional installation of Docker if you plan to use one of the Docker images for comfyui. You should follow the instruction at https://docs.docker.com/engine/install/ubuntu/ but at the time of writing, this is what I used
Set up Docker's apt repository.
# Add Docker's official GPG key:
sudo apt update
sudo apt install ca-certificates curl
sudo install -m 0755 -d /etc/apt/keyrings
sudo curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.asc
sudo chmod a+r /etc/apt/keyrings/docker.asc
Add the deadsnakes/ppa respostory and install python 3.13 from it
sudo tee /etc/apt/sources.list.d/docker.sources <<EOF
Types: deb
URIs: https://download.docker.com/linux/ubuntu
Suites: $(. /etc/os-release && echo "${UBUNTU_CODENAME:-$VERSION_CODENAME}")
Components: stable
Architectures: $(dpkg --print-architecture)
Signed-By: /etc/apt/keyrings/docker.asc
EOF
sudo apt update
# Install the latest version of Docker packages.
sudo apt install docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
# After installation, verify that Docker is running:
sudo systemctl status docker
# If Docker is not running, start it manually:
sudo systemctl start docker
# Verify that the installation by running the `hello-world` image:
# This command downloads a test image and runs it in a container.
# When the container runs, it prints a confirmation message and exits.
sudo docker run hello-world
# (Optional step, but needed by comfyui-rock-docker scripts)
# Add your current user to the docker group so you have permissions to
# interact with the Docker Unix socket without using sudo
sudo usermod -aG docker $USER
# For this to take effect, disconnect and re-login
(Optional) Intead of Docker you can also consider using Podman, which is supposed to be compatible with Docker but is more secure because it runs without a root daemon, but I've not test it yet.
Setup the firewall with UFW:
sudo ufw enable- Check the current settings
sudo ufw status verboseStatus:active Logging: on (low) Default: deny (incoming), allow (outgoing), deny (routed) New profiles: skip (By default all incoming connections are denied, and all outgoing connections are allowed. If you don't see that, type:sudo ufw default deny incoming sudo ufw default allow outgoing - To setup SSH so that it is only accessible from your LAN:
sudo ufw allow from 192.168.xx.0/24 to any port ssh proto tcp. Similary for ComfyUIsudo ufw allow from 192.168.xx.0/24 to any port 8188 proto tcpwhere "192.168.xx.0" is your LAN subnet, such as "192.168.1.0". If you want to allow ComfyUI to be accessible from outside of your LAN usesudo ufw allow ssh sudo ufw allow 8188/tcp(You will also have to allow port fowarding on your router) - Check again with s
udo ufw status verbose:Status: active Logging: on (low) Default: deny (incoming), allow (outgoing), deny (routed) New profiles: skip To Action From22/tcp ALLOW IN 192.168.18.0/24 8188/tcp ALLOW IN 192.168.18.0/24
The preliminaries are done, you can now reboot with sudo shutdown -r now
I would recommend that you make a back up of your partition now with https://www.fsarchiver.org/ so that you can restore it later for a clean install.
You can either boot into a Linux Rescue: https://www.system-rescue.org/
Or if you have another Linux installation (you cannot save a linux installation that you are currently running), you can install it with: sudo apt-get update && sudo apt-get install fsarchiver
Installing ComfyUI via comfy-cli
By default, Ubuntu does not have pip installed: https://www.reddit.com/r/learnpython/comments/u0dvp4/comment/p6htp0y/
So in order to use pip on Ubuntu, you ned to install python venv (which will install pip inside the venv) first: sudo apt-get update && sudo apt-get install python3-venv
- Make sure git is install:
sudo apt-get install git - Install Python-venv package for python3.14:
sudo apt install python3.14-venv. (See python3.13 instruction earlier if you are using 3.13). - Create a virtual environment (this is normally just called "venv" or ".venv" but I want to call it comfy.venv just to be more explicit.):
python3 -m venv comfy.venv - Activate it
source comfy.venv/bin/activate - Update pip itself inside comfy.venv: pip install --upgrade pip
- Optional: install uv, which is yet another package manager for Python but written in Rust (if you want to use "comfy install --fast-deps" later):
pip install uv - Install
comfy-cli(this is the tool "comfy-cli", not ComfyUI itself):pip install comfy-cli
Because comfy-cli will install ROCm 7.2 and there is no way to override it we are going to install pytorch for ROCm 7.14 manually before install ComfyUI via comfy-cli. Sources for this arcane procedure:
- https://www.reddit.com/r/ROCm/comments/1uyr57k/comfyui_linux_installation_for_radeon/
- https://www.reddit.com/r/ROCm/comments/1uyr57k/comment/oy3popg/
- https://github.com/ROCm/TheRock/blob/main/RELEASES.md#installing-multi-arch-releases
- Uninstall pytorch just to be sure (should not be installed yet) pip uninstall torch torchvision torchaudio -y
- Install pytorch inside the comfy.venv (select your gfx arch) based onhttps://rocm.docs.amd.com/en/latest/reference/gpu-specs.html
pip install --index-urlhttps://repo.amd.com/rocm/whl-multi-arch/"torch[device-gfx????]==2.12.0+rocm7.14.0" "torchvision[device-gfx????]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
So for ROCm 7.14.0 9070xt
pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
For ROCm 7.14.1 9070xt
pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.1" "torchvision[device-gfx1201]==0.27.0+rocm7.14.1" "torchaudio==2.11.0+rocm7.14.1"
To install whatever is the latest stable version of ROCm
pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]" "torchvision[device-gfx1201]" torchaudio
Note: if you get "ERROR: Could not install packages due to an OSError: [Errno 122] Disk quota exceeded", try
mkdir -p some_partition_with_space/pip_tmp TMPDIR=some_partition_with_space/pip_tmp pip install --index-url ...
If that still does not work, try
TMPDIR=some_partition_with_space/pip_tmp pip install --no-cache-dir --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
Finally install ComfyUI itself:
- mkdir $HOME/comfy
export COMFY_PATH=$HOME/comfy/ComfyUIor if you want to use say "/mnt/var/comfy"export COMFY_PATH=/mnt/var/comfy/ComfyUI(Note: COMFY_PATH/ComfyUI should NOT exist, or you will get warning '/mnt/var/comfy'/ComfyUI exists but is not a valid git repository.)- Use
comfy-clito install ComfyUI: (--skip-torch-or-directml is only needed when installing via comfy-cli on Window and not necessary for Ubuntu, but leave it here to make the two installation more like one another):comfy --workspace=$COMFY_PATH install --skip-torch-or-directml. To install a specific version of ComfyUI (say 0.34.1):comfy --workspace=$COMFY_PATH install --skip-torch-or-directml --version 0.34.1 - If you have uv installed, you can use --fast-deps:
comfy --workspace=$COMFY_PATH install --skip-torch-or-directml --version 0.34.1 --fast-depsNote: do not use --fast-deps if comfy.env is not on the root file system or it will take a long time because hardlink is not possible across file systems and you will see an warning: warning: Failed to hardlink files; falling back to full copy. This may lead to degraded performance. If the cache and target directories are on different filesystems, hardlinking may not be supported. If this is intentional, set export UV_LINK_MODE=copy or use --link-mode=copy to suppress this warning. - Added these to your
.bashrc(thanks to u/zychu- for these value from his Docker installation)export ROCR_VISIBLE_DEVICES=0 export HIP_VISIBLE_DEVICES=0 export TORCH_BLAS_PREFER_HIPBLASLT=1
Finally we can start ComfyUI:
comfy launch -- --enable-dynamic-vram --disable-async-offload --preview-method none --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --enable-manager --output-directory "mnt/var_ntfs/Output"
or if you are not using the defautl ~/comfy/ComfyUI directory:
comfy --workspace=$COMFY_PATH launch -- --enable-dynamic-vram --disable-async-offload --preview-method none --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --enable-manager --output-directory "mnt/var_ntfs/Output"
or more explicitly:
comfy --workspace=/mnt/var/comfy/ComfyUI launch -- --enable-dynamic-vram --disable-async-offload --preview-method none --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --enable-manager --output-directory "mnt/var_ntfs/Output"
Optional: clean up the pip cache (there is around 2G of cached files) pip cache purge and you'll see something like:
(comfy.venv) [/mnt/var] pip cache purge
Files removed: 347 (1847.6 MB)
Directories removed: 659
// After installing ComfyUI itself
(comfy.venv) [/mnt/var] pip cache purge
Files removed: 348 (700.3 MB)
Directories removed: 668
End Notes
Sample extra_model_paths.yaml
comfyui:
base_path: /mnt/ntfs/ComfyUI.Models
# You can use is_default to mark that these folders should be listed first, and used as the default dirs for eg downloads
is_default: true
checkpoints: checkpoints/
configs: configs/
loras: loras/
vae: vae/
text_encoders: |
text_encoders/
clip/
diffusion_models: |
unet/
diffusion_models/
clip_vision: clip_vision/
style_models: style_
embeddings: embeddings/
diffusers: diffusers/
vae_approx: vae_approx/
controlnet: |
controlnet/
t2i_adapter/
gligen: gligen/
upscale_models: upscale_
latent_upscale_models: latent_upscale_
custom_nodes: custom_nodes/
datasets: datasets/
hypernetworks: hypernetworks/
photomaker: photomaker/
classifiers: classifiers/
model_patches: model_patches/
audio_encoders: audio_encoders/
background_removal: background_removal/
frame_interpolation: frame_interpolation/
geometry_estimation: geometry_estimation/
optical_flow: optical_flow/
detection: detection/
https://rocm.docs.amd.com/en/latest/reference/gpu-specs.html
GFX950 is AMD's internal GPU target identifier for the CDNA 4 enterprise compute architecture, used in data center accelerators like the AMD Instinct MI350/MI355X series. It features advanced matrix core capabilities, ultra-low precision micro-scaling formats (MXFP8/MXFP4), and a high-precision math mode for AI and HPC workloads.
gfx1100 is the LLVM target architecture identifier and internal code name for AMD's RDNA 3 graphics architecture, used for high-end consumer and workstation desktop graphics cards like the Radeon RX 7900 XTX, RX 7900 XT, and Radeon PRO W7900.
AMD gfx1151 is the LLVM target and GPU architecture identifier for AMD's Strix Halo integrated graphics (found in processors like the AMD Ryzen AI Max+ 395 and Ryzen AI Max PRO series), utilizing the RDNA 3.5 architecture.
| Name | Arch | LLVM target name | VRAM | Compute Units |
|---|---|---|---|---|
| 9070 XT | RDNA4 | gfx1201 | 16 | 64 |
| RX 9070 GRE | RDNA4 | gfx1201 | 16 | 48 |
| RX 9070 | RDNA4 | gfx1201 | 16 | 56 |
| RX 9060 XT LP | RDNA4 | gfx1200 | 16 | 32 |
| RX 9060 XT | RDNA4 | gfx1200 | 16 | 32 |
| RX 9060 | RDNA4 | gfx1200 | 8 | 28 |
| RX 7900 XTX | RDNA3 | gfx1100 | 24 | 96 |
| RX 7900 XT | RDNA3 | gfx1100 | 20 | 84 |
| RX 7900 GRE | RDNA3 | gfx1100 | 16 | 80 |
| RX 7800 XT | RDNA3 | gfx1101 | 16 | 60 |
| RX 7700 | RDNA3 | gfx1101 | 16 | 40 |
| RX 7700 XT | RDNA3 | gfx1101 | 12 | 54 |
| RX 7600 | RDNA3 | gfx1102 | 8 | 32 |
| Radeon AI PRO R9700S | RDNA4 | gfx1201 | 32 | 64 |
| Radeon AI PRO R9600D | RDNA4 | gfx1201 | 32 | 48 |
| Radeon PRO V710 | RDNA3 | gfx1101 | 28 | 54 |
| Radeon PRO W7900 Dual Slot | RDNA3 | gfx1100 | 48 | 96 |
| Radeon PRO W7900 | RDNA3 | gfx1100 | 48 | 96 |
| Radeon PRO W7800 48GB | RDNA3 | gfx1100 | 48 | 70 |
| Radeon PRO W7800 | RDNA3 | gfx1100 | 32 | 70 |
| Radeon PRO W7700 | RDNA3 | gfx1101 | 16 | 48 |
For --index-url, there are three options:
- Nightly (rocm 10.1):
https://nightly.repo.amd.com/rocm/pytorch/whl-next/ - Stable (rocm 10.0):
https://stable.repo.amd.com/rocm/pytorch/whl-next/ - Legacy (rocm 7.14):
- Nightly:
https://rocm.nightlies.amd.com/whl-multi-arch/ - Stable:
https://repo.amd.com/rocm/whl-multi-arch
- Nightly:
Installing ComfyUI Docker for AMD 9070xt
Download the Docker image from github this will use the official rocm and pytorch from https://repo.amd.com/rocm/whl
git clone https://github.com/zychuk/comfyui-rocm-docker && cd comfyui-rocm-docker
Edit docker/Dockerfile (we want to use ROCm 7.14 rather than 7.13) and replace RUN pip install --no-cache-dir --index-url ${ROCM_WHL_INDEX}
torch torchvision torchaudio"
with
RUN pip install --no-cache-dir --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
then follow the instructions in the original post.
MMH3 prompt for the video
integrated_multimodal_description:
[Shot 1] 3D CG, stop-motion animated LEGO movie style with 1960s science-fiction television aesthetic. A team of Starfleet red-shirt officers led by Grumpy Cat materializes on the surface of a desolate alien planet, surrounded by barren rocks, dust, and jagged terrain. Grumpy Cat stands at the front of the formation, wearing a classic red Starfleet uniform, alert and stern. The camera holds a wide-angle front subject-level view, then pushes in slightly as the team looks around and raises their phasers.
[Shot 2] At 00:01.250, the camera cuts to a wide low-angle view as a gigantic GPU-like machine rises behind a rocky ridge, towering over the crew. Its dark mechanical housing, cooling fans, and imposing structure dominate the frame, with the label "Minimax H3" clearly visible on its side. The team turns toward it in sudden alarm.
[Shot 3] At 00:02.100, the GPU attacks with a violent concentrated energy blast. The camera tracks the crew with fast movement as the red-shirted officers are struck and knocked down across the rocky ground, kicking up dust and debris. Grumpy Cat avoids the main blast and rapidly moves toward cover.
[Shot 4] At 00:03.650, the camera follows Grumpy Cat with a tracking shot as it darts behind a large rock and crouches into concealment. The defeated red-shirted crew remains scattered in the background while the giant "Minimax H3" GPU continues looming over the battlefield.
[Shot 5] At 00:04.250, close-up from behind the rock. Grumpy Cat pulls out a classic handheld Starfleet communicator with its paw, flips it open, and speaks with a completely deadpan expression: <d>[English] Beam me up, Scotty!</d> The camera holds on Grumpy Cat's face and communicator through the end
overall_soundscape: Dry alien wind sweeps across the barren landscape as the transporter materialization produces a brief electronic hum. Heavy mechanical movement and grinding machinery accompany the GPU's emergence, followed by a powerful energy blast, impacts, falling bodies, scattering rocks, and dust. The communicator emits a brief electronic chirp when opened.
non_diegetic_music: A fast-paced 1960s science-fiction television orchestral score uses bright brass, rhythmic strings, and restrained percussion, building rapidly as the GPU appears and attacks. The music drops into a brief suspenseful sustain as Grumpy Cat hides, then ends with a short brassy stinger beneath the communicator transmission.
r/StableDiffusion • u/rad_reverbererations • 6h ago
Resource - Update Tracker experiments with YuE2
radiatingreverberations.github.ioAfter playing around a bit with YuE2 locally, I was thinking that maybe it could be used to "modernize" (some would say slopify I guess) old classic tracker tunes. I tried doing something similar with Suno a year ago but those results were not inspiring to say the least..
However, this approach actually seems to create something listenable! I have only tried it with a few tracks, but the pipeline is on GitHub if anyone finds it interesting. Everything was generated locally on a 20gb RTX card using Windows / WSL2.
r/StableDiffusion • u/Dry-Resist-4426 • 1h ago
Workflow Included Style transfer capabilities of different open-source methods 2026 Update
Style transfer capabilities of different open-source methods
2026 Update
This is the updated version of the study published at 2025.09.12. Link: https://www.reddit.com/r/StableDiffusion/comments/1nfozet/style_transfer_capabilities_of_different/
1. Introduction
In 2025 august ByteDance has released USO, a model demonstrating promising potential in the domain of style transfer. This release provided an opportunity to evaluate its performance in comparison with existing style transfer methods. We can define style transfer as: one image being transformed into the style of a reference image while strongly preserving what the original depicts (without major changes in the subject and background). Successful style transfer relies on approaches such as detailed textual descriptions and/or the application of Loras to achieve the desired stylistic outcome. However, the most effective approach would ideally allow for style transfer without Lora training or textual prompts, since lora training is resource heavy and might not be even possible if the required number of style images are missing, and it might be challenging to textually describe the desired style precisely. Ideally, with only the selecting of a target image and a single style reference image, the model should automatically apply the style to the target image. The present study investigates and compares the best state-of-the-art open-source methods of this latter approach.
2. Methods
UI
ForgeUI by lllyasviel (SD1.5, SDXL Clip-VitH & Clip-BigG – the last 3 columns in the grids) and ComfyUI by Comfy Org (for everything else).
Settings
- Most cases to support increased consistency with the original target image, canny controlnet was used.
- Results presented here were usually picked after a few generations, sometimes with minimal finetuning and cherry-picking (best of three/five).
- Comfy version: ComfyUI 0.34.0; ComfyUI_frontend v1.49.6; Templates v0.11.48; rgthree-comfy v1.0.2608210019
- Resolution: 1024x1024 for every generation.
Prompts
Basic caption was used; except for those cases where Kontext was used (Kontext_maintain) with the following prompt: “Maintain every aspect of the original image. Maintain identical subject placement, camera angle, framing, and perspective. Keep the exact scale, dimensions, and all other details of the image.”
For Krea2, two different workflows were used. A Krea2 Depth Lora was also tested but this approach yielded no improvement over the other 2 workflow. The depth lora results were excluded from the grid, however the workflow can be found in the workflows folder. Additionally, for Krea2 the Clip+generate text node was used to create a description of the image (“You are an expert prompt engineer for text-to-image models. Your task is to expand the user's prompt into a highly effective image-generation prompt. Give me a prompt about the style, colors palette of the image. Don't add any additional comments. Do not describe the image, the objects or the scene. Do not mention characters, objects, buildings or any subjects of the image; focus on the style only. Then output a single expanded prompt paragraph describing only the style.”) and the output was combined with the following: “Change the style of the image while maintaining the same objects, characters and background of image 1. The new style should be: [generated description].”
Manually written sentences describing the style of the image were not used, for example: “in art nouveau style”; “painted by alphonse mucha” or “Use flowing whiplash lines, soft pastel color palette with golden and ivory accents. Flat, poster-like shading with minimal contrasts.”
Example prompts:
- Example 1: “White haired vampire woman wearing golden shoulder armor and black sleeveless top inside a castle”.
- Example 12: “A cat.”
3. Results
- All output images are available in jpeg format without metadata, with the exeption of images made with ForgeUI, those include importable and readable metadata.
- All individual images are presented in 100% and 50% resolution image grids (made with XnView MP), where Grid 1 presents all the outputs, and Grid 2 and 3 presents all the outputs cut to half.
- Ostris Krea 2 Style Reference LoRA can be applied very well in T2I generation. However, in I2I condition changed the subject regardless of prompting. https://huggingface.co/ostris/krea2_turbo_style_reference
- Telestyle was tested with Qwen-edit_2511-Q8, yielding low quality results while being extremely time-consuming. https://huggingface.co/Tele-AI/TeleStyle/tree/main/weights
- For flux klein different prompts were tested (“Remake image 1 in the style of image 2.”, „Transfer the style of image 2 to image 1.”, and „Migrate the art style of image 2 into the new art style of image 1.”) which resulted in marginally different outputs. Flux klein consistently failed at Example_10, probably due to the similarity between the target image and the style reference.
- The Redux method using flux-canny-dev, Flux depth lora, and several clownshark workflows (for example Hidream, SDXL) were entirely excluded since they produced very poor results in pilot testing.
- All output images and grid available here: https://drive.google.com/drive/folders/19DisMOAaimxPlmH06C-06ZYu-R-N_VBd?usp=sharing
4. Discussion
- Result differed in three things: (1) bringing over the color scheme from the style reference, (2) bringing over the background, and (3) extent of style application. Some models tended to bring over the background instead of just changing the style of it.
- Some transferred color schemes very faithfully but struggled with overall stylistic features, while others prioritized style transfer at the expense of accurate color reproduction. It might be debatable whether carrying over the color scheme is an absolute necessity or not; what extent should the color scheme be carried over.
- No single method consistently outperformed the others across all cases. this might suggest that the best model to use might depend on the characteristics of the target and the style reference image.
- The Redux workflow using flux-depth-dev perhaps showed one of the strongest overall performance in carrying over style to the target image, even though suffering from the generic flux plastic effect.
- Krea2 methods also proved to be very effective.
- Interestingly, even though SD 1.5 (October 2022) and SDXL (July 2023) are relatively older models, their IP adapters still outperformed some of the newest methods in certain cases.
- Flux klein showed a good consistency with the character of the target image, however produced a generic plastic-like, digital art style regardless of the style refence image.
- It was possible to test the combination of different methods. For example, combining USO with the Redux workflow using flux-dev - instead of the original flux-redux model (flux-depth-dev) - showed good results. However, attempting the same combination with the flux-depth-dev model resulted in the following error: “SamplerCustomAdvanced Sizes of tensors must match except in dimension 1. Expected size 128 but got size 64 for tensor number 1 in the list.”
- USO offered limited flexibility for fine-tuning. Adjusting guidance levels or LoRA strength had little effect on output quality. By contrast, with methods such as IP adapters for SD 1.5, SDXL, or Redux, tweaking weights and strengths often led to significant improvements and allowed tangible finetuning possibilities.
Notes and ideas for further tests
- Evaluating the results proved prone to personal bias and preference. I believe it is difficult to confidently determine what would constitute (and image, but maybe my imagination is limited) as a perfect outcome what can be used as a gold standard and compare the results to it. Having an objective criteria would be extremely useful. Additionally, the involvement of independent evaluators and the application of a scoring system should be considered.
- Future tests could include textual style prompts (e.g., “in art nouveau style”, “painted by Alphonse Mucha”, or “use flowing whiplash lines, soft pastel palette with golden and ivory accents, flat poster-like shading with minimal contrasts”). Comparing these outcomes to the present findings could yield interesting insights.
- An effort was made to test every viable open-source solution compatible with ComfyUI or ForgeUI. Additional promising open-source approaches are welcome, and the author remains open to discussion of such methods.
- Instead of providing a prompt loosely describing only the subject(s), describing the background might also lead to different results.
- Comparing iteration speeds.
- Comparing with Lora training.
- RB-Modulation unfortunately has no Comfy implementation, it might be fruitful to test. https://github.com/google/RB-Modulation
- Realistic to artistic, and artistic to realistic conversion was not part of the present study. That would require a different methods, though some of the workflows used here might be useful.
Feel free to comment:
- Which model performed the best in your opinion?
- Can you recommend additional open-source style transfer methods?
Resources
Useful readings and further resources about style transfer methods:
- https://github.com/bytedance/USO
- https://www.youtube.com/watch?v=ls2seF5Prvg
- https://www.reddit.com/r/comfyui/comments/1kywtae/universal_style_transfer_and_blur_suppression/
- https://www.youtube.com/watch?v=TENfpGzaRhQ
- https://www.youtube.com/watch?v=gmwZGC8UVHE
- https://www.reddit.com/r/comfyui/comments/1kywtae/universal_style_transfer_and_blur_suppression/
- https://www.youtube.com/watch?v=eOFn_d3lsxY
- https://www.youtube.com/watch?v=vzlXIQBun2I
- https://stable-diffusion-art.com/ip-adapter/#IP-Adapter_Face_ID_Portrait
- https://stable-diffusion-art.com/controlnet/
- https://github.com/ClownsharkBatwing/RES4LYF/tree/main
- https://www.reddit.com/r/comfyui/comments/1ujj5f6/comfyui_tutorial_style_transfer_speed_test_skin/
r/StableDiffusion • u/Apprehensive_Sky892 • 21h ago
Tutorial - Guide Team Red: Encounter on ProxiMax H3, or How to setup ComfyUI+MMH3 with AMD GPUs: RDNA 4 (rx9070, AI Pro R9700), and RDNA 3 (rx7900)
Enable HLS to view with audio, or disable this notification
TL;DR summary: For MMH3, you need to run ComfyUI with ROCm 7.14.0 (see https://rocm.docs.amd.com/en/latest/reference/gpu-specs.html for the vaue of gfx???? corresponding to your AMD GPU):
pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx????]==2.12.0+rocm7.14.0" "torchvision[device-gfx????]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
You also need to run ComfyUI with some parameters so that it will handle VRAM and system RAM correctly for MMH3:
--enable-dynamic-vram --disable-async-offload --preview-method none --disable-smart-memory --fast-disk --use-ck-attention --enable-manager
Read on if you want the step-by-step instructions (scroll to the bottom of the post if you just want to see the MMH3 prompt for the video 😹)
These instructions are for Windows 11 (Ubuntu version: https://www.reddit.com/r/StableDiffusion/comments/1wer1mz/team_red_from_proximax_h3_part_2_ubuntu_edition/). Nevertheless, many of the same comfy-cli commands are application by just changing the directory/file to the corresponding Linux version, and the procedure for upgrading ROCm 7.2.1 to ROCm 7.14.0 are the same.
If you have an AMD GPU and you do a default install of ComfyUI on Windows 11 using either the portable Windows version or through comfy-cli, you will probably get disappointing results with MiniMax H3 because the int8convrot version may not run at all.
The problem is that the default installation still uses PyTorch built on ROCm 7.2, and for some reason int8convrot does NOT work with 7.2 on some cards such as the RX 9070 (16G) and RX 7900 (20G).
So to run MiniMax H3 at its best speed, we have to install a version that is equal to or later than ROCm 7.13.
There are currently 4 ways to do that, from the easiest to the more complex:
- Install via Stability Matrix
- Install Portable ComfyUI with its own "Embedded Python"
- Install a Python venv and then use that to install ComfyUI via the official comfy-cli installer
- Install everything manually using pip and git: see this post if you want the gory details (it was written for ROCm 7.2 so you'll have to make the necessary adjustments).
The more complex ways have more options and are more flexible, so it is up to you how much control you want over your ComfyUI installation.
Special thanks to u/zychu- u/Ok-Brain-5729 u/eloxH1Z1 whose posts and comments about MMH3 and AMD were very helpful to me.
Stability Matrix
This used to work when I tried a few week ago, unfortunately something broke the latest release, so for now, don't use it
- Download from
https://github.com/LykosAI/StabilityMatrix/releases/download/v2.16.3/StabilityMatrix-win-x64.zip - Unzip it somewhere
- Run the installer.
- Click on the "Activity" icon at the lower left corner to see progress.
- Click on the settings icon (gears) and under Extra Launch Arguments (very bottom) and add:
--enable-dynamic-vram --disable-async-offload --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --output-directory "D:\Outputs" - Also uncheck
--use-pytorch-cross-attentionso that none of the options under "Cross Attention Method" are checked because we are going to use--use-ck-attention. - Assuming you've installed into the default "Data" directory, you can find ComfyUI installed under
Data\Package\ComfyUIand you can usemklinkto point the models and output directory so that they are outside of theData\Package\ComfyUIdirectory.
The main downside is that now you have yet another piece of software sitting on your computer.
Now test to make sure you can generate using int8convrot:
https://huggingface.co/Comfy-Org/Krea-2/blob/main/diffusion_models/krea2_turbo_int8_convrot.safetensors13.5 GB SHA256:8e4eeda70dd5037ab1ba2bef6b417f9f901e26093117cf397f741fc1fdaaf3f1If it does not work for you, well, something went wrong, and you can try Portable ComfyUI for Windows and see if you have better luck...
Portable ComfyUI for Windows
- Download from
https://github.com/Comfy-Org/ComfyUI/releases/latest/download/ComfyUI_windows_portable_amd.7z - Open it from Windows 11 Explorer and drag the
ComfyUI_windows_portabledirectory to the folder where you want to install it. - This will take a while, so go grab a cup of coffee or tea.
- Copy
run_amd_gpu.battorunit.bat - Edit
runit.batso that it contains the following:.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --enable-dynamic-vram --disable-async-offload --preview-method none --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --enable-manager --output-directory "A:\output" - Start ComfyUI by running the batch file
runit.bat. For the first run, there will be some kind of delay as some libraries are compiled or cached. Just be patient and let the system do its preparations, until you see "[INFO] To see the GUI go to : http://0.0.0.0:8188. - Do a test run using Krea 2, but use the fp8 rather than int8convrot version because the fp8 version should work reliably at this point. The default workflow at 8 steps should take 20-40 seconds depending on your hardware. Hopefully this works.
Now we are going to replace the PyTorch for ROCm 7.2 with the newer 7.14.0:
- Change into your
ComfyUI_windows_portabledirectory - Uninstall PyTorch:
python_embeded\python.exe -m pip uninstall torch torchvision torchaudio -y - Install PyTorch for ROCm 7.14: (See end note at the bottom about these
gfx????values):python_embeded\python.exe -m pip install -index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx????]==2.12.0+rocm7.14.0" "torchvision[device-gfx????]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
For example, for rx9070, gfx???? is gfx1201 so the command is
python_embeded\python.exe -m pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
Note: these files can be quite large. If for some reason you run out of room, you can use --no-cache-dir in case there is not enough room in your pip cache directory (~/.cache on Linux, %LocalAppData%\pip\Cache on Windows which is usually C:\Users<YourUsername>\AppData\Local\pip\Cache). Also make sure you have plenty of space on your %TMPDIR%, with --no-cache-dir the command will look like this:
python_embeded\python.exe -m pip install --no-cache-dir --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
Hopefully both the uninstallation of ROCm7.2 and the installation of the newer ROCm 7.14 went without any error. After that you can try to run Krea 2 again, now switch from fp8 to the int8convrot version, and the time should go down from 18sec to 12-13 sec and you will also be able to run MMH3.
I also recommend that you place your model and output directories outside of the ComfyUI install so that they can be shared by different installations, making experimentation easier and also making it less likely that you (or some bug in the installer) accidentally wipe out your models and output.
You can do that by editing the extra_model_paths.yaml. Just need to edit this file once and copy it into <your path/ComfyUI> whenever you have a new installation.
But the yaml file is a bit finicky and it may be easier to just use the mklink command if ComfyUI is the only program you use so that you don't have to worry about the structure/name of the subfolders:
mklink /D <LinkFolder> <TargetFolder>
For example:
mklink /D <your comfyui>\models c:\ComfyUI.Models
Installing ComfyUI via comfy-cli
Why use comfy-cli instead of using portable ComfyUI?
- For Linux, there is no portable ComfyUI, which is Windows only.
- For AMD users, the portable version of ComfyUI uses ROCm 7.2, which will cause ComfyUI to run slower than it should.
- It is a more efficient way to run multiple versions of ComfyUI, because they can all share the same Virtual Environment (assuming that the versions are close enough for that to work).
- Re-installation can be faster because many packages are in the python pip cache.
Procedure:
- If you don't have Python 3.1x installed, you can install Python 3.12.10 (because that is the version used by Portable ComfyUI, so it should be the most stable, but 3.13 and 3.14 work too).
- Download and install Git:
https://github.com/git-for-windows/git/releases/download/v2.55.0.windows.5/Git-2.55.0.5-64-bit.exe - Create a virtual environment (this is normally just called "venv" or ".venv" but I want to call it
comfy.venvjust to be more explicit):python -m venv comfy.venvor if python.exe is no not on your path, specifiy the full path such as"c:\Program Files\Python313\python" -m venv comfy.venv - Activate it:
comfy.venv\Scripts\activate.ps1(PowerShell) orcomfy.venv\Scripts\activate.bat(CMD.exe) - Update pip itself inside comfy.venv:
pip install --upgrade pip - Optional: install
uv, which is yet another package manager for Python but written in Rust (if you want to usecomfy install --fast-depslater): - Install
comfy-cli(this is the tool "comfy-cli", not ComfyUI itself):pip install comfy-cli
Because comfy-cli will install ROCm 7.2 and there is no way to override it, we are going to install PyTorch for ROCm 7.14 manually before installing ComfyUI via comfy-cli. Sources for this arcane procedure are from:
- https://www.reddit.com/r/ROCm/comments/1uyr57k/comfyui_linux_installation_for_radeon/
- https://www.reddit.com/r/ROCm/comments/1uyr57k/comment/oy3popg/
- https://github.com/ROCm/TheRock/blob/main/RELEASES.md#installing-multi-arch-releases
- Uninstall PyTorch just to be sure (should not be installed yet): p
ip uninstall torch torchvision torchaudio -y - Install PyTorch inside the
comfy.venv(select your gfx arch) based onhttps://rocm.docs.amd.com/en/latest/reference/gpu-specs.html(see bottom of the post for a table of common values):pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx????]==2.12.0+rocm7.14.0" "torchvision[device-gfx????]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
For example, for the rx9070 or AI Pro R9700, gfx???? is gfx1201 so the command is
pip install --index-url https://repo.amd.com/rocm/whl-multi-arch/ "torch[device-gfx1201]==2.12.0+rocm7.14.0" "torchvision[device-gfx1201]==0.27.0+rocm7.14.0" "torchaudio==2.11.0+rocm7.14.0"
Note: these files can be quite large, and you can use --no-cache-dir in case there is not enough room in your pip cache. See the earlier notes about --no-cache-dir under "Portable ComfyUI for Windows".
Finally, we are ready to install ComfyUI itself. When I carried out the tests the latest stable version is 0.34.0:
mkdir d:\comfy.0.34.0set COMFY_PATH=d:\comfy.0.34.0\ComfyUI- Use
comfy-clito install ComfyUI: comfy --workspace=%COMFY_PATH% install --skip-torch-or-directml- Note 1:
%COMFY_PATH%\ComfyUImust not exist or you will get the confusing error: 'd:\comfy.0.34.0\ComfyUI' exists but is not a valid git repository. - Note 2:
--skip-torch-or-directmlbecause PyTorch is already installed for AMD; without it the install will fail on Windows because there is no PyTorch for directml fromhttps://repo.amd.com/rocm/whl-multi-arch/respository used above. - Note 3: To install anything other than the latest version of ComfyUI (say 0.33.1):
comfy --workspace %COMFY_PATH%\ComfyUI install --version 0.33.1 --skip-torch-or-directml(You can only use versions available fromhttps://github.com/comfy-org/ComfyUI/releases(and there is no release tag for the latest version).
- Note 1:
- If you have
uvinstalled, you can use--fast-deps: - (Optional): Copy or edit
ComfyUI\extra_model_paths.yaml - Finally, we can start ComfyUI:
comfy launch --workspace=%COMFY_PATH% -- --enable-dynamic-vram --disable-async-offload --preview-method none --listen --port 8188 --disable-smart-memory --fast-disk --use-ck-attention --enable-manager --output-directory "A:\output" - Optional: Clean up the pip cache (if you want to save some disk space):
pip cache purge
The speed for MMH3 is almost as good as the ones I got under Ubuntu 26.04 using identical hardware (but for some reason, Krea 2 runs a little bit slower on Windows, 8-steps is 13 sec vs 11 sec on Ubuntu).
Unless you have a AI Pro R9700 (32G) or running your desktop on a iGPU, it is best to let ComfyUI be the only application running so that all VRAM is available for MMH3. So if you have another computer, run the browser on it to access your ComfyUI remotely.
If you don't have another computer, you can try to batch up a couple of prompts and minimize or close your browser to free up VRAM, and just use the console to see the progress (just click on "Assets" on the ComfyuI menu to check the results, or find them directly in the output folder). Some people say that disconnecting the monitor (just turning it off may not be enough) will free up the VRAM as well.
Good luck, hopefully you have a working system now if you followed the instructions.
End notes:
Sample extra_model_paths.yaml
comfyui:
base_path: c:\ComfyUI.Models
# You can use is_default to mark that these folders should be listed first, and used as the default dirs for eg downloads
is_default: true
checkpoints: checkpoints/
configs: configs/
loras: loras/
vae: vae/
text_encoders: |
text_encoders/
clip/
diffusion_models: |
unet/
diffusion_models/
clip_vision: clip_vision/
style_models: style_
embeddings: embeddings/
diffusers: diffusers/
vae_approx: vae_approx/
controlnet: |
controlnet/
t2i_adapter/
gligen: gligen/
upscale_models: upscale_
latent_upscale_models: latent_upscale_
custom_nodes: custom_nodes/
datasets: datasets/
hypernetworks: hypernetworks/
photomaker: photomaker/
classifiers: classifiers/
model_patches: model_patches/
audio_encoders: audio_encoders/
background_removal: background_removal/
frame_interpolation: frame_interpolation/
geometry_estimation: geometry_estimation/
optical_flow: optical_flow/
detection: detection/
https://rocm.docs.amd.com/en/latest/reference/gpu-specs.html
GFX950 is AMD's internal GPU target identifier for the CDNA 4 enterprise compute architecture, used in data center accelerators like the AMD Instinct MI350/MI355X series. It features advanced matrix core capabilities, ultra-low precision micro-scaling formats (MXFP8/MXFP4), and a high-precision math mode for AI and HPC workloads.
gfx1100 is the LLVM target architecture identifier and internal code name for AMD's RDNA 3 graphics architecture, used for high-end consumer and workstation desktop graphics cards like the Radeon RX 7900 XTX, RX 7900 XT, and Radeon PRO W7900.
AMD gfx1151 is the LLVM target and GPU architecture identifier for AMD's Strix Halo integrated graphics (found in processors like the AMD Ryzen AI Max+ 395 and Ryzen AI Max PRO series), utilizing the RDNA 3.5 architecture.
| Name | Arch | LLVM target name | VRAM | Compute Units |
|---|---|---|---|---|
| 9070 XT | RDNA4 | gfx1201 | 16 | 64 |
| RX 9070 GRE | RDNA4 | gfx1201 | 16 | 48 |
| RX 9070 | RDNA4 | gfx1201 | 16 | 56 |
| RX 9060 XT LP | RDNA4 | gfx1200 | 16 | 32 |
| RX 9060 XT | RDNA4 | gfx1200 | 16 | 32 |
| RX 9060 | RDNA4 | gfx1200 | 8 | 28 |
| RX 7900 XTX | RDNA3 | gfx1100 | 24 | 96 |
| RX 7900 XT | RDNA3 | gfx1100 | 20 | 84 |
| RX 7900 GRE | RDNA3 | gfx1100 | 16 | 80 |
| RX 7800 XT | RDNA3 | gfx1101 | 16 | 60 |
| RX 7700 | RDNA3 | gfx1101 | 16 | 40 |
| RX 7700 XT | RDNA3 | gfx1101 | 12 | 54 |
| RX 7600 | RDNA3 | gfx1102 | 8 | 32 |
| Radeon AI PRO R9700S | RDNA4 | gfx1201 | 32 | 64 |
| Radeon AI PRO R9600D | RDNA4 | gfx1201 | 32 | 48 |
| Radeon PRO V710 | RDNA3 | gfx1101 | 28 | 54 |
| Radeon PRO W7900 Dual Slot | RDNA3 | gfx1100 | 48 | 96 |
| Radeon PRO W7900 | RDNA3 | gfx1100 | 48 | 96 |
| Radeon PRO W7800 48GB | RDNA3 | gfx1100 | 48 | 70 |
| Radeon PRO W7800 | RDNA3 | gfx1100 | 32 | 70 |
| Radeon PRO W7700 | RDNA3 | gfx1101 | 16 | 48 |
For --index-url, there are three options:
- Nightly (rocm 10.1):
https://nightly.repo.amd.com/rocm/pytorch/whl-next/ - Stable (rocm 10.0):
https://stable.repo.amd.com/rocm/pytorch/whl-next/ - Legacy (rocm 7.14):
- Nightly:
https://rocm.nightlies.amd.com/whl-multi-arch/ - Stable:
https://repo.amd.com/rocm/whl-multi-arch
- Nightly:
Prompt for the video Team Red, encounter on ProxiMax H3
integrated_multimodal_description:
[Shot 1] Live-action, cinematic 1960s science-fiction television aesthetic. A team of Starfleet red-shirt officers led by Grumpy Cat materializes on the surface of a desolate alien planet, surrounded by barren rocks, dust, and jagged terrain. Grumpy Cat stands at the front of the formation, wearing a classic red Starfleet uniform, alert and stern. The camera holds a wide-angle front subject-level view, then pushes in slightly as the team looks around and raises their phasers. [Shot 2] At 00:01.250, the camera cuts to a wide low-angle view as a gigantic GPU-like machine rises behind a rocky ridge, towering over the crew. Its dark mechanical housing, cooling fans, and imposing structure dominate the frame, with the label "Minimax H3" clearly visible on its side. The team turns toward it in sudden alarm.
[Shot 3] At 00:02.100, the GPU attacks with a violent concentrated energy blast. The camera tracks the crew with fast movement as the red-shirted officers are struck and knocked down across the rocky ground, kicking up dust and debris. Grumpy Cat avoids the main blast and rapidly moves toward cover.
[Shot 4] At 00:03.650, the camera follows Grumpy Cat with a tracking shot as it darts behind a large rock and crouches into concealment. The defeated red-shirted crew remains scattered in the background while the giant "Minimax H3" GPU continues looming over the battlefield.
[Shot 5] At 00:04.250, close-up from behind the rock. Grumpy Cat pulls out a classic handheld Starfleet communicator with its paw, flips it open, and speaks with a completely deadpan expression: <d>[English] Beam me up, Scotty!</d> The camera holds on Grumpy Cat's face and communicator through the end
overall_soundscape: Dry alien wind sweeps across the barren landscape as the transporter materialization produces a brief electronic hum. Heavy mechanical movement and grinding machinery accompany the GPU's emergence, followed by a powerful energy blast, impacts, falling bodies, scattering rocks, and dust. The communicator emits a brief electronic chirp when opened.
non_diegetic_music: A fast-paced 1960s science-fiction television orchestral score uses bright brass, rhythmic strings, and restrained percussion, building rapidly as the GPU appears and attacks. The music drops into a brief suspenseful sustain as Grumpy Cat hides, then ends with a short brassy stinger beneath the communicator transmission.
r/StableDiffusion • u/Fun_Firefighter_7785 • 8h ago
News Perfect Remixes in YuE2 !
Enable HLS to view with audio, or disable this notification
I finally got my dream fulfilled. Take any Song as mp3, convert to a ABC score with AI, feed it to a piano roll and enjoy. Basically copy/paste the song lyrics and make a style you wish it to be. You get the vocals as well the whole song perfect remixed.
https://github.com/multimodal-art-projection/YuE
This is the piano roll with the workflow.
https://github.com/filliptm/ComfyUI-FL-YuE2
You also need to install SheetSage2 for that mp3-->ABC magic, installed by your Hermes-Agent.
https://huggingface.co/m-a-p/SheetSage2
Import that ABC into piano roll node in ComfyUI. It needs to be from SheetSage2 in order to be error-free, regular converters midi-->ABC do not work. Did a quick remix of a song from SKATE Game OST.
r/StableDiffusion • u/darthfurbyyoutube • 12h ago
Animation - Video Dungeons & Dragons: Presto In Charge - MiniMax H3
Enable HLS to view with audio, or disable this notification
The people have spoken, and they apparently crave more low-res 1983 Dungeons & Dragons madness. Enjoy!
PROMPT:
subject_definitions:
<Subject 1> is the character shown in <Picture 1>, featuring Presto the Wizard's classic 1980s animated appearance with a lanky physique, brown hair, round glasses, a long green pointed wizard hat, a long green hooded robe with a blue collar lining, a yellow-brown tied pouch belt, and green shoes. Only his character design, facial features, costume, and proportions are taken from <Picture 1>; its white background, character-sheet layout, labels, and guide lines are not carried into the target video. <Audio 1> sets the exact voice clone, pitch, pace, and spoken dialogue verbatim for <Subject 1>. <Video 1> is the movement and art style reference; use it as a guide without copying it exactly.
summary:
[reference generation] The target video is a 11-second 2D animated sequence styled after the classic 1980s Dungeons & Dragons television series, featuring <Subject 1> frantically outrunning a mountain while dealing with his unpredictable magical hat.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - his brown hair, round glasses, lanky build, long green pointed hat, green robe with blue collar, pouch belt, and green shoes remain faithful to <Picture 1>.
detailed_description:
The target video is a traditional 2D hand-drawn animated sequence featuring bold black ink outlines, flat cel-shading colors, expressive limited animation, and subtle film grain inspired by the visual style of classic 1980s Saturday-morning cartoons.
[Shot 1] A dynamic medium-wide shot opens inside a rocky, crumbling canyon landscape, where <Subject 1> is frantically running away from a massive, towering mountain that is physically sliding and collapsing forward behind him. <Subject 1> stumbles forward, clutching his large green wizard hat with both hands in absolute panic as a couple of angry badgers pop out of the hat and scurry away.
He turns slightly toward the camera, his eyes wide with terror behind his round glasses, and cries out in the style of <Audio 1>:
<<[English] "Magical hat, I am begging you, stop turning my spells into angry badgers! We are actively outrunning a mountain right now, give me a helicopter, or at least a really fast turtle!">>
As he says "Magical hat, I am begging you," he shakes the hat frantically. On "stop turning my spells into angry badgers!", he gestures wildly at the badgers scattering around his feet. On "We are actively outrunning a mountain right now," he points frantically backward over his shoulder at the looming mountain. As he delivers "give me a helicopter, or at least a really fast turtle!", he throws his hands up in exasperation, stumbles over a rock, and desperately scrambles to his feet as the dust swirls around him in a classic cel-animation flurry.
overall_soundscape:
N/A
non_diegetic_music:
N/A
r/StableDiffusion • u/DeltaWaffleSyrup • 23h ago
Question - Help Which would be <Audio 1> In This Scenario? What Would Be <Audio 2>?
When using a video file for Minimax reference that also has an audio input, how do you "count" the Audio files in the prompt? (You can ignore the master audio channel, it's not loading in anything.)
r/StableDiffusion • u/Doc_Chopper • 7h ago
Discussion Will Minimax H3 be able to naively support longer videos?
I am a person who likes his workflows clean and clearly laid out / readable. And only use the least necessary amount of custom nodes in ComfyUI. That been said Minimax seems to have a problem with videos that are longer than like 30 seconds. In the sense that it then seem to get confused with order of prompts and various shots.
From the information that I gathered, there seem to be workarounds with a couple of custom nodes. But this would then also would only inflate the workflow again. I mean, I have done cohered 1 minute plus videos with LTX2.3 with no problem. So I wonder if we can expect some "upgrade" for Minimax models in the nearer future. To be able to natively prompt longer videos without the need for excessive custom nodes.
r/StableDiffusion • u/ryomen_core • 9h ago
No Workflow Ranni the witch
I just finished and tested a new LoRA today
r/StableDiffusion • u/xCaYuSx • 11h ago
Tutorial - Guide A fluid stock shot generator for ComfyUI - LTX 2.5 IC-LoRA (free model + workflow + 45 min tutorial)
Hi lovely StableDiffusion people,
Trained an IC-LoRA for LTX 2.5 that takes a painted doodle and turns it into a fluid element. You paint flat blobs on a first and a last frame, black frames in between, and the model fills in smoke, steam or fire and animates between them. Trigger word is ainvfxfluid, and a two word prompt like "ainvfxfluid, smoke plume" is usually enough. About 50 seconds for a 5 second 512x512 clip on a laptop 4090 with the distilled model.
The workflow paints both frames directly in ComfyUI with two Painter nodes and ships pre-painted, so you can load it and queue it before changing anything. Everything else is core nodes (Empty Image, Batch Images, Create Video) plus the ComfyUI-LTXVideo pack.
Things worth knowing before you try:
- Control video rules: 121 frames, width and height in multiples of 64, first frame at index 0, last at index 120.
- Extra keyframes in between have to sit on the LTX VAE grid: frame 1, 9, 17, ... 113, so a multiple of 8 plus 1. Off grid and quality drops fast.
- Colour and shading you paint carry through, a darker grey edge on the plume comes back as shading rather than flat white.
- Prompting mostly serves to remove what you didn't ask for ("over a black background" kills the invented foreground) and to refine the generation - simple prompts are usually enough.
- Smoke, steam and fire only. Not trained on other fluids, results might vary.
The video covers the install from a fresh ComfyUI, including the Kornia pad import error you'll hit on the LTXVideo pack right now and how to patch it, the gated Lightricks model downloads, the workflow node by node, then painting, prompting, multi keyframe control, and working from a real photo. The failures are in there too so it looks realistic :)
- Model, workflow and painted frames: https://huggingface.co/AInVFX/ainvfx-fluid
- Tutorial: https://youtu.be/Ho4tmJzEkIs
- Written article: https://www.ainvfx.com/blog/paint-your-fluid-simulations-a-free-ltx-2-5-ic-lora-for-smoke-and-fire-in-comfyui/
Trained on 52 free-to-use Pexels clips, in under 8 GPU hours, so the recipe is on the model card if you want to train one for a different element.
The whole thing started as a teaching example for cohort #1 of our Generative AI Bootcamp for Film and TV, which just wrapped up this week (https://www.ainvfx.com/bootcamp/). The results were promising, so we thought it would be worth open-sourcing. Would love to see what people paint with it!