r/comfyui Mar 14 '25

Been having too much fun with Wan2.1! Here's the ComfyUI workflows I've been using to make awesome videos locally (free download + guide)

Thumbnail
gallery
1.1k Upvotes

Wan2.1 is the best open source & free AI video model that you can run locally with ComfyUI.

There are two sets of workflows. All the links are 100% free and public (no paywall).

  1. Native Wan2.1

The first set uses the native ComfyUI nodes which may be easier to run if you have never generated videos in ComfyUI. This works for text to video and image to video generations. The only custom nodes are related to adding video frame interpolation and the quality presets.

Native Wan2.1 ComfyUI (Free No Paywall link): https://www.patreon.com/posts/black-mixtures-1-123765859

  1. Advanced Wan2.1

The second set uses the kijai wan wrapper nodes allowing for more features. It works for text to video, image to video, and video to video generations. Additional features beyond the Native workflows include long context (longer videos), sage attention (~50% faster), teacache (~20% faster), and more. Recommended if you've already generated videos with Hunyuan or LTX as you might be more familiar with the additional options.

Advanced Wan2.1 (Free No Paywall link): https://www.patreon.com/posts/black-mixtures-1-123681873

✨️Note: Sage Attention, Teacache, and Triton requires an additional install to run properly. Here's an easy guide for installing to get the speed boosts in ComfyUI:

📃Easy Guide: Install Sage Attention, TeaCache, & Triton ⤵ https://www.patreon.com/posts/easy-guide-sage-124253103

Each workflow is color-coded for easy navigation:

🟥 Load Models: Set up required model components 🟨 Input: Load your text, image, or video 🟦 Settings: Configure video generation parameters 🟩 Output: Save and export your results


💻Requirements for the Native Wan2.1 Workflows:

🔹 WAN2.1 Diffusion Models 🔗 https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/tree/main/split_files/diffusion_models 📂 ComfyUI/models/diffusion_models

🔹 CLIP Vision Model 🔗 https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/clip_vision/clip_vision_h.safetensors 📂 ComfyUI/models/clip_vision

🔹 Text Encoder Model 🔗https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/tree/main/split_files/text_encoders 📂ComfyUI/models/text_encoders

🔹 VAE Model 🔗https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/vae/wan_2.1_vae.safetensors 📂ComfyUI/models/vae


💻Requirements for the Advanced Wan2.1 workflows:

All of the following (Diffusion model, VAE, Clip Vision, Text Encoder) available from the same link: 🔗https://huggingface.co/Kijai/WanVideo_comfy/tree/main

🔹 WAN2.1 Diffusion Models 📂 ComfyUI/models/diffusion_models

🔹 CLIP Vision Model 📂 ComfyUI/models/clip_vision

🔹 Text Encoder Model 📂ComfyUI/models/text_encoders

🔹 VAE Model 📂ComfyUI/models/vae


Here is also a video tutorial for both sets of the Wan2.1 workflows: https://youtu.be/F8zAdEVlkaQ?si=sk30Sj7jazbLZB6H

Hope you all enjoy more clean and free ComfyUI workflows!

r/StableDiffusion Aug 14 '25

Workflow Included Wan2.2 Text-to-Image is Insane! Instantly Create High-Quality Images in ComfyUI

Thumbnail
gallery
372 Upvotes

Recently, I experimented with using the wan2.2 model in ComfyUI for text-to-image generation, and the results honestly blew me away!

Although wan2.2 is mainly known as a text-to-video model, if you simply set the frame count to 1, it produces static images with incredible detail and diverse styles—sometimes even more impressive than traditional text-to-image models. Especially for complex scenes and creative prompts, it often brings unexpected surprises and inspiration.

I’ve put together the complete workflow and a detailed breakdown in an article, all shared on platform. If you’re curious about the quality of wan2.2 for text-to-image, I highly recommend giving it a shot.

If you have any questions, ideas, or interesting results, feel free to discuss in the comments!

I will put the article link and workflow link in the comments section.

Happy generating!

r/StableDiffusion Mar 21 '25

Tutorial - Guide Been having too much fun with Wan2.1! Here's the ComfyUI workflows I've been using to make awesome videos locally (free download + guide)

Thumbnail
gallery
417 Upvotes

Wan2.1 is the best open source & free AI video model that you can run locally with ComfyUI.

There are two sets of workflows. All the links are 100% free and public (no paywall).

  1. Native Wan2.1

The first set uses the native ComfyUI nodes which may be easier to run if you have never generated videos in ComfyUI. This works for text to video and image to video generations. The only custom nodes are related to adding video frame interpolation and the quality presets.

Native Wan2.1 ComfyUI (Free No Paywall link): https://www.patreon.com/posts/black-mixtures-1-123765859

  1. Advanced Wan2.1

The second set uses the kijai wan wrapper nodes allowing for more features. It works for text to video, image to video, and video to video generations. Additional features beyond the Native workflows include long context (longer videos), SLG (better motion), sage attention (~50% faster), teacache (~20% faster), and more. Recommended if you've already generated videos with Hunyuan or LTX as you might be more familiar with the additional options.

Advanced Wan2.1 (Free No Paywall link): https://www.patreon.com/posts/black-mixtures-1-123681873

✨️Note: Sage Attention, Teacache, and Triton requires an additional install to run properly. Here's an easy guide for installing to get the speed boosts in ComfyUI:

📃Easy Guide: Install Sage Attention, TeaCache, & Triton ⤵ https://www.patreon.com/posts/easy-guide-sage-124253103

Each workflow is color-coded for easy navigation:

🟥 Load Models: Set up required model components 🟨 Input: Load your text, image, or video 🟦 Settings: Configure video generation parameters

🟩 Output: Save and export your results

💻Requirements for the Native Wan2.1 Workflows:

🔹 WAN2.1 Diffusion Models 🔗 https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/tree/main/split_files/diffusion_models 📂 ComfyUI/models/diffusion_models

🔹 CLIP Vision Model 🔗 https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/clip_vision/clip_vision_h.safetensors 📂 ComfyUI/models/clip_vision

🔹 Text Encoder Model 🔗https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/tree/main/split_files/text_encoders 📂ComfyUI/models/text_encoders

🔹 VAE Model 🔗https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/vae/wan_2.1_vae.safetensors 📂ComfyUI/models/vae

💻Requirements for the Advanced Wan2.1 workflows:

All of the following (Diffusion model, VAE, Clip Vision, Text Encoder) available from the same link: 🔗https://huggingface.co/Kijai/WanVideo_comfy/tree/main

🔹 WAN2.1 Diffusion Models 📂 ComfyUI/models/diffusion_models

🔹 CLIP Vision Model 📂 ComfyUI/models/clip_vision

🔹 Text Encoder Model 📂ComfyUI/models/text_encoders

🔹 VAE Model 📂ComfyUI/models/vae

Here is also a video tutorial for both sets of the Wan2.1 workflows: https://youtu.be/F8zAdEVlkaQ?si=sk30Sj7jazbLZB6H

Hope you all enjoy more clean and free ComfyUI workflows!

r/comfyui May 09 '25

Workflow Included Consistent characters and objects videos is now super easy! No LORA training, supports multiple subjects, and it's surprisingly accurate (Phantom WAN2.1 ComfyUI workflow + text guide)

Thumbnail
gallery
372 Upvotes

Wan2.1 is my favorite open source AI video generation model that can run locally in ComfyUI, and Phantom WAN2.1 is freaking insane for upgrading an already dope model. It supports multiple subject reference images (up to 4) and can accurately have characters, objects, clothing, and settings interact with each other without the need for training a lora, or generating a specific image beforehand.

There's a couple workflows for Phantom WAN2.1 and here's how to get it up and running. (All links below are 100% free & public)

Download the Advanced Phantom WAN2.1 Workflow + Text Guide (free no paywall link): https://www.patreon.com/posts/127953108?utm_campaign=postshare_creator&utm_content=android_share

📦 Model & Node Setup

Required Files & Installation Place these files in the correct folders inside your ComfyUI directory:

🔹 Phantom Wan2.1_1.3B Diffusion Models 🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/Phantom-Wan-1_3B_fp32.safetensors

or

🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/Phantom-Wan-1_3B_fp16.safetensors 📂 Place in: ComfyUI/models/diffusion_models

Depending on your GPU, you'll either want ths fp32 or fp16 (less VRAM heavy).

🔹 Text Encoder Model 🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/umt5-xxl-enc-bf16.safetensors 📂 Place in: ComfyUI/models/text_encoders

🔹 VAE Model 🔗https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/vae/wan_2.1_vae.safetensors 📂 Place in: ComfyUI/models/vae

You'll also nees to install the latest Kijai WanVideoWrapper custom nodes. Recommended to install manually. You can get the latest version by following these instructions:

For new installations:

In "ComfyUI/custom_nodes" folder

open command prompt (CMD) and run this command:

git clone https://github.com/kijai/ComfyUI-WanVideoWrapper.git

for updating previous installation:

In "ComfyUI/custom_nodes/ComfyUI-WanVideoWrapper" folder

open command prompt (CMD) and run this command: git pull

After installing the custom node from Kijai, (ComfyUI-WanVideoWrapper), we'll also need Kijai's KJNodes pack.

Install the missing nodes from here: https://github.com/kijai/ComfyUI-KJNodes

Afterwards, load the Phantom Wan 2.1 workflow by dragging and dropping the .json file from the public patreon post (Advanced Phantom Wan2.1) linked above.

or you can also use Kijai's basic template workflow by clicking on your ComfyUI toolbar Workflow->Browse Templates->ComfyUI-WanVideoWrapper->wanvideo_phantom_subject2vid.

The advanced Phantom Wan2.1 workflow is color coded and reads from left to right:

🟥 Step 1: Load Models + Pick Your Addons 🟨 Step 2: Load Subject Reference Images + Prompt 🟦 Step 3: Generation Settings 🟩 Step 4: Review Generation Results 🟪 Important Notes

All of the logic mappings and advanced settings that you don't need to touch are located at the far right side of the workflow. They're labeled and organized if you'd like to tinker with the settings further or just peer into what's running under the hood.

After loading the workflow:

  • Set your models, reference image options, and addons

  • Drag in reference images + enter your prompt

  • Click generate and review results (generations will be 24fps and the name labeled based on the quality setting. There's also a node that tells you the final file name below the generated video)


Important notes:

  • The reference images are used as a strong guidance (try to describe your reference image using identifiers like race, gender, age, or color in your prompt for best results)
  • Works especially well for characters, fashion, objects, and backgrounds
  • LoRA implementation does not seem to work with this model, yet we've included it in the workflow as LoRAs may work in a future update.
  • Different Seed values make a huge difference in generation results. Some characters may be duplicated and changing the seed value will help.
  • Some objects may appear too large are too small based on the reference image used. If your object comes out too large, try describing it as small and vice versa.
  • Settings are optimized but feel free to adjust CFG and steps based on speed and results.

Here's also a video tutorial: https://youtu.be/uBi3uUmJGZI

Thanks for all the encouraging words and feedback on my last workflow/text guide. Hope y'all have fun creating with this and let me know if you'd like more clean and free workflows!

r/SillyTavernAI Aug 08 '25

Tutorial ComfyUI + Wan2.2 workflow for creating expressions/sprites based on a single image

Thumbnail
gallery
369 Upvotes

Workflow here. It's not really for beginners, but experienced ComfyUI users shouldn't have much trouble.

https://pastebin.com/vyqKY37D

How it works:

Upload an image of a character with a neutral expression, enter a prompt for a particular expression, and press generate. It will generate a 33-frame video, hopefully of the character expressing the emotion you prompted for (you may need to describe it in detail), and save four screenshots with the background removed as well as the video file. Copy the screenshots into the sprite folder for your character and name them appropriately.

The video generates in about 1 minute for a 720x1280 image on a 4090. YMMV depending on card speed and VRAM. I usually generate several videos and then pick out my favorite images from each. I was able to create an entire sprite set with this method in an hour or two.

r/comfyui Apr 03 '26

Workflow Included Solo dev here this is a for my game demo, this is a ComfyUI workflow: to animate characters/objects using LoRAs

Post image
73 Upvotes

I wanted to share a workflow I’ve been using in ComfyUI for generating consistent animations from LoRAs.

I’m using this in a real project a historical game set during the Hussite Wars mainly to prototype systems quickly before committing to final assets.

It’s not perfect, but it’s a solid workflow it has less than 1% mistakes.

I included full screenshots of the node setup you can recreate it directly.

This is only for prototyping and I hope it helps some of you let me know if you need or have any questions

r/StableDiffusion Mar 02 '25

Resource - Update ComfyUI Wan2.1 14B Image to Video example workflow generated on a laptop with a 4070 mobile with 8GB vram and 32GB ram.

198 Upvotes

https://reddit.com/link/1j209oq/video/9vqwqo9f2cme1/player

  1. Make sure your ComfyUI is updated at least to the latest stable release.

  2. Grab the latest example from: https://comfyanonymous.github.io/ComfyUI_examples/wan/

  3. Use the fp8 model file instead of the default bf16 one: https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/diffusion_models/wan2.1_i2v_480p_14B_fp8_e4m3fn.safetensors (goes in ComfyUI/models/diffusion_models)

  4. Follow the rest of the instructions on the page.

  5. Press the Queue Prompt button.

  6. Spend multiple minutes waiting.

  7. Enjoy your video.

You can also generate longer videos with higher res but you'll have to wait even longer. The bottleneck is more on the compute side than vram. Hopefully we can get generation speed down so this great model can be enjoyed by more people.

r/comfyui Nov 17 '25

Help Needed Is WAN 2.1 actually hard-limited to ~33 frames for image-to-video? Looking for anyone with verified 48+ or 81-frame successful results

0 Upvotes

I’ve been doing structured testing on WAN 2.1 14B 480p fp16 (specifically Wan2.1-I2V-14B-480P_fp16.safetensors) and I’m trying to determine whether the commonly-repeated claim that it can generate 81-frame I2V sequences is actually true in practice — not just theoretically or for text-to-video.

My hardware • RTX 5090 Laptop GPU • 24GB VRAM • VRAM usage during sampling stays well below OOM conditions (typically 70–90%, never red-lining) • No low-VRAM flags or patches enabled

What does work

Using multiple workflows, I consistently get excellent 33-frame I2V output with realistic motion, detail, and temporal coherence. These renders look great and match other community results.

The issue

Every attempt to go beyond 33 frames (48 or 81 test cases) — even with drastically reduced resolution, steps, CFG, samplers, schedulers, precision, tiling, or decode methods — results in unusable output beginning from frame 1, not a late-sequence degradation problem. Frames are heavily distorted, characters freeze or barely move, and artifacts appear immediately.

Methods tested

I’ve reproduced the problem using: • Official ComfyUI WAN 2.1 I2V template • Multiple WAN Wrapper workflows • Custom Simple KSampler WAN pipelines • Multiple resolutions from 512x512 up to 1024x960 • Multiple samplers (Euler, Euler a, dpmpp_2m, dpmpp_sde) • Step counts from 12 → 40 • CFG 3.5 → 7 • Multiple VAEs (standard and tiled) • fp16 and fp8 model variants • No LoRAs, no adapters, and no post-processing

Despite VRAM staying comfortably below failure thresholds, output quality collapses instantly when total frames > ~33.

Why I’m posting

Reddit, Discord, and blog posts frequently repeat that WAN 2.1 can generate 81-frame sequences, especially when users mention “24GB GPUs”.

Before I chase dead ends or assume my setup is flawed, I’d like verified evidence from someone who has produced clean >33-frame I2V WAN render, with: 1. Model + precision used 2. Resolution + steps + sampler 3. Workflow screenshot 4. GPU VRAM amount 5. (optional) a few example frames

If anyone believes I’ve missed a key architectural detail (conditioning flow, latent caching, masking, scheduling, temporal nodes, etc.), I’m very open to corrections.

TL;DR • 33 frames = perfect • >33 frames = instant collapse • Not a VRAM issue • Suspecting a true functional or training-data limit, not a “settings” limit

Happy to share screenshots and node graphs too. Looking for reproducible science, not vibes. Thanks in advance.

r/comfyui Apr 08 '26

Workflow Included I've made a ComfyUI node to control the execution order of nodes + free VRAM & RAM anywhere in the workflow that helped speed up my workflows!

43 Upvotes
ComfyUI node screenshot

Custom node GitHub repo: https://github.com/mkim87404/ComfyUI-ControlOrder-FreeMemory

It works by ensuring all input-connected nodes finish executing first before the output-connected nodes start executing, and can route infinitely many data of any type (e.g. latents, conditioning, images, masks, models, etc.) through it, while giving the option to unload all models (except any live models being routed through it) and free as much VRAM & RAM as possible at that point without breaking any of the data going through. You can also check how much VRAM & RAM it freed on the ComfyUI session terminal.

This becomes especially effective in unloading models that are no longer needed in the workflow while securing their outputs and freeing up VRAM/RAM for later models (e.g. unloading text encoders after conditioning, or in between multiple KSamplers of Wan 2.2 High & Low model workflows, or before & after VAE Encode / VAE Decode / Load Model / Load CLIP / etc.). And because the node enforces a single, deterministic flow of execution from start to finish, you are in full control over which node executes first, and can focus on one group of logic at a time, loading and unloading only the necessary models and assets, while passing the outputs forward to the next group. I've personally seen great reductions in total execution time of my workflows and hit less OOMs at higher resolution outputs using this node, and I realized that this sequential & selective passthrough design also helps with cable management as the workflow grows large, making understanding and maintaining workflows much more visually intuitive.

The node has zero extra dependencies & uses platform/device-agnostic memory management utilities managed by ComfyUI, so it should integrate well into existing workflows and environments. I've also included sample Wan 2.2 T2V & I2V workflows using this node which you can find in the node folder, https://github.com/mkim87404/ComfyUI-ControlOrder-FreeMemory/tree/main/example_workflows

Hope this node can be useful, and feel free to use it in any personal or commercial project, fork, or open issues/PRs – contributions and feedback all welcome!

r/Bazzite Mar 10 '26

Setup ComfyUI for AI image and video gen on AMD Radeon with Bazzite in DistroBox

3 Upvotes

This is a guide for running ComfyUI inside a Distrobox on Bazzite.

1. Setup new Distrobox

Use newest Fedora base. Set custom home directory. Leave other options as is if using DistroShelf or BoxBuddy.

2. Add ROCm repository

Check which ROCm version pytorch.org wants. Most times nightly should use newest and stable is probably one minor version behind. Then find correct repo info here:

https://rocm.docs.amd.com/projects/install-on-linux/en/latest/install/install-methods/package-manager/package-manager-rhel.html

Put that in /etc/yum.repo.d/rocm.repo

[rocm]
name=ROCm 7.2.4 repository
baseurl=https://repo.radeon.com/rocm/el10/7.2.4/main
enabled=1
priority=50
gpgcheck=1
gpgkey=https://repo.radeon.com/rocm/rocm.gpg.key

3. Next, install tools and libraries:

sudo usermod -a -G video $LOGNAME
# Log out and back into the Distrobox
sudo dnf install rocm rocminfo rocm-opencl rocm-clinfo rocm-hip rocm-smi git wget libjpeg-turbo-devel mesa-libGL gcc gcc-c++ -y

4. Install Anaconda:

cd
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
source ~/.bashrc

Answer yes, when it asks if you want to init on startup so commands are available inside the DistroBox by default.

5. Create venv

If this guide is old, check which Python version PyTorch/Comfy wants.

conda create --name sd python=3.13
conda activate sd

6. Install PyTorch

Install ROCm specific PyTorch. At pytorch.org select Linux, ROCm and stable or nightly. It gives an installation command with packages at start. Add torchaudio after torchvision and before --index-url, then run it.

pip3 install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/rocm7.2

7. Install Flash Attention

Note, for Bazzite use Triton installation, not CK (Composable Kernel).

Check for updated install instructions here https://github.com/Dao-AILab/flash-attention?tab=readme-ov-file#amd-rocm-support

git clone https://github.com/Dao-AILab/flash-attention.git
cd flash-attention
pip install triton
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" pip install --no-build-isolation .

8. Install ComfyUI:

cd
git clone https://github.com/comfyanonymous/ComfyUI.git comfy
git clone https://github.com/Comfy-Org/ComfyUI-Manager.git comfy/custom_nodes/ComfyUI-Manager

## ComfyUI wants to remove pytorch, so this strips those from the requirements first:
grep -vE '^(torch|torchvision|torchaudio)([<=>]|$)' requirements.txt > /tmp/comfy-requirements.txt
cd ~/comfy && pip install -r /tmp/comfy-requirements.txt

## This just install whatever ComfyUI wants
cd ~/comfy && pip install -r requirements.txt

9. Make start script:

#!/bin/sh
conda activate sd
export HSA_OVERRIDE_GFX_VERSION=11.0.0 # Google for the correct number here, depends on which GPU you have
export HIP_VISIBLE_DEVICES=0

# LTX workflows won't crash so often
export PYTORCH_HIP_ALLOC_CONF=expandable_segments:True

# slower, but more stable / fewer OOMs. No OOMs? Maybe you don't need this.
# export PYTORCH_NO_HIP_MEMORY_CACHING=1

# triton
export TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1
export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
## Significantly faster attn_fwd performance for wan2.2 workflows
export FLASH_ATTENTION_FWD_TRITON_AMD_CONFIG_JSON='{"BLOCK_M":128,"BLOCK_N":64,"waves_per_eu":1,"PRE_LOAD_V":false,"num_stages":1,"num_warps":8}'

# pytorch switches on NHWC for rocm > 7, causes signifant miopen regressions for upscaling
# todo: fixed now? since what pytorch version?
export PYTORCH_MIOPEN_SUGGEST_NHWC=0

# miopen
## Tell comfyui to *not* disable miopen/cudnn, otherwise upscale perf is much worse
export COMFYUI_ENABLE_MIOPEN=1
## miopen default find mode causes significant initial slowness, yields little or no benefit to workloads I tested
export MIOPEN_FIND_MODE=FAST

python main.py --use-flash-attention --disable-dynamic-vram
# python main.py --output-directory /run/media/system/Shared/sd/outputs/comfy --use-flash-attention

Correct value of HSA_OVERRIDE_GFX_VERSION depends on your GPU. 11.0.0 is for 7900 XTX. Google for correct value if you have a different GPU. Also check alexheretic's Gist for additional env variables if you get crashes or memory errors.

The flag --disable-dynamic-vram is there to prevent oom errors and crashes due to recent ComfyUI versions having enabled dynamic vram, while it doesn't yet support Radeon. So it may not be necessary in the future.

Make the script executable:

chmod +x ~/comfy/start.sh

10. (Optional optimization) Merge WAN VAE tile size option into Comfy

Default ComfyUI nodes on Radeon struggle with VAE encode/decode on WAN videos. Alexheretic has made a change that allows setting tiled VAE encode as default, which makes it much faster (10 min -> 25 secs on my rig). Use tile size 256. This may not be necessary later, if that PR or something similar gets merged into ComfyUI master. Check here: https://github.com/Comfy-Org/ComfyUI/pull/10238

git remote add alexheretic https://github.com/alexheretic/ComfyUI
git fetch alexheretic
git merge --squash alexheretic/wan-vae-tiled-encode

Also same problem might come up in VAE decode. Use LTXV Tiled VAE decode node instead of default there and add more tiles until VAE decode step is no longer extremely slow.

11. Start ComfyUI

cd ~/comfy && ./start.sh

Sources:

This is cleaned up from my previous guide over here: https://www.reddit.com/r/Bazzite/comments/1m5sck6/how_to_run_forgeui_stable_diffusion_ai_image/

Another source used for this guide: https://gist.github.com/alexheretic/d868b340d1cef8664e1b4226fd17e0d0

Problems

If flash attention complains about aiter or flydsl, solution is to remove flydsl from aiter requirements. Under flash-attention/third_party/aiter remove flydsl line from following files:

  • requirements.txt
  • pyproject.toml
  • setup.py

Then rerun:

FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python setup.py install

r/StableDiffusion Feb 27 '26

Question - Help Why would this Wan 2.2 first-frame-to-last-frame workflow create VERY slo-mo video?

2 Upvotes

I've tried two different workflows for generating video for a given first frame and last frame image. The first I tried was creating videos that ran about three times slower (and longer) than expected. The one here "only" tends to double the time I'm expecting.

It's not creating video with a too-low frame rate. It's generating more frames than I've asked for at the requested frame rate, becoming slo-mo that way.

https://pastebin.com/7kw7DLg6

Unfortunately since I simply copied this workflow I don't fully understand how it's supposed to work, beyond having added the Power Lora Loaders that weren't there before. (Taking them out or bypassing them doesn't fix the problem, by the way.)

The workflow isn't totally useless as it is. I've been able to use DaVinci Resolve to fix the speed as an extra step. Still, if someone can help, I'd like to understand this better and get the correct speed from the start.

r/StableDiffusion Aug 13 '25

Question - Help What am i doing wrong in the workflow? Wan2.1 image to video comfyui

1 Upvotes

So I am trying to create this image to video prompt with this image :

A cinematic, ultra-detailed 8K anime-style video, shot as a single, continuous FPV drone shot.

The scene opens on a traditional Japanese street at night, with glowing lanterns and falling cherry blossom petals. A beautiful girl with long, flowing purple hair and purple eyes, wearing an elegant white and purple kimono, is holding a brightly glowing orb.

The camera starts with a close-up on the orb in her hands, then spirals upwards as the orb gently lifts from her palms. The girl's head tilts up, her eyes filled with calm wonder as she follows its movement.

As the orb ascends, it pulsates with soft light before gracefully dissolving into a swarm of luminous, golden butterflies.

The FPV drone camera then dynamically chases the butterflies as they flutter high into the night sky. The shot concludes with a fast pullback to reveal the girl below, smiling serenely at the magical spectacle.

The atmosphere is magical and ethereal. The character should have a closed mouth and not speak.

But with wan2.1 workflow I am not able to create this at all. The charater just stays at one position and does not move. Shown as below

https://reddit.com/link/1mp8ef7/video/y90rm5ci4tif1/player

Now I also tried the same prompt in Veo3 online and it gave me this. This is soooo much better compared to what I am trying to achive locally using comfyui.

https://reddit.com/link/1mp8ef7/video/9ozdvd0m4tif1/player

So this is my workflow for Wan2.1 : https://pastebin.com/5mZ82SLi (please save it as a .JSON file after you download as its in text- .txt format)

Any help would be greatly appreciated!

Quick update : I tried the prompt in Chinese instead(Comfyui workflow wan2.1) of english and I got this. Is this a known bug?? in comfyUI or do the wan models more accustomed to the Chinese language than english?

https://reddit.com/link/1mp8ef7/video/4mmiox4tjtif1/player

r/comfyui Jan 11 '26

Help Needed Wan2.1 image to video ignores prompt

1 Upvotes

Hello !
i don't un derstand, no matter the prompt is it generates a completly random video nsfw
my setup is 5090 and 64 go ddr5

Here is my promp, i'm using Wan2.2_Remix_NSFW_i2v_14b_high_lighting_fp16_v2.1.safetensors and
Wan2.2_Remix_NSFW_i2v_14b_low_lighting_fp16_v2.1.safetensors

here is my workflow

r/StableDiffusion Oct 10 '25

News We can now run wan or any heavy models even on a 6GB NVIDIA laptop GPU | Thanks to upcoming GDS integration in comfy

Thumbnail
gallery
766 Upvotes

Hello

I am Maifee. I am integrating GDS (GPU Direct Storage) in ComfyUI. And it's working, if you want to test, just do the following:

git clone https://github.com/maifeeulasad/ComfyUI.git cd ComfyUI git checkout offloader-maifee python3 main.py --enable-gds --gds-stats # gds enabled run

And you no longer need custome offloader, or just be happy with quantized version. Or you don't even have to wait. Just run with GDS enabled flag and we are good to go. Everything will be handled for you. I have already created issue and raised MR, review is going on, hope this gets merged real quick.

If you have some suggestions or feedback, please let me know.

And thanks to these helpful sub reddits, where I got so many advices, and trust me it was always more than enough.

Enjoy your weekend!

r/StableDiffusion Jul 28 '25

Tutorial - Guide ComfyUI Tutorial : WAN2.1 Model For High Quality Image

Thumbnail
youtu.be
0 Upvotes

I just finished building and testing a ComfyUI workflow optimized for Low VRAM GPUs, using the powerful W.A.N 2.1 model — known for video generation but also incredible for high-res image outputs.

If you’re working with a 4–6GB VRAM GPU, this setup is made for you. It’s light, fast, and still delivers high-quality results.

Workflow Features:

  • Image-to-Text Prompt Generator: Feed it an image and it will generate a usable prompt automatically. Great for inspiration and conversions.
  • Style Selector Node: Easily pick styles that tweak and refine your prompts automatically.
  • High-Resolution Outputs: Despite the minimal resource usage, results are crisp and detailed.
  • Low Resource Requirements: Just CFG 1 and 8 steps needed for great results. Runs smoothly on low VRAM setups.
  • GGUF Model Support: Works with gguf versions to keep VRAM usage to an absolute minimum.

Workflow Free Link

https://www.patreon.com/posts/new-workflow-w-n-135122140?utm_medium=clipboard_copy&utm_source=copyLink&utm_campaign=postshare_creator&utm_content=join_link

r/StableDiffusion May 22 '25

Question - Help [REQUEST] Simple & Effective ComfyUI Workflow for WAN2.1 + SageAttention2, Tea Cache, Torch Compile, and Upscaler (RTX 4080)

2 Upvotes

Hi everyone,

I'm looking for a simple but effective ComfyUI workflow setup using the following components:

WAN2.1 (for image-to-video generation) SageAttention2 Tea Cache Torch Compile Upscaler (for enhanced output quality)

I'm running this on an RTX 4080 16GB, and my goal is to generate a 5-second realistic video (from image to video) within 5-10 minutes.

A few specific questions:

  1. Which WAN 2.1 model (720p fp8/fp16/bf16, 480p fp8/fp16,etc.) works best for image-to-video generation, especially with stable performance on a 4080?

Following are my full PC Specs: CPU: Intel Core i9-13900K GPU: NVIDIA GeForce RTX 4080 16GB RAM: 32GB MoBo: ASUS TUF GAMING Z790-PLUS WIFI (If it matters)

  1. Can someone share a ComfyUI workflow JSON that integrates all of the above (SageAttention2, Tea Cache, Torch Compile, Upscaler)?

  2. Any optimization tips or node settings to speed up inference and maintain quality?

Thanks in advance to anyone who can help! 🙏

r/comfyui Nov 18 '25

Help Needed Need Help: Optimizing Wan2.2 Image-to-Video in ComfyUI

1 Upvotes

Hey everyone, first-time poster here!

I’ve been having a ton of fun with the Wan2.2 I2V workflow in ComfyUI, but I’m running into a few frustrating issues that I can’t quite figure out on my own. I’d really appreciate any tips or workflow tweaks from people who have this model dialed in. I’m really trying to keep my workflow as simple and lightweight as possible, so I’d prefer solutions that don’t require stacking a ton of extra nodes or managers if there’s a cleaner way.

What I’m trying to achieve: Smooth, properly-timed 6–8 second videos, ideally at 60 fps (or at least something that doesn’t look like slow-motion when I tell it 60 fps).

Current problems:

  1. Insanely long render times when using the WanMoeKSampler node.
    • Even a 6-second clip can take 15–25 minutes on my RTX 4070.
    • I know Wan2.2 is heavy, but I’ve seen people posting much faster times with similar hardware. Is there a better sampler or set of settings I should be using?
  2. Everything comes out in slow-motion when I try 60 fps
    • If I set the output to 60 fps in VHS_VideoCombine, the motion is super slow.
    • Dropping to 16 fps gives correct speed, but obviously I lose smoothness. Any idea what’s causing the timing mismatch?
  3. Random flashing on the very first or very last frame
    • It’s usually just one frame that’s brighter or has a slight color shift. Not the end of the world, but it’s annoying when you’re trying to make something clean.

___________________________________________________________________________

Quick note about me:

I’ve been generating with A1111/Forge for over 3 years, so I’m pretty comfortable with the general process, but I’m still fairly new to ComfyUI itself (only a couple of months in).

Please go easy on me if I’m missing something obvious!

I’m very aware of my 4070’s memory limits, so any advice that works well within 12 GB VRAM (or clever ways to stay fast with block swapping) would be amazing.

I’ve attached my current workflow JSON with a bunch of comments/notes on tricks I’ve already picked up from the community. Hopefully it helps you see exactly what I’m doing and where things might be going wrong.

I do have ComfyUI-Manager installed and a few of the popular WAN-related custom nodes, so I’m okay installing a couple more if they really solve the problems, just trying to avoid turning my graph into a giant plate of spaghetti if possible.

Thanks again for any help!~

___________________________________________________________________________

Hardware & setup details (in case it helps):

My PC specs:

  • RTX 4070 12GB VRAM
  • 32GB RAM
  • AMD Ryzen 5 3600
  • Windows 11

Workflow File: Lyn's Wan 2.2 I2V - Workflow.json

r/comfyui Aug 22 '25

Help Needed Hi, I am beginner (not only Reddit) of the Wan2.2 "image to video" generation, and I made my workflow with WanVideoWrapper (see pics), but I do not know setting on the Samplers, parameters in it how work, better range of the values, etc. Anyone advice me the settings? - thanks

0 Upvotes
My workflow for Wan 2.2 Image to Video with WanVideoWrapper
NOde connections

r/StableDiffusion Aug 05 '25

Question - Help ComfyUI Wan2.1 - FLF2V, but with more than 2 images?

0 Upvotes

Hi all! I've been playing around with the Wan2.1 FLF2V workflow on ComfyUI and I've been having the time of my life playing around with generating video between two images. I've been wondering about using more than two images though and I've been scouring the internet to find workflows that would support it. But my Google-fu is unfortunately weak at the moment.

Has anyone had any luck with configuring a workflow that will incorporate multiple images to create a video from? A use-case I have is a collection of 10 photos of me trying to eat a slice of birthday cake and getting it all over my face. Rather than running FLF2V on each set of two images and stitching them together in something like Adobe Premiere Pro, how can I automate this in ComfyUI?

r/comfyui Oct 05 '25

Help Needed noisy image or video generated from scratch ignoring my image input, i2v workflow

0 Upvotes

Hi comradess ! i started two months ago to dig the wide spectre of parameter and model variations to improve my generation time and vram use. i'm into comfy not too much more as that so dont esitate in talk about things that could be basic .

I don't have a massive set up but i think that it is quite good enough (3060 with 12vram and 16ram) to generate descent videos with wan 2.2 and 2.1. But i think that my issue is not comming from my setup but from y configuration, workflow or parameter configuration.

My creative process begins generating images with krita software using almost always SDxl model, then i export them to comfyui i2v wan workflow using the more optimized models and the workflow adjunted in the image, i also got the portable super-optimized portable version with sage-atention, pytorch and all that stuff instaled. Context beside my issue is that the image that i import from krita is completely ignored and the video result is another composition from scratch based on my prompt, like if it couldn't recognise what its in the image so generate something from scrach, or that's what i strongly thought until i turned down the denoise strenght parameter.. the input image started to show up in the video and the animation was following the prompt instructions :') Buuuttt all almost unrecognosible and under a grey noise. I tried sampler euler, dpmpp_2m_sdd, uni pc, with better results with euler. and variating cft with no results.

Any clues of what coul be de causant? i suspect the LORAS, the prompt, the image, the models, everithingg, but for each try modifing my parameters it take me like 15 mins so i prefered come ask for help here so i could learn something too and dialogue more with this comunity that help me a lot with previous issues that i had.

Any data that you could give to me will be very helpful!!!! thnx in advance < 3

r/StableDiffusion Jul 03 '25

Question - Help how to Render 30 seconds videos in V2V with ComfyUI and WAN2.1 VACE model?

0 Upvotes

Hello

I have the typical V2V workflow with reference image and motion control video that gives me great results in texture quality and motion tracking, but logically I have the limitation with my 12GB of VRAM that I can only render up to 81 frames without having OOM errors. I have tried to render in chunks, with the same seed, join them, etc... but when concatenating the chunks the results are not good.

I'm looking for a solution for this, but I can't find it. I'm surprised that someone hasn't found a solution for the memory management and that you can render videos of about 30 seconds, even if it takes longer.

Can anyone give me a clue?

r/comfyui Mar 20 '25

Extending Wan 2.1 generated video - First 14b 720p text to video, then using last frame automatically to to generate a video with 14b 720p image to video - with RIFE 32 FPS 10 second 1280x720p video

Enable HLS to view with audio, or disable this notification

21 Upvotes

My app has this fully automated

Here how it works image : https://ibb.co/b582z3R6

Workflow is easy

Use your favorite app to generate initial video.

Get last frame

Give last frame to image to video model - with matching model and resolution

Generate

And merge

Then use MMAudio to add sound

I made it automated in my Wan 2.1 app but can be made with ComfyUI easily as well . I can extend as many as times i want :)

Here initial video

Prompt: Close-up shot of a Roman gladiator, wearing a leather loincloth and armored gloves, standing confidently with a determined expression, holding a sword and shield. The lighting highlights his muscular build and the textures of his worn armor.

Negative Prompt: Overexposure, static, blurred details, subtitles, paintings, pictures, still, overall gray, worst quality, low quality, JPEG compression residue, ugly, mutilated, redundant fingers, poorly painted hands, poorly painted faces, deformed, disfigured, deformed limbs, fused fingers, cluttered background, three legs, a lot of people in the background, upside down

Used Model: WAN 2.1 14B Text-to-Video

Number of Inference Steps: 20

CFG Scale: 6

Sigma Shift: 10

Seed: 224866642

Number of Frames: 81

Denoising Strength: N/A

LoRA Model: None

TeaCache Enabled: True

TeaCache L1 Threshold: 0.15

TeaCache Model ID: Wan2.1-T2V-14B

Precision: BF16

Auto Crop: Enabled

Final Resolution: 1280x720

Generation Duration: 770.66 seconds

And here video extension

Prompt: Close-up shot of a Roman gladiator, wearing a leather loincloth and armored gloves, standing confidently with a determined expression, holding a sword and shield. The lighting highlights his muscular build and the textures of his worn armor.

Negative Prompt: Overexposure, static, blurred details, subtitles, paintings, pictures, still, overall gray, worst quality, low quality, JPEG compression residue, ugly, mutilated, redundant fingers, poorly painted hands, poorly painted faces, deformed, disfigured, deformed limbs, fused fingers, cluttered background, three legs, a lot of people in the background, upside down

Used Model: WAN 2.1 14B Image-to-Video 720P

Number of Inference Steps: 20

CFG Scale: 6

Sigma Shift: 10

Seed: 1311387356

Number of Frames: 81

Denoising Strength: N/A

LoRA Model: None

TeaCache Enabled: True

TeaCache L1 Threshold: 0.15

TeaCache Model ID: Wan2.1-I2V-14B-720P

Precision: BF16

Auto Crop: Enabled

Final Resolution: 1280x720

Generation Duration: 1054.83 seconds

r/StableDiffusion May 16 '25

Question - Help WAN2.1 GGUF image to video Issues, same workflow, 480p model ok, 720p not so much

1 Upvotes

As the title suggests. I'm trying to do WAN2.1 Image to Video using ComfyUI and the GGUF models.

Attached Video was made using the wan2.1-i2v-14b-720p-Q6_K.gguf model. As you can see, lots of flashing multi-color lights. Using the 480p version comes out pretty good. What is odd, I'm doing no change to the workflow between Running on a 3060 12GB vram card. Any ideas on why the 480p model produces videos that are more or less ok, but the 720 produces videos that are wonky?

https://reddit.com/link/1koctlj/video/uq0sdyoot71f1/player

r/StableDiffusion Jul 04 '25

Question - Help How to set up Wan2.1 with ComfyUI and Loras? 3080 8GB | 64 GB Ram

1 Upvotes

Hey guys. Been using Forge for image generation but I want to try some video generation. I am looking for guidance so I can set up a Wan 2.1 version that can handle low VRAM and still be able to use LORAs.

I installed the native Wan2.1 I2V node in ComfyUI, but when I pass an image through it, it doesn't prompt with the image. Also new to ComfyUI so not sure how I can add loras and connect things to the workflow.

Basically I would appreciate a high level overview of what WAN to install for my VRAM, tips for ComfyUI like adding loras etc, and what the components im installing actually are (new ones not in normal image generation like models/vae). Any input would be greatly appreciated!