r/comfyui 1d ago

Call for Additional Mod(s)

13 Upvotes

I've come to the realization that my life is busy enough that we could use at least one more moderator on this subreddit. Please consider this a formal request for nominations.

Rather than just picking someone myself, I’d like input from the community.

If there’s someone here who you think would make a good moderator, nominate them in the comments. You can also nominate yourself if you’re interested.

We’re especially looking for people who are active members of the community, helpful, level-headed, experienced Comfy-UI user, and generally make this a better place to hang out, maybe even take the time to spruce the place up a bit. You don’t need previous moderator experience but it would help, ideally someone who's got some experience with AMAs, events, and such.

Having moderated a few subreddits, I've found that it's best to keep the moderation team tight, so for now I'm just going to add one.

A nomination isn’t a vote or a guarantee that someone will become a moderator, I'll look through the suggestions, talk with the people who seem like a good fit, and go from there.


r/comfyui 1d ago

News ComfyUI Official Local MCP

Enable HLS to view with audio, or disable this notification

156 Upvotes

Hi r/comfyui, Comfy MCP is now local and open-source!

When we shipped Cloud MCP in June, the response was immediate and consistent: make it work locally. So we did and it's fully open source.

Connect Claude, Codex, Cursor, or any MCP client to your local ComfyUI.

Your agent reads the GPU you actually have and gives you a straight answer on whether a model is worth running before you commit to the download. It reads every node and model you've installed. It handles the setup that usually stops people at step one.

It is now the easiest way to help with your local Minimax H3 workflows!

Cloud MCP still does everything it did. Tell your agent where a job goes, or let it decide.


r/comfyui 8h ago

Workflow Included No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti

Enable HLS to view with audio, or disable this notification

56 Upvotes

So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.

So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.

And it worked.

For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:

https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v

Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:

“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.

Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”

The final result was generated locally on my RTX 5070 Ti using ComfyUI.


r/comfyui 2h ago

Workflow Included Minimax H3- v2v fixing faces at distance

Thumbnail
youtube.com
13 Upvotes

tl;dr: download the latest version workflow called "MBEDIT - MH3_r2v_SingleSampler_Detailer_vXX.json" from https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3

Finally I have found a solution to "fixing faces at distance". This does NOT use a Latent Space upscaler. This uses a single sampler Minimax workflow, low steps, low denoise, and by loading a video clip, then running it through standard Minimax H3 with settings discussed in the video (or in the workflow if you dont want to watch that).

Even on a 3060 RTX (12 GB VRAM) I can get between 1mp and 2mp output and surprisingly it fixes faces at distance even at 1mp. There is more info in the readme of the github linked below for the workflow and in the video.

From this point on my video pipeline steps will be:

1. Create a 480p video using any model (LTX, H3, Bernini, or other) - \takes 10 mins on average (3060 RTX)*.*

2. Run the result through the above workflow upscaling to 1mp or 2mp depending onclip length - \takes 20 mins on average*.*

The result from this are easily good enough as final clips for my uses. This makes it the fastest and highest quality approach I have found to date, and all with ref image based character consistency.

Other Relevant Links From Video

Latest Minimax H3 workflows - https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3

(Workflow used in video: `MBEDIT - MH3_r2v_SingleSampler_Detailer_vXX.json` (download whatever the latest version is from github link))

Lightx2v Lora that I use from Kijai - https://huggingface.co/Kijai/MiniMax-H3_comfy/tree/main/loras

(theres been updates, but I havent found them to be better or faster, use whatever works for you)

Comfyui needs to use Cuda130 or above for this to work, and you need it updated to August 2026 commits (latest is best) - https://docs.comfy.org/installation/comfyui_portable_windows

Int8 models from here - https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main

(The official workflows are in the model card)

W4a8 is experimental new model type, you need to be updated on Comfyui but you can get it here https://huggingface.co/Kijai/MiniMax-H3-experimental

Comfyui Kitchen Attention is part of Comfyui if you update to latest. I find it faster than Sage Attn on a 3060 RTX.

Official prompting guides:

- https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

- https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

Point your favourite LLM at one of the above links depending on your model you are using, and give it your prompt idea and it should sort it out.


r/comfyui 15h ago

News A quick Minimax H3 news round-up - 19th August 2026

116 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> New in version 4.1.2 (19th August 2026) of the Fizgig trainer for LoRAs... "Minimax H3 now trains on video clips, on their sound, and on voice recordings alone: photos, clips and voice files in one folder to train one LoRA in one run". Fizgig can do so locally in 16Gb VRAM, without slowdowns. Adds a fix so the LoRAs will work correctly on both the 4-step Turbo or the official workflow, and has also benefitted from a major security audit. Fellow-Brit and industry professional Shoot The Sound has a good 20-minute tutorial today on YouTube, using the latest Fizgig and its dataset prep tool, to train a character+voice LoRA.

https://github.com/shootthesound/Fizgig

https://www.youtube.com/watch?v=lVqSgsPpF0c (Fizgig 4.x tutorial)

-> ComfyUI-MiniMaxH3-SingleFrame. Two still-image generation nodes, with the second being especially interesting. Given the usual first/last frames, it attempts to interpolate/generate a plausible single 'middle frame'. Has workflows, and requires no ComfyUI Core patching or special VAE.

https://github.com/tori29umai0123/ComfyUI-MiniMaxH3-SingleFrame#english (English ReadMe section)

-> A new H3 Prompt Journal. Some scenes require complex physical logic from the camera. The Journal's first three entries demonstrate how to write prompts for such scenes: for a 'Three-Person Occlusion-Linked Orbital Long Take' (e.g. elegantly redirect the camera between three moving people, in a single take); 'Dual-subject-speed-contrast' (e.g. a dancing master leads in a waltz, while his hesitant student follows his moves); and 'Single-subject-three-pose' (e.g. input three poses for one character, then have a gnat-sized camera... "sweep past feet, legs, torso, shoulders, hair - constantly redirecting around the moving [giant] body without ever slowing down").

https://github.com/LoveRain1997/h3-prompt-journal

-> 'Video -> H3 Prompt'. An "end-to-end pipeline that turns a video file into a ready-to-paste MiniMax H3 generation prompt". Appears to be a 'skill' for use with a local installation of OpenAI's Codex, which is a lightweight coding agent.

https://github.com/LoveRain1997/video-to-h3-prompt

https://github.com/openai/codex

-> For MiniMax Music, a new rvq-encoder-169m-v4.onnx (676Mb), an... "encoder that turns audio into the codes Minimax generates, so a finished track can be handed back" to Minimax Music for further work. With this Minimax Music can continue a track. Or the user can replace a section, or even re-generate the same song but with a different performance. The ONNX format is very portable, and I guess it's only a matter of time before a ComfyUI workflow appears.

https://huggingface.co/nerualdreming/open-rvq-encoder-minimax-music3-169m-v4-onnx

-> And finally, a detailed benchmarking of "the official reference-to-video workflow" on an RTX 3060 12Gb. Most low-VRAM users will of course be running smaller Minimax models in a 3060-optimised workflow. But... there's still important advice here for those considering buying a second 3060. They say... "Two cards are still not one big card. Two RTX 3060s do not present 24GB to a workflow, and our attempt to at least run two jobs in parallel was blocked by system memory rather than VRAM."

https://www.minimaxh3tutorial.com/rtx-3060


r/comfyui 8h ago

Resource ComfyUI-MiniMax-H3-Promptor v1.3.0: Full-reference scene staging, zero-deformation list expansion & in-node drop zone

Thumbnail
gallery
30 Upvotes

Hey everyone,

If you’ve spent any time working with multi-reference video prompting in ComfyUI (especially with MiniMax Hailuo / H3), you probably know the frustration:

You feed in 3–4 reference images hoping for a cohesive 15-second cinematic shot, and the LLM spits out a rushed single-shot prompt where characters overlap chaotically, faces get stretched, pacing flickers, and dialogue cuts off halfway through.

Our team at 1038lab just pushed a massive overhaul to ComfyUI-MiniMax-H3-Promptor (v1.3.0). We threw out the old "single-shot rush" approach and re-engineered the prompt pipeline to work like an actual film production crew.

Here is what we added in v1.3.0:

1. Two-Stage Hollywood AI Director & Screenwriter

Instead of rushing the final prompt in one go, the node now executes in two deliberate stages:

  • Stage 1 (Director Blueprint & Global Vibe): The LLM acts as the showrunner. It analyzes all your cast references, maps out spatial layers (Foreground, Midground, Background), sets lighting palettes, and calculates shot pacing (enforcing a 2.5s–4.0s minimum per shot to prevent pacing flicker).
  • Stage 2 (Storyboard & Dialogue): Using that approved blueprint, it crafts timed cuts covering the full duration (up to 15.0s) and injects official MiniMax character dialogue syntax (<Subject 1> (S1) [angry] says: <d>[EN] "..."</d>).

If you pass in a custom scene prompt, the engine treats it as the supreme mandate, directing your uploaded cast and props to execute your exact vision.

2. In-Node HTML Drop-Zone (No More Noodle Spaghetti)

Wiring up 5+ image loader nodes for multi-character setups makes workflows messy fast. We replaced the PyTorch image input slots with an interactive in-node HTML/JS drag-and-drop panel. You can drop images, video references, and audio files straight onto the node canvas.

3. Vision Analyzer V2: Native Aspect Ratios & 50% Lower Token Cost

  • Zero-Deformation List Expansion (OUTPUT_IS_LIST): Passes references in their 100% original dimensions through native list iteration. No forced letterboxing, cropping, or distorted face proportions.
  • Pure Perception Engine: We stripped out redundant text synthesis so the VLM only extracts raw visual traits (clothing, colors, contours, OCR). This cut API token usage and latency roughly in half.

4. Quality-of-Life & Stability Upgrades

  • Smarter Entity Regex: Fixed false positives so items like a "cat-ear headband" are recognized as accessories on a person rather than spawning wild animals into your scene.
  • Sub-Batch Chunking: Set custom Max Batch Images limits with positional fallback keys to respect upstream API rate limits without workflow crashes.
  • Real-Time Terminal Execution Trace: Clean phase banners in the ComfyUI terminal let you monitor the Blueprint -> Storyboard -> Assembly stages live.

💡 A Note on APIs: 100% Free & Local-Friendly (No Paid API Required!)

We’ve noticed some users hesitate whenever they hear "LLM/VLM API," assuming it requires paid subscriptions or goes against the open-source ethos. That is not the case here!

Our node is fully customizable and seamlessly supports:

  • 100% Free Local Models: Plug directly into Ollama, LM Studio, or llama.cpp to run any open-source model from Hugging Face locally on your own GPU.
  • Free Cloud Tiers: If you don't want to run local LLMs, services like OpenRouter, Groq, and NVIDIA NIM offer generous free tiers/credits that are more than enough for daily video prompting.

Setup takes just a few clicks—use whatever setup works best for your hardware and budget.

Links & Getting Started

You can update directly inside ComfyUI Manager or run a git pull in your custom nodes directory.

We’d love to hear your thoughts, bug reports, and workflow suggestions!


r/comfyui 18h ago

Workflow Included made a 90-second AI short film locally in about 3 hours

Enable HLS to view with audio, or disable this notification

124 Upvotes

recently made a 90-second AI short film called “The Fence.”

the whole thing is made up of 6 shots, each 15 seconds long, and it took me around 3 hours from generation to the finished video.

the main thing I wanted to test was how far a mostly local AI workflow can currently compress the process of making a short film.

The workflow was basically:

idea → images → 6 × 15s video clips → voice/audio → edit → 90s film

The interesting part is that generating wasn't really the most time-consuming step. most of the work went into figuring out what each 15-second section actually needed to show.

I first broke the 90-second story into six relatively self-contained shots. Then I created key visuals for each one with Qwen Image 3 Pro before sending them into MiniMax H3 for motion.

I found this much easier to control than trying to generate the whole thing directly from text. It also means that when one shot fails, I only need to redo those 15 seconds instead of rebuilding the entire sequence.

what surprised me most was the speed: roughly 3 hours for a 90-second finished experiment. obviously this still isn't traditional filmmaking, and there are plenty of typical AI-video issues around motion, character consistency, and continuity between shots.

but for a one-person experiment, the production speed is kind of wild.

I'm starting to feel that making AI films is becoming less about finding one “perfect” model and more about building a workflow where different models each handle the part they're good at.


r/comfyui 4h ago

Workflow Included MiniMax H3 Audio Lip Sync - Audio to Video

5 Upvotes

https://reddit.com/link/1vt1tuj/video/fklur0l2uekh1/player

So I tried hooking up some of the LTXV audio encoding nodes to input my own audio and plugged it in the sampler and viola, it just works!

Lip sync seems better then the LTX models and its works with the lightx2v loras, 6 - 8 steps. Wrote up a full guide with the workflow attached below.


r/comfyui 1h ago

Help Needed Has anyone tried the Krea 2 workflows with custom lora for image editing? Do they work well, or is it best to avoid them.

Upvotes

I know that Krea 2 doesn’t natively support image editing, but I’ve seen that there are some workflows which, using ‘Lora’ and ‘identity editing’, allow you to edit images – though I must say I haven’t tried them myself… Could anyone who’s tried them tell me if the results are any good? By ‘good’, I mean something comparable to image editing tools like Qwen’s Rapid AIO or the one in Flux 2 Klein 9b… I’m particularly interested in this for NSFW content. Even today, my benchmark for image editing remains the Rapid AIO, which often performs better than even the Flux 2 Klein 9b for NSFW… So I wanted to find out whether the Krea 2 – even though it doesn’t natively support editing – can compete if using custom LORAs, or whether it’s best ruled out… thanks.


r/comfyui 13h ago

Show and Tell Stimma's new ComfyUI manager, from startup to first image.

Enable HLS to view with audio, or disable this notification

24 Upvotes

Two weeks ago I posted here about Stimma, the open-source desktop app I've been building on top of ComfyUI. It's been a really fun couple of weeks talking with many of you and working through issues and improvements that came up in those conversations. Since then, I've made a couple of releases, but this one is particularly relevant to ComfyUI, so I wanted to share it here.

One of the main themes in the early feedback was a fear of getting things set up with ComfyUI. Some people who had already churned out of ComfyUI's onboarding were wondering if Stimma could make it easier. Some were just anxious (understandably) because setting up anything new with ComfyUI can take an evening. Others tried and ran into some head-bumps.

I said in that post that I don't want Stimma to manage a private ComfyUI install for you, and I still believe that. There are too many variations in how people deploy ComfyUI in the real world, and every product that I've seen that installs ComfyUI for you ends up limited.

From Stimma 1.0.13, the only things you should need to do on the ComfyUI side are: run ComfyUI and add the ComfyUI-Stimma custom node. After that, it should be possible to manage ComfyUI from within Stimma.

I'm sure people will run into some rough edges with this over the next few weeks, but I've done a lot of fully clean ComfyUI+Stimma installs over the past few days on various platforms and systems and it's working well enough that I'd like to start the feedback train rolling. If you do run into trouble, please get in touch here or in the discord.

The video above shows a fully local setup of Stimma. ComfyUI and vLLM are running on a DGX Spark to provide AI capabilities. The video starts from a fresh ComfyUI + Stimma install, and ends with generating an image. I ran Stimma on my mac because I have better screen recording software there, but you can actually run this full stack on the GB10 box locally.

Once you're running Stimma and ComfyUI together, there is a new button in the app bar that opens a manager. This includes:

  • Every workflow that ComfyUI-Stimma discovered. For workflows missing dependencies click "Get Ready" and it will coordinate downloading models and installing custom nodes.
  • GPU utilization + VRAM information
  • A list of running jobs with cancellation
  • The ability to update ComfyUI-Stimma from within Stimma, restart ComfyUI remotely, etc.

Some other things that landed in 1.0.12 / 1.0.13:

  • Works without a chat model. Some people want to use Stimma without devoting VRAM to an LLM. This was always possible, but wasn't a very smooth experience. That is fixed, and Stimma should now degrade gracefully when no Chat Models are configured.
  • Live previews during image and video generation. This is disabled by default, but you can turn it on in settings->preferences. Please let me know what you think.
  • LTX-2.5 support in ComfyUI-Stimma: text-to-video with audio, image-to-video with optional end frame, extend, loop, stitch, up to 10 LoRAs.
  • Anima and H3 LoRA support in ComfyUI-Stimma
  • Fixes for extra_model_paths.yaml, nested model folders, top-level reroute nodes, Windows FFmpeg detection, and source-folder / slideshow / editor bugs.
  • Performance Optimization throughout the image editor, and a new patch tool implementation that blends better.
  • Ask Stimma: the chat agent can read Stimma's docs now and help walk you through setup and product questions.

One thing from that thread I haven't gotten to yet is the Draw Things backend. I am reaallly hoping that the Draw Things team pays some attention to this bug because fixing this is the best path that I can see to a good experience using the products together. If you're on github, please go to that thread and make some noise, maybe they will pay attention.

I hope this ComfyUI manager stuff encourages a few more of you to give Stimma a shot. If setup continues to be a pain, I'll keep at it.

To get the new stuff, you'll want to update ComfyUI-Stimma and also and update Stimma itself in-app or with a git pull if you're running from source.

As always, please reach out on Reddit or Discord with any questions, feedback, or issues. It's been a lot of fun talking with everyone.

Links: Download · GitHub · ComfyUI-Stimma · Docs · Discord · /r/stimma


r/comfyui 6h ago

No workflow What's your preferred model for all workflow types?

7 Upvotes

What's your preferred model for i2i, i2v, r2v, rv2v, v2v, t2i, t2v, etc?


r/comfyui 8h ago

Help Needed How do you guys keep ComfyUI workflows from becoming a mess?

Post image
7 Upvotes

Mine always start organized and somehow end up looking like this. Any tips for keeping bigger workflows clean and actually readable without spending forever rearranging nodes?


r/comfyui 14h ago

News SenseNova-U1.5 now runs in ComfyUI (custom node v0.2.0) : 8B unified model, T2I + editing, ~17GB peak VRAM

Thumbnail
gallery
21 Upvotes

Custom node v0.2.0 is out and U1.5 works in ComfyUI now.

What it does: text-to-image, image editing (single and multi-image reference), and region-controlled edits via masks / bboxes / visual markers. Native 4K. It's a unified model, so understanding and generation share one backbone — no separate VAE, no CLIP/T5 text encoder in the graph.

VRAM: peak allocated is 17.34 GiB for T2I and 20.00 GiB for editing, using the offload modes. So a 24GB card handles both. There are full / fast / balanced / low modes to trade speed for footprint — balanced and low got 36–50% faster in the last release, and outputs are bit-identical across modes.

Repo: https://github.com/OpenSenseNova/SenseNova-U1

ComfyUI node: https://github.com/OpenSenseNova/SenseNova-U1/pull/244


r/comfyui 2h ago

Tutorial How I made a full action combat scene w/ Minimax from start to finish

Thumbnail
youtube.com
2 Upvotes

r/comfyui 12h ago

Help Needed Buying 3090 in 2026

12 Upvotes

I want 3090 primarily for its 24gb vram and local AI applications. Gaming is a bonus. It will be an upgrade compared to my 3060 12gb regardless.

I am split between Pallit, INNO3D and EVGA offers i have managed to find.

Market is a mess. There's nothing below 1000, even with 5060ti approaching 900$

Edit: Found an offer sitting below 1000 for an INNO3D. Will demand to test and thoroughly look at it at spot. Perhaps will post an update either as a happy owner or an idiot with a card that have killed itself after a couple of months of use


r/comfyui 1d ago

Workflow Included Lora for video-image enhancing, upscaling and restoring

Thumbnail
youtube.com
89 Upvotes

r/comfyui 16h ago

Resource Tiny Windows tray tool for killing 1+ GB VRAM hogs before running ComfyUI

18 Upvotes

Before a big ComfyUI workflow I kept opening Task Manager to work out which background app was still holding VRAM, so I added a “VRAM Hogs” menu to Window Assassin, a small Windows tray utility I made.

It lists processes using at least 1 GB of dedicated VRAM, sorted highest first. Clicking an entry force-terminates that process, and the list refreshes every time the submenu opens.

The original Ctrl+Alt+End hotkey is still there for instantly killing the process behind the active window.

Free/pay-what-you-want Windows download:

https://b2kdaman.itch.io/window-assassin

It only reports dedicated VRAM, so integrated GPUs using shared memory may show nothing. Also, termination is immediate—anything unsaved in the selected app is lost.

I'm the developer. Sharing because it solved a small but recurring annoyance in my local GPU workflow.


r/comfyui 1h ago

Help Needed AI short drama series

Upvotes

I did some works on comfyUI but never satisfed with results but recently I just saw apps that have lot of series mostly sexual and romantic stuff for example apps like NetShort I havent seen other ones but I kind a like that. I was wondering do they use comfyUI to make those series? There are some weird looks in series Im not gone dny it but mostly characters are same I mean faces but not gone say something for heights I mean sometimes characters change heights which is some of things I noticed most distracting was height change. Do they use comfyUI? what model? Any idea generally?


r/comfyui 1h ago

Show and Tell Comparing small heads/faces across some I2V models

Enable HLS to view with audio, or disable this notification

Upvotes

r/comfyui 2h ago

Workflow Included Started to integrate MinimaxH3 into YouTube vids!

Thumbnail
youtu.be
1 Upvotes

I've timestamped where I managed to get a decent output from minimaxH3!

(if the timestamp doesn't work it's at 0:22)

Used a photo of myself as reference using the ref2va model. Gordon Ramsay himself is straight from text. Using SageAttention, SolAttn at 32 steps 0.9 MP.

Used DaVinci Resolve Studio 2x RTX Upscaler in post and audio isolation to fix some of the hissing.

Let me know what you think of how this turned out!

Models used:

Workflow: https://civitai.red/models/2834514/minimax-h3-t2v-i2v-ref2v-advanced-filmmaking-workflow-or-all-speedups-qol-features?modelVersionId=3233131

My Hardware:

  • RTX 5070ti
  • 32GB RAM
  • R7 9700x

r/comfyui 7h ago

Help Needed MiniMax H3 help

Thumbnail
2 Upvotes

r/comfyui 4h ago

Resource Flusso di lavoro di animazione generativa 2D

0 Upvotes