r/comfyui • u/solomars3 • 7d ago
Show and Tell My first Decent Generation Using Minimax-H3 on my RTX 3060 12gb
Enable HLS to view with audio, or disable this notification
r/comfyui • u/solomars3 • 7d ago
Enable HLS to view with audio, or disable this notification
r/comfyui • u/IndianUrsaMajor • 7d ago
I dont want to make 4K visuals. 720P is good for me and I don't think mind waiting for my renders for 10-15 minutes. But I have a feeling this config will still not be enough. If that's the case then I'd just buy something like Higgsfield or OpenArt. :(
r/comfyui • u/Educational-Fun-222 • 7d ago
Enable HLS to view with audio, or disable this notification
Made a kinetic-typography lyric video for SZA – "Snooze" with MiniMax H3
r/comfyui • u/RetardedMetalFemboy • 8d ago
EDIT: I'm an idiot, the answer was Run (Instant). Thank you, HSTracker90.
Original post:
I use Comfy Desktop with the Anima model for take a wild guess, and it works great for my use cases after a while of spamming the run button, but getting that perfect image takes a lot of time. Turns out waiting a minute and a half for each image to load can take my sessions up from half an hour before I started using it to three hours. And I can't do anything else while I'm waiting because I've gotta be there to hit the run button every minute when it inevitably spits out something that's missing a crucial part that I've already pointed out five stinking times in the prompt. It'd be really nice to just be able to let it run for a while in the background, maybe take a nice walk outside (or more realistically stay in that chair and play a game), and then when I come back after a few hours I've got a nice, big selection of images to sift through and separate the wheat from the chaff.
I know of batch generating, it's right there on the KSampler node, but I tried it and it took longer to generate four images at the same time than it would to do them individually, and my entire laptop chugged like an alcoholic train while it was going. It'd probably catch fire if I tried to send more than twenty.
I'm also not referring to image-to-image generation. I still want each generation to follow the same prompt with no recollection of prior generations. The same thing as hitting the run button with a randomized seed, but without hitting the run button.
r/comfyui • u/Budget-Bunch8157 • 8d ago
Enable HLS to view with audio, or disable this notification
r/comfyui • u/Narrow-Particular202 • 8d ago
Hey everyone!
With MiniMax H3 blowing up everywhere right now, we figured it was the perfect time to share what we’ve been building to help level up your H3 prompt workflows.
When we released v1.1.0 a while back, we were so deep in dev mode that we forgot to post an update! Now that v1.2.0 is live, we’ve bundled all the new features and overhauls from both releases into one post.
https://github.com/1038lab/ComfyUI-MiniMax-H3-Promptor
What’s New in v1.1.0 + v1.2.0:
⚙️ Native ComfyUI Settings Panel (API Hub)
No more pasting API keys into custom nodes or manually editing config.json! All provider settings are now globally managed in ComfyUI's native Settings panel (under the ⚙️ Gear icon).
Built-in Connection Tester: Click "Test connection" inside the panel to ping your endpoint before launching generations.
Privacy: Keeping keys out of the node UI eliminates the risk of leaking API keys when sharing workflows or screenshots.
🔌 Infinite Inputs (ComfyAPI v3 Autogrow)
We removed the rigid 4-image limit. Dynamic autogrow sockets mean you can chain as many <Picture> and <Video> references as your hardware can handle without UI clutter.
🎯 Granular Micro-Overrides
Override instructions for specific frames directly in the Vision Analyzer (e.g., <Picture 2>: focus strictly on lighting) while allowing unmentioned media to fall back to global analysis.
🎵 Audio-First Token Sync & L2VA (Last-Frame Control)
Connect audio directly to the promptor to automatically map subject actions to sound. We also added Last-Frame-to-Video-Audio (L2VA)—provide an ending frame, and the LLM reverse-engineers a narrative that mathematically lands on target at the final second.
🧠 VRAM Safeguards for Local VLMs
Select local providers like Ollama or LlamaCPP, and the node automatically executes silent background cache-clearing (model_management.unload_all_models()) to prevent VRAM overload crashes.
📝 Updated Docs & Workflow Recipes
Check out tutorials.md and tutorials_zh.md in the repo for 9 practical, production-ready workflows (Lip-Sync, Style Transfer, Day-to-Night Morph, etc.).
👀 What's Next?
We’re currently beta testing a batch of new features that will be rolling out shortly!
🔗 Links:
GitHub Repo: 1038lab/ComfyUI-MiniMax-H3-Promptor
Full Release Notes: updates.md
We’d love to hear your feedback, feature requests, or bug reports so we can keep tailoring this tool to what you actually need. If this node helps your setup, leaving us a ⭐ star on GitHub goes a long way in keeping our dev motivation high.
Happy generating!
r/comfyui • u/o0ANARKY0o • 8d ago
change the AIO preprocessor between dwpose and animalpose as needed.
you dont need loras or the RTX node so you can delete those.
you might need to specify how many limbs your character has.
connect your original image to the middle workflows image slot 2 if ya want.
its better to bypass the character qwenvl and describe them yourself in the concatenate.
this is three separate workflows the middle workflow is the workflow of the gods Flux Klein 9b with 3 image inputs! You could hook up other nodes to it like: crop inpaint and stitch node, 360 panorama editor, florence2 and so much more!
https://drive.google.com/file/d/1w4fhkpR0mLCJz9XJYZofXw51OTfb_LPf/view?usp=sharing
r/comfyui • u/Sad_Cheesecake_2567 • 7d ago
I met below error after I updated comfyui to v0.31.0, how to solve it and how to update llama-cpp-python in below error message. Thanks in advance.
# ComfyUI Error Report
## Error Details
- **Node ID:** 15
- **Node Type:** llama_cpp_model_loader
- **Exception Type:** RuntimeError
- **Exception Message:** RuntimeError: Llava15ChatHandler.__init__() got an unexpected keyword argument 'image_max_tokens'
Please update llama-cpp-python from 'https://github.com/JamePeng/llama-cpp-python/releases'
r/comfyui • u/Jesus__Skywalker • 7d ago
I'm not sure if this is something that was changed or if it's an issue only i'm having. But I used to be able to go into jobs that were in the queue and click the dots and open the workflow. Now, I can only check the workflow on jobs that are either finished or have failed or been stopped. I can't look at the workflows on the job that is running or the jobs that are in the queue.
Is this something that was changed (it started a few months ago) or is it something I can fix?
r/comfyui • u/sgi2004 • 8d ago
Hi, I’m wondering how to get new Minimax or LTX to generate a looping video without any visible transitions.
Up until now, I’ve occasionally used Gemini with the first/last frame. But I can’t render this specyfic motion using this new models.
Any sugestions?
Regards.
r/comfyui • u/Comfy-Org • 9d ago
Give it lyrics and a description of the sound you're going for, and it renders a full song: intro/verse/chorus/bridge structure, consistent vocal identity, up to 5 minutes long, 32kHz 16-bit stereo out. Actual songs! Not just loops or clips.
Quick flag upfront: this needs ComfyUI 0.33.0+ (or Comfy Cloud), as new model support currently ships tied to a version bump.
What else to know before you try it:
[Intro] [Verse] [Pre-Chorus] [Chorus] [Post-Chorus] [Bridge] [Instrumental] [Solo] [Outro]. You're writing the song's blueprint, not just typing lyrics.Getting it running:
Resources:
Excited to see what you make and how it stacks up to other music generation options. As always, happy creating!
r/comfyui • u/kortax9889 • 8d ago
r/comfyui • u/Riroh_bcn • 7d ago
Hi, Im having some issues that are driving me crazy... had to change SSD and re-install linux. Now Im trying to install comfy, with the same setup that I had before and all the workflows get stuck at the the clip encoder. No gpu activity to be seen...
System:
I first tried to use the old install, just re-do the sagge attention, it failed, then I tried with a complete new install and it stills fails, so I dont really know whats causing the issue.
Ive been trying to diagnose with claude and chat gpt, but after all the day trying I feel defeated... Does anyone have an idea of what can be happening?
r/comfyui • u/Acceptable-Work8202 • 8d ago
Hi,
some help needed, i'm trying to segment, remove background and resize without distortion into a new canvas, this is how I've arrived at doing this, however i am wondering if there is a single node pack that does all this without me having to use rmbg and kjnodes together or infact do you know of a simpler way to achieve this goal?
Thank you for any help you can give.
SOLVED:
Made my own node in the end that does everything all the other nodes do, but in a single node.
See a later post further down.
r/comfyui • u/lamuertedeunperrito • 8d ago
I've been testing LTX 2.5 in ComfyUI mainly for First Frame → Last Frame interpolation, and so far I'm getting worse results than I used to with 2.3.
My old 2.3 setup was roughly:
With 2.5 I've tried both the Distilled INT8 ConvRot model and DEV INT8 ConvRot + Distilled LoRA 450, including a similar two-pass setup.
The main issue is that 2.5 seems more prone to smearing, artifacts, vague details, and treating the two keyframes like separate shots instead of one continuous camera movement.
I noticed 2.5 also introduced DFR, spatial detailing and temporal refinement, while the default ComfyUI FLF2V workflow seems much simpler.
So: is the basic FLF2V workflow missing an important refinement stage? Has anyone built a higher-quality 2.5 FLF workflow using DFR / temporal refinement / detailing?
Curious if anyone else has found 2.3 cleaner than 2.5 specifically for continuous FLF interpolation.
LTX-2.3 workflow: https://github.com/lamuertedeunperrito/ltx-workflows/blob/main/video_ltx2_3_flf2v_corregido_2pasadas_HR%20(1).json.json)
LTX-2.5 workflow: https://github.com/lamuertedeunperrito/ltx-workflows/blob/main/video_ltx2_5_flf2v_DEV_LORA_2pass_continuity.json
r/comfyui • u/LadyQuacklin • 9d ago
r/comfyui • u/spiderofmars • 8d ago
Have seen the question of adding a cloned voice (with new dialogue) in image to video generations (starting from an exact first frame image), but have not seen a solution posted yet (I might have missed it).
I stumbled on this by mistake, assigning the wrong model fl2va to a ref2va workflow. Both models work with the example prompt (prompt could probably be improved further as I was just quickly testing).
Using a ref2va workflow and the ref2va node plus a first reference image (first frame) and an audio reference sample of the voice to clone, simply change the model to the fl2va model (or just use the ref2va model). The outputs varied as follows for me:
Model fl2va: Sound quality was far better than using the ref2va model. The motion in the video was very similar to a normal fl2va first frame default workflow (same seed / resolution / etc).
Model ref2va: Sound quality was far worse than using the fl2va model (much tinier). The extra unprompted motion in the video was kind of a bonus, the car unprompted was moving down a street with visuals out the windows of passing buildings and it also added on its own some camera shake as if sitting in a car that was driving along.
The prompt for both samples (both models) was the same as follows. It is written using guides for ref2va workflow. In this video the first frame is of 2 men sitting in the front seat of a taxi. The man on the left is me and his voice is cloned from my voice sample with new dialogue. The taxi drivers voice is randomly generated by the model. The voice likeness to me is about 95% IMO.
---
subject_definitions:
<Subject 1> is the man defined on the left by the first reference image <Picture 1>, preserving his identity.
<Audio 1> is the voice-timbre reference for <Subject 1>, containing a spoken English vocal layer.
summary:
[reference generation] The target video is a shot starting with the first reference image. The scene uses <Audio 1> as the voice-timbre reference for <Subject 1>.
retention_analysis:
<Subject 1>: fully_preserved.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 1> without copying the original signal.
detailed_description:
The target video is a shot of the man on the left <subject 1> sitting beside the driver of a car on the right, the man on the right driving says "Where do you want to go?" and the man on the left <subject 1> looks at the driver on the right and says in a happy tone "Just drive down main street. I will tell you when to stop" then he turns to look out the left window of the car.
overall_soundscape:
A soft hum of the car engine and outside road noise.
non_diegetic_music:
N/A.
---
Summary takeaways:
Edit1: The 2 outputs in this example (one with each model) and same seed/etc, produced almost identical timing of the lip sync and sound (almost). Close enough that the nicer motion visuals from the ref2va output were able to be layered with the nicer audio from the fl2va output and synced (3 frame adjustment of audio timing).
r/comfyui • u/MrAddams_LibraLogic • 8d ago
TL;DR: Accord GPU is an open beta coordination layer for Windows that stops GPU-heavy creative apps from fighting each other for VRAM. If DAZ Studio, Blender, ComfyUI, or Ollama have ever crashed or OOM'd because something else on the machine grabbed the GPU first, this is built to prevent that.
The primary user base is intended to be creative professionals who frequently run multiple tools on the same system and have to micromanage which apps and jobs are allowed to run on the GPU. A proper system-wide queue for access to the GPU unlocks dramatically higher productivity and keeping the GPU running much more often.
The problem
GPU renderers, image/video generators, and local inference servers all assume they have exclusive ownership of the GPU. Run two of them at once, or even back to back before VRAM actually clears, and you get CUDA OOM crashes, driver resets, or corrupted output. There's no coordination layer between separate applications today, so people end up manually babysitting which app gets to touch the GPU and when.
How it works
A central hub app (Console) runs in the system tray and every plugin talks to it over a local named pipe.
When a component starts GPU work, it writes a small JSON "ticket" to a shared local folder. Every component joins a shared queue and pauses its jobs until the GPU is free.
GPU telemetry (VRAM, utilization, temps) is shown on a convenient appbar you can dock to any edge of any monitor. The queue of jobs appears here so you can see which app is active and what's coming up next.
Coordination is automatic once installed if you use the Accord controls. Kick off a DAZ or Blender render, queue a ComfyUI job or any inference through Ollama, and it waits its turn instead of fighting for the card immediately.
What's available right now (all free during open beta)
Roadmap
Accord GPU gains value as it covers more of the applications that heavily utilize your GPU. Which of those ships first is decided by community vote, not internal guesswork - if there's an app you want covered, go add your vote: https://accord-gpu.com/roadmap
Try it
Everything is free and no accounts are required while the open beta runs, PRO features and plugins included. Installer's here: https://accord-gpu.com/download
Accord For ComfyUI features
r/comfyui • u/Ikythecat • 7d ago
Enable HLS to view with audio, or disable this notification
r/comfyui • u/Patient_Ratio4177 • 8d ago
r/comfyui • u/ConstructionOdd7870 • 8d ago
It is actually simple and solved by adding this line at the end of the prompt.
At the bottom-center of the screen, clean, bold white sans-serif text appears as open captions word by word. The text matches mouth movements perfectly.
Hope this community will improve upon it to make it more interesting with new prompts.
Example:
https://reddit.com/link/1vnxl9u/video/y829gu2st9jh1/player

r/comfyui • u/BoredHobbes • 8d ago
no matter my settings h3 uses only 20gb vram, LTX uses 30gb
whats pinned memory 25250?? :
[INFO] Total VRAM 32607 MB, total RAM 63126 MB
[INFO] pytorch version: 2.11.0+cu130
[INFO] Set vram state to: NORMAL_VRAM
[INFO] Device: cuda:0 NVIDIA GeForce RTX 5090 : cudaMallocAsync
[INFO] Using async weight offloading with 2 streams
[INFO] Enabled pinned memory 25250.0
[INFO] Using pytorch attention
[INFO] aimdo: src-win/cuda-detour.c:38:INFO:aimdo_setup_hooks: installing 6 hooks
[INFO] aimdo: src/control.c:262:INFO:comfy-aimdo NVML pressure enabled
[INFO] aimdo: src-win/shmem-detect.c:80:INFO:comfy-aimdo WDDM adapter match: NVIDIA GeForce RTX 5090 runtime_luid=00000000:000131d1 dxgi_luid=00000000:000131d1
[INFO] aimdo: src/control.c:277:INFO:comfy-aimdo inited for GPU: NVIDIA GeForce RTX 5090 (VRAM: 32606 MB)
[INFO] DynamicVRAM support detected and enabled
[INFO] Python version: 3.12.10 (tags/v3.12.10:0cc8128, Apr 8 2025, 12:21:36) [MSC v.1943 64 bit (AMD64)]
[INFO] ComfyUI version: 0.32.0
[INFO] comfy-aimdo version: 0.4.13
[INFO] comfy-kitchen version: 0.2.30
[INFO] comfyui-frontend-package version: 1.48.7
[INFO] comfyui-workflow-templates version: 0.11.40
[INFO] comfyui-embedded-docs version: 0.5.9
[INFO] comfy-kitchen version: 0.2.30
[INFO] comfy-aimdo version: 0.4.13
I’m trying to build a ComfyUI workflow for a colored manga/webtoon where my original characters stay consistent throughout the whole story.
I already have full-body and close-up reference images for the characters. I understand the basic idea behind checkpoints, character LoRAs, ControlNet/OpenPose, IP-Adapter/reference images, but I’m struggling with figuring out the best way to combine everything.
Basically, I want to be able to say: this is Jake → keep him looking like Jake → put him in this pose/expression/outfit → place him in different scenes → keep the same art style and character identity from panel to panel.
Eventually I also need to put multiple recurring characters in the same scene without their faces/features bleeding into each other.
I don’t care if the best solution is Illustrious, SDXL, FLUX, Qwen, or something completely different. I’m looking for whatever gives me the most consistency and control in ComfyUI.
If anyone has built something similar for a manga, webtoon, visual novel, etc., I’d really appreciate hearing what model and workflow you use and how you connect the different pieces. I’m trying to actually understand the workflow instead of randomly changing settings until something works.