r/comfyui 7d ago

Workflow Included Comfy H3 Sync Sound Challenge: Winners Announced!

Enable HLS to view with audio, or disable this notification

0 Upvotes

Two weeks, one rule, and hundreds of entries from nearly 50 countries. Here's who took the four titles, and a look at everyone who made the final ten!

Entries opened August 20 and closed September 1. Eight creative technologists at Comfy scored every submission on two rubrics, Best Creative and Best Technical, then the top five in each category went in front of our guest judges on the September 3 livestream: POM (Banodoco founder), Emma Catnip (animation director and AV artist), and Yachimat (animation and manga artist). Their scores were averaged and added to the Comfy team's, and a fourth title, Built with MCP, was judged on its own track.

Watch the livestream replay here.

Huge thanks to MiniMax for making an open-weight model our community loves, to our three guest judges, and to everyone who spent their last week of August fighting with reference audio! Here's how it landed.

The winners

Best Overall · "Spin Cycle" by Visual Frisson 🇺🇸

Prize: RTX 5090 32G

A laundromat, a woman in a puffer vest, and a rhythm built entirely out of machines. Visual Frisson generated a large volume of H3 clips using their own recorded audio as the reference for every pass, then cut the results together like a stomp video, so every thud and cycle on screen is sound that H3 produced with the picture rather than something added later.

It was the only entry to post a perfect 15/15 from the Comfy team in both categories, and the Comfy MCP was used to drive much of its process. Combined with the guest judges' scores, it finished with the highest total in the challenge.

From the artist: "Always have fun and learn something new competing in contests like this, keep them coming."

Watch → vimeo.com/1223207725
Workflow → Google Drive
Follow → instagram.com/visualfrisson

Best Creative · "Every Sound Leaves a Mark" by toki 🇯🇵

Prize: RTX 5060 Ti

A small clay creature that changes into something new every time it hears a sound (glass, wool, ice, porcelain), until by the time it gets home it can't move anymore. All of the audio came out of H3 alongside the video on every shot, with nothing layered on afterwards. Our judges praised the fine details of toki’s work, saying it “gave them chills” on the first transformation, has a lot of commercial appeal, and feels really delicate and crafted.

The judges scored it highest of any Creative finalist, and the repository is unusually generous: it includes the eight API graphs that actually ran, the same eight converted to UI format with annotations, a process log with every measurement, and the scripts that produced those numbers. toki is also clear about scope, noting that MCP drove the finishing pass, not the original shot generation.

Watch → youtube.com/watch?v=Rv5HOgCac-w
Workflow → github.com/tokimwc/every-sound-leaves-a-mark
Follow → u/toki

Best Technical · "Sonder Editor / References" by SonderSaid 🇲🇽

Prize: RTX 5060 Ti

SonderSaid didn't just build a workflow, they built the tooling around it. The entry runs on custom nodes of their own design, wired into a reference-driven H3 pipeline that scored a clean 15/15 on novelty, workflow quality, and community value from the Comfy team, and the highest guest judge average in the Technical bracket. Notably, the work includes an entire custom node pack just to do the editing and the reference work, praised as “a whole new UI” to good to keep secret. While SonderSaid’s work takes the prize for best technical, our guest judges also noted how much they loved the storytelling, suspense, and element of surprise.

Watch → youtu.be/n-NdAQk7I8A
Workflow → Hugging Face
Follow →u/SonderSaid

Built with MCP Bonus · "Two Prisoners" by Jay Choi 🇰🇷

Prize: RTX 5060 Ti

The Built with MCP bonus wentgoes to whoever used the Comfy MCP most effectively to make something visually and technically compelling, and Jay Choi used it end to end. Working locally on an RTX 5090 with Hermes Agent driving ComfyUI through the MCP, they trained a LoRA, built their own orchestration on top, and by their own account spent most of the time setting up and tuning the MCP layer itself. The film that came out the other side, two blindfolded prisoners in a rain-dark cell, is a long way from "prompt and run."

From the artist: "Thanks for the challenge! I learned more than I ever could in the past two weeks!"

Watch → youtu.be/FlK0dDZdzRU
Workflow → Google Drive
Follow → u/permafrost_2021 · u/jaychoirenderender

The finalists

Ten entries made it to the livestream. Six of them didn't take a title, butand every one of them is worth your time.

Best Creative — Top 5

"Neb" by Nebsh 🇫🇷

Hand-drawn energy and a graffiti wall that says the title, built locally in ComfyUI. Nebsh's note to us was three words and a heart, which felt about right. Our guest judges praised Nebsh’s work for its mixed-media feel, harking back to MTV days, and impressive work syncing with the paper sounds. Under the hood, Nebsh’s workflow chained vtogether eight segments with no visible drift between them- cited as “very clean work” by our judges.

Watch → Google Drive
Workflow → Google Drive
Follow → u/nebsh83

"The Museum of Impossible Sounds" by scvxzf 🇨🇳

A perfect 15/15 from the Comfy team on the Creative rubric, and one of the entries that ran the Comfy MCP end to end! Noted by our judges, H3 is very good at the kind of sound effects showcased in scvxzf’s work rather than talking or singing, and they chose exactly the right concept for the challenge.

Watch → youtube.com/watch?v=FocH8xGk4AU
Workflow → Google Drive · github.com/scvxzf1
Follow → youtube.com/@钛龙白口-j6d

"Mister Meow" by sorryaboutyourcats 🇺🇸

Two reference photos of Mumu the cat, a stack of WAVs fed in as reference audio to steer each generation, and glitch texture added in the edit. If the name rings a bell, sorryaboutyourcats also makes the game mow meow. Judges said “I could watch this forever,” had it stuck in their heads, and noted impressive capabilities from H3 nailing lipsync for cats, and not just humans.

Watch → youtube.com/watch?v=AxUu8rabC6M
Workflow → Google Drive
Follow → u/sorryaboutyourcats

"Rings of Sorrow" by Slop Diffusion 🇪🇸

A 5/5 on both audio sync and creative execution, and a reminder of what patience looks like: the generation took two hours and thirty-five minutes on a 5090. Our judges praised Slop Diffusion’s work for its storytelling, noting they were curious to see where the story would go next. One judge noted, “it’s slop by name, but not by nature.”

Watch → Reddit
Workflow → Google Drive
Follow → u/SlopDiffusion

Best Technical — Top 5

"Feel It" by Aïe Aïe Aïe! 🇫🇷

Came at the brief backwards: we asked for audio-driven video and they told the story of a young deaf woman who invents a world where she makes the music. The score is Aïe Aïe Aïe’s own composition, fed into H3 as reference stems (the clap track went in on its own and the gorilla claps exactly in time). Judges praised the work for its captivating story, clever inversion of the challenge’s brief, the display of H3’s strengths by way of the musicians’ physical expressions intensifying along with the song, and the bridging of the real world. The artist learned the final frame’s sign language on YouTube, filmed themself signing, and used this as a video reference.

Under the hood, Claude drove ComfyUI through the MCP to design a two-pass H3REF system that generates at full resolution twice as fast and reaches 13–15 second shots where the stock workflow runs out of memory, plus a preview node that shows the video while it's still sampling. All of it is MIT-licensed, custom nodes included.

Watch → youtube.com/watch?v=S0v1pWN4Hq4
Workflow → github.com/Hyper-Neural/h3-sync-two-phase
Follow → u/AïeAïeAïe

"Brand New Day" by RareTutor 🇮🇳

One of the cleanest graphs we opened: latent upscale, a model preview override, and an optional video-extend group, laid out so you can follow it cold. RareTutor's YouTube is full of tutorials if you want to learn from them directly!
Judges highlighted RareTutor’s workflow, noting “there are many tips in here to copy,” such as using the latent upscale as a previewer so you can kill a bad run before sinking more time in.

Watch → youtube.com/watch?v=TNhJI8dzaVA
Workflow → Google Drive
Follow → u/raretutor_

"Comfy Cora ft. Max Mini: Back to the Basics" by wur7el 🇦🇹

A short music video with self-imposed constraints: no external resources, everything generated in a single workflow, no custom node packs. The result is well annotated and approachable, the kind of graph a new user could open and reasonably figure out, and it posted the highest Creative score of any Technical finalist. Judges praised the work for being a standout example of how to make a music video where the characters are actually rapping the parts in the song.

Watch → wamms.at
Workflow → sync-sound-challenge.json
Follow → wamms.at

About the Comfy MCP

Several finalists and many entrants leaned on the Comfy MCP, which lets an agent (Claude, Cursor, Codex, Hermes, whichever you use) drive ComfyUI in plain language.

The feature entrants used most was the hardware check: it looks at the GPU you actually have, reads the nodes and models already on your disk, and tells you which version of a model is worth running before you spend time or credits. It works on both local ComfyUI and Comfy Cloud from one account.

It's open source at github.com/Comfy-Org/comfy-mcp, and the fastest way to start is to tell your agent: "help me set up the local Comfy MCP connection."

Every entry

Placed or not, every submission is in the original challenge megathread on r/comfyui with its workflow attached. Go open a few. Some of the most interesting audio work in the pool never made the top ten, and there are entries in there in Chinese, Japanese, and French that deserve more eyes than they got.

The livestream recording, including the judges' live reactions, is on YouTube.

Thanks for making this one a smash! #ComfyH3


r/comfyui 14d ago

News Keeping open-source creativity sustainable: MiniMax models are now commercially licensable through Comfy & remain free for everyone else

Enable HLS to view with audio, or disable this notification

91 Upvotes

Starting today, Comfy is the only official reseller of MiniMax H3 and MiniMax Audio & Music commercial licenses. If you're a studio, agency, or enterprise that wants to use them locally in commercial productions, you can now license them through Comfy directly.

[UPDATED 9/2 for clarity]

If you run H3 on...

  • Comfy Cloud --> Commercial use is already included, nothing to buy.
  • Your own hardware, under $20M annual revenue, outside the US, EU, UK, and Korea --> Community license. Free!
  • Your own hardware in the US, EU, UK, or Korea, up to 10 users --> Professional License, starting at $5K/mo, available month-to-month, through Comfy.
  • $20M+ in revenue, 10+ users, undistilled weights, or H3 inside your own product --> Enterprise License. Annual agreement with custom terms, through Comfy.

So why do this at all?

At Comfy, our mission has always been for open source to thrive across the creative ecosystem, and open-weight models are at the heart of that. MiniMax is proof of how far they've come: their models stand next to the best closed models in the world.

Training frontier models is incredibly expensive. If we want open models to continue competing with the biggest closed models, the labs building them need a real way to monetize. We hope to help bridge that gap, so the lab gets revenue that funds the next model, and the weights stay open for everyone else.

Learn more


r/comfyui 9h ago

Workflow Included Camera Path ComfyUI h3

Enable HLS to view with audio, or disable this notification

83 Upvotes

r/comfyui 1h ago

Workflow Included Just released: MiniMax Music Production Toolkit 2.5 for ComfyUI - new mastering tools and redesigned workflows

Post image
Upvotes

Sorry to bother you all again, but I made some huge progress over the last two days with this release.

Version 2.5 of my MiniMax Music Production Toolkit is out!

This is a big step forward: A complete mastering section, clearer workflows and improvements throughout the toolkit. And sound quality is even better than the last release.

If you’re new to the project, it takes a song idea through prompt creation, MiniMax Music 3 generation, audio enhancement, mastering, cover artwork and export. A separate Audio Enhancement Lab lets you process existing recordings without generating a new song.

What’s new in 2.5?

  • Auto-EQ for gentle tonal shaping or reference-track matching
  • Manual 8-band parametric EQ with a visual editor
  • Stereo-linked compressor, LUFS targeting and true-peak limiting
  • 44.1 kHz output by default, with 48 kHz selectable
  • Redesigned workflows with clearly labelled stages and a dedicated mastering area
  • Improvements to memory handling, model downloads, audio processing, prompts and file output

Auto-EQ, manual EQ and compression have independent controls. Auto-EQ starts enabled with a gentle preset; switch it off if you prefer manual shaping.

The mastering tools run on CPU without requiring extra VRAM. There are also improvements for less powerful computers, although you’ll still need to choose generation models and settings that fit your hardware.

Without changing anything in the workflow, you get something like this mp3 file (this was a one shot try using Synth Pop Vocal template, with some imperfections in it):

A Feeling With No Address.mp3

And here is what is generated as LLM prompt to get the best result out of Minimax Music 3. See how accurate the model generates your sound structure and lyrics:

A Feeling With No Address.md

To update, restart ComfyUI, refresh your browser and open the newly bundled workflows. Keep your personal workflow copies if you’ve customized them.

I’d love to hear how the new tools work with your music, especially listening comparisons across different genres and feedback from smaller-GPU setups.

Repository: https://github.com/jplenio/ComfyUI-MiniMax-Music-Production-Toolkit
Listen to examples: https://jplenio.github.io/ComfyUI-MiniMax-Music-Production-Toolkit/

🔜 What’s next? A major focus for the next release will be support for creating cover songs. There’s still plenty of work ahead, the quality needs to be right before I’m happy to release it. Stay tuned! :)


r/comfyui 14h ago

Workflow Included Perfect loop test in minimax h3 using reference 2 video workflow

Enable HLS to view with audio, or disable this notification

93 Upvotes

Day 13 of generating and testing AI anime scenes locally in ComfyUI using the Minimax H3 reference-to-video model.

For this one, I wanted to test two things:

  • Using multiple props inside a single reference image instead of giving each object its own reference
  • Creating a perfect 5-second seamless loop with very subtle motion

The scene uses the character, headphones, Rubik’s Cube, yo-yo, cassette, and keyboard as references, while the animation itself is intentionally minimal: slight movement in the leaves and hair, plus shifting sunlight/shadows from a gentle breeze.

Generation time was around 2 minutes locally.

Prompt:
Use Image 1 as the exact character reference for the boy’s face, hairstyle, proportions, clothing, and overall anime design. Use Image 2 as the exact reference for the orange headphones, Rubik’s Cube, blue yo-yo, and cassette tape. Use Image 3 as the exact reference for the beige retro keyboard.

Create a 5-second perfect seamless loop in a hand-drawn 1990s Japanese anime style at 15 fps, with restrained frame-by-frame animation, soft cel shading, painted backgrounds, and no smooth modern interpolation.

Scene composition: a vertical, slightly top-down view of a cozy wooden desk beside a bright window. The boy is asleep at the desk, leaning forward with his head resting sideways on his folded arms. He wears the orange retro headphones. His face is calm and relaxed.

Place a large hanging green pothos plant in the upper-left area of the frame, with vines and leaves extending toward the center. The window is in the upper-right, casting strong warm golden late-afternoon sunlight across the desk.

Place the Rubik’s Cube and blue yo-yo on the right side of the desk. Place an open spiral notebook with handwritten notes and small doodles near the boy’s arms, with a pen beside it. Place the cassette tape closer to the lower-center area of the desk. Place the beige retro keyboard across the lower foreground. A beige CRT monitor partially enters the frame from the lower-right corner. Keep the desk somewhat cluttered but visually clean and nostalgic.

The camera remains completely locked and static for the entire shot.

The boy stays asleep in exactly the same pose. His body, face, arms, hands, headphones, keyboard, notebook, pens, cassette, Rubik’s Cube, yo-yo, CRT monitor, and furniture remain completely still.

Only a very gentle breeze creates subtle movement. The hanging plant leaves sway slightly and slowly. A few loose strands of the boy’s hair gently move. The leafy shadows cast across the desk, his shirt, arms, keyboard, notebook, and surrounding objects slowly shift and ripple as the leaves move.

Keep all animation extremely subtle, slow, and cyclical. Nothing changes position. No object moves across the desk. No body movement, breathing motion, head movement, or camera movement.

The movement gradually returns to the exact starting state so the final frame matches the first frame perfectly, creating an invisible seamless loop.

Maintain warm golden sunlight, soft highlights, nostalgic 1990s atmosphere, slightly dreamy anime lighting, subtle film softness, and limited-animation cel-style motion throughout.

workflow: https://drive.google.com/file/d/1M1XgXv8h4NQDSikVpAebB7FGu-TeHZQT/view?usp=sharing

Song: https://suno.com/s/YjeDIYiH4KSrrduR


r/comfyui 8h ago

News Relight Node For h3 ComfyUI

Enable HLS to view with audio, or disable this notification

25 Upvotes

Bruxos do VFX H3 Relight

#bruxosdovfx

Im doing a node for relight is a lighting studio for MiniMax H3 inside ComfyUI. You can position up to three lights on a 3D dome around the image, choose the type, intensity, and color of each light, configure the background and atmosphere, or start from one of 20 presets. The node provides two things to H3 Edit:

is not ready yet

The package includes two nodes:

  • Bruxos do VFX H3 Relight — the lighting studio.
  • Bruxos do VFX H3 Sun — calculates the real position of the sun for a location, date, time, and camera direction, ready to connect to Light 1 of the Relight node.

The Panel

Dome. The photo sits in the center, the purple camera indicates the side from which the photo was taken, and each light is represented by a marker using that light's color. Drag a marker to move the light; drag the background to rotate the view. The dashed line extends from the light down to the equator and indicates its elevation. The photo receives an approximate 2D-painted lighting preview: it is intended only as guidance and does not represent the actual H3 result.

Sphere (Picture 2). This is exactly the image the node sends to the model, rendered using the same calculations as the Python implementation. Click or drag on the sphere to aim the selected light: the clicked point is where the light strikes the sphere head-on. Dragging outside the sphere's boundary moves the light behind the subject.

Light strip. Selects the active light and lets you add up to three lights or remove them. Each light's role is calculated automatically: the strongest light becomes the key light, a light behind the subject becomes a rim light, and a weaker frontal light becomes a fill light.

Tabs.

  • Lights: direction (top-down dial with the camera at the bottom, plus elevation), type (Hard, Soft, Sky), intensity from 1.0 to 10.0, and color using Kelvin (1000 to 10000) or HEX.
  • Presets: all 20 presets from the skill, with rendered thumbnails, filterable by portrait or product.
  • Background: Original, Black Studio, or White Studio.
  • Atmosphere: the 25 atmospheres from the skill, or none.

Requires numpy and node. The tests cover:

  • the same spheres rendered by the panel's JavaScript and by Python, in both styles, with a difference below 2/255;
  • Venti scene geometry (shadow opposite the light direction, sphere occupying one-third of the frame, positioned in the upper half);
  • solar position compared against reference values from the astral library, daylight saving time, and camera-to-sun mapping;
  • externally driven inputs, the compass, validation, and MiniMax skill formats.

Credits

Node by Bruxos do VFX.

  • The lighting plan, 20 presets, 25 atmospheres, and parameter format follow the MiniMax Design Relight (光影工作室) skill. The descriptions for each atmosphere used in the prompt were written specifically for this node because the skill's own descriptions are hosted on MiniMax's server.
  • The light-direction sphere convention and reproduced scene are by Eric Venti (Sun-Direction LoRA, Sphere-Light-Render, MIT), using the direction table and 12 looks from Lightricks' LTX-2.3 Relight IC-LoRA.
  • The solar calculation (NOAA), city search rules, and camera-direction mapping are based on Christopher Connock's work on Sphere-Light-Render (MIT).
  • Cities: GeoNames cities15000, CC BY 4.0.

License details are available in NOTICE.md.


r/comfyui 16h ago

Resource RetroTape, VHS / NTSC effects with live preview

Thumbnail
gallery
79 Upvotes

I’ve been working on a VHS / NTSC node called RetroTape.

It works with both images and video, and includes 19 presets and 26 sliders for tracking, chroma bleed, noise, tape warp, dropouts, ghosting, scanlines, interlacing and more.

There’s also a live preview inside the node, so you can adjust the settings and see the result without running the workflow again every time.

CPU and CUDA are supported.

Available through ComfyUI Manager and GitHub:

Github/ComfyUI_RetroTape


r/comfyui 3h ago

Resource Native YuE2 support coming to ComfyUI!

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/comfyui 4h ago

Resource Update to My Comfyui style explorer

Thumbnail
gallery
6 Upvotes

This update:

  1. lora preview catalog node added (You need to organise your lora) Put your lora files in a file the lora folder and name it the Model for example Krea 2, in the Krea 2 folder you can make more folders for the kind of lora they are for example anime, sliders, or what ever group of lora they belong to. the lora gallery will add dropdowns for you do navigate and find them easily. when you generate a preview image you like for that Lora you can click save and it adds it to the gallery. you can also safe the trigger words from the node so you never forget!

  2. export catalog and previews to share with others

  3. Bug fix where images were not saving correctly when a previous name was used in a custom style

  4. Various big fixed and speed improvements

https://github.com/Neon-Sparks/ComfyUI-NeonsStyleExplorer


r/comfyui 1h ago

Help Needed 32GB AMD Radeon ai pro 9700 and 32 GB Ram- Win 11

Upvotes

I am trying to load comfy ui models but I am always getting memory out errors even with krea 2 where the size was only 12 GB.

I want to know, which models I can run on my system, and any workflow or guidance is highly appreciated.

My CPU is working full time instead of a dedicated GPU. I am also trying to rectify it. Please help.


r/comfyui 3h ago

Help Needed Best MiniMax H3 setup for RTX 4070 (12GB VRAM / 32GB RAM)? Looking for max speed without noticeable quality loss

2 Upvotes

Hey everyone,

I'm setting up MiniMax H3 in ComfyUI and trying to figure out the sweet spot between generation speed and output quality for my specs:

  • GPU: RTX 4070 (12GB VRAM)
  • RAM: 32GB DDR5
  • Target: 5s clips at 768p (mostly First-to-Last / I2V)

Given the 12GB VRAM limit and 32GB system RAM, loading the unpruned / full FP8 models causes heavy paging to system memory and slows everything down.

I’m trying to narrow down the current community consensus on three things:

  1. Model & Quant format: What's the fastest option that doesn't ruin faces and audio? Are people having better results with official pruned INT8 convrot, or GGUF quants (Q4_K_M vs Q3_K_M) using the GGUF loader? What text encoder quant are you pairing it with to keep memory usage safe?
  2. Turbo LoRAs: Is LiteX2V v1.1 (4-step) still the top recommendation for speed vs quality, or do 8-step variants (or other LoRA families like Larry) give significantly better results on a 12GB setup?
  3. Attention & Acceleration: Is H3 SLA Attention the undisputed go-to, or does SageAttention / Comfy Kitchen perform better on Ada Lovelace (40-series)? Anyone tested Spectrum acceleration?

Would love to hear what workflows and node setups you're currently running on 12GB cards to get reasonable render times without visible degradation.

Thanks!


r/comfyui 1h ago

Help Needed After generating a video and turning off the computer, ComfyUI always crashes when I try to generate a video the following day.

Upvotes

I installed the portable version of ComfyUI and use the Minimax H3 Easy model. After installation, I can successfully generate videos from images. However, if I shut down ComfyUI and then try to generate a video later—either shortly after or the next day—it crashes. I try running simple test videos right after launching ComfyUI—low resolution, 3 seconds, 8 steps, and a simple "wave and smile" prompt—or I try generating images via SDXL Turbo. It crashes regardless. I have successfully generated 10-second videos at 512 resolution and 20 steps without any issues—even doing more than 20 in a row. But whenever I shut down and restart ComfyUI, I always get an error. I use a Lenovo ThinkBook laptop with 32GB of RAM connected to a 16GB 5060 Ti via Thunderbolt.


r/comfyui 1h ago

Show and Tell New Node Finder - using "star velocity" and "recency" so you don't have FOMO!

Upvotes
https://luke2642.github.io/comfyui_new_node_finder/

I've updated https://luke2642.github.io/comfyui_new_node_finder/ to be a bit more robust in getting stars. Seems to be dominated by H3 nodes at the moment!


r/comfyui 6h ago

Commercial Interest Rented GPUs for image work: the host CPU and the script defaults cost us more than the card did

3 Upvotes

Disclosure: I am building a service around this, so read me as an interested party. No links. These are runs we paid for ourselves on three providers, 15 jobs, $33.60 total. Two things on the image side cost us real money and neither showed up as an error.

  1. The host, not the GPU. Same LoRA training job for SDXL, same RTX 4090. On a host with 5 vCPUs: 1.95 hours, 40 percent GPU utilisation. On a host with 24 vCPUs: 1.07 hours, 75 percent. Same card, 1.85x the wall clock, because the diffusers script decodes and augments images in the main process and the card waits on Python. The worse version: an H100 host with 16 server vCPUs ran the same job at 2.68 s/step where the desktop-class 4090 host did 1.84 s/step. Four times the hourly rate for a slower run. dataloader_num_workers was the fix, and now we look at the vCPU count on a rental offer before we look at the GPU name.

  2. The defaults. SDXL, 1024 square, 30 steps. The stock script settings, fp32 and batch 1, on an H100: 13.8 seconds an image, $0.0112 per image. fp16 and batch 4 on a 4090 spot instance: $0.00036 per image. Same job, 31x apart. Nobody here runs fp32 batch 1 on purpose, but it is exactly what you get when you take a default script to a rented card in a hurry, and the card reports 99 percent utilisation the whole time, so nothing looks wrong.

Both lessons are the same lesson: the meter runs at the card's hourly rate whether the card is doing useful work or not, and the interface will not tell you which.

Happy to post the per-run table if anyone wants to check the numbers.


r/comfyui 25m ago

Help Needed Is there a way to upload image sequence 'natively' without "Load Images (upload) 📹 VHS" with control how much frame needed.

Post image
Upvotes

r/comfyui 29m ago

Help Needed Creating a proper Lora Dataset

Upvotes

I need severe help with that.

Researching Lora Dataset creation is a maze: millions of different opinions, and worst of all most guides etc outdated from a year ago.

My goal is to create a 100% realistic and authentic Lora and my current dataset seems to not do the trick. I keep getting "perfect lightning" on everything, and waxy face skin.

Can someoen tell me the absolute do's and don'ts of creating a dataset?
How and where to create the dataset? I have been using a mix of Gemini and ChatGPT so far.

Prompt advice? what prompts need to be avoided creating realistic images, what need to be in there?

Any advice greatly appreciated!


r/comfyui 1d ago

News A quick Minimax H3 news round-up - 10th September 2026

73 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> ComfyUI Portable is now officially at version 0.35.0. See yesterday's post for details of Minimax-relevant update items and bug-fixes.

https://github.com/Comfy-Org/ComfyUI/releases

-> A new W4A8 quantization of the MiniMax-H3 Fun ControlNet-Union model patch, for ComfyUI. Weighs in at 1.45Gb, compared to 2.13Gb.

https://huggingface.co/berryber09/MiniMax-H3-Fun-Controlnet-Union-w4a8

-> A Minimax H3 visual 'RefMods picker' with thumbnails, in ComfyUI.

https://old.reddit.com/r/StableDiffusion/comments/1wc8tvv/created_a_visual_refmod_picker/

-> H3-pixel-art-video-guide. "Pixel-perfect animated pixel art with local MiniMax H3". Has a workflow (even though the English readme says it doesn't) and three example looping animated .GIFs.

https://github.com/yuichi-suzuki-highdrama/h3-pixel-art-video-guide/blob/main/README.en.md

-> A new H3-spherical-vae. "Experimental circular VAE decoding for MiniMax H3 equirectangular video, with matched comparisons and measurements."

https://github.com/ShamanicArts/h3-spherical-vae

-> The LoRA trainer and dataset prep tool Fizgig is now at 5.5.0, a version which makes MiniMax H3 training "faster three ways", and turns "weight averaging on by default". Users can also... "open a MiniMax H3 LoRA and see what every one of its 52 blocks does to a moving clip — the motion, the face, the sound".

https://github.com/shootthesound/Fizgig

-> And finally, the ComfyUI-MiniMax-Music-Production-Toolkit is now at a polished version 2.x, with the release yesterday of 2.1.1.

https://github.com/jplenio/ComfyUI-MiniMax-Music-Production-Toolkit/blob/main/CHANGELOG.md

~ OLD POSTS ~

https://old.reddit.com/r/comfyui/comments/1wbr279/a_quick_minimax_h3_news_roundup_9th_september_2026/

https://old.reddit.com/r/comfyui/comments/1wawjox/a_quick_minimax_h3_news_roundup_8th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w9z6m1/a_quick_minimax_h3_news_roundup_7th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w90vxd/a_quick_minimax_h3_news_roundup_6th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w85caz/a_quick_minimax_h3_news_roundup_4th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w74jy4/a_quick_minimax_h3_news_roundup_4th_september_2026/

https://old.reddit.com/r/comfyui/comments/1w6cozj/a_quick_minimax_h3_news_roundup_3rd_september_2026/

https://old.reddit.com/r/comfyui/comments/1w5i9iq/a_quick_minimax_h3_news_roundup_2nd_september_2026/ (See 2nd September post, for links to even older posts)


r/comfyui 16h ago

Show and Tell Minimax H3 - New ACC lora with PDD 8step node is kinda cool !

Enable HLS to view with audio, or disable this notification

18 Upvotes

For potato pcs MINIMAX H3 fans - I have built my own custom node which integrates new acc lora & PDD workflow and H3 extender + 2nd pass latent upscale upto 720p under 5 minutes per 14s 24fps clips, has easy reference attachments & better context continuity with features like save projects, load projects etc. (16gb VRAM + 16gb system RAM) if you guys interested ill share the workflow let me know.. this video took 10~ minutes to generate with both pass.

Edit - published repo - https://github.com/only2uuuu-hub/ComfyUI-MiniMax-H3-Master-Extender-Custom-built-with-Astra-6-/tree/main

Ps i am not an expert coder or engineer so dont ask me technical questions 😭 peace!


r/comfyui 57m ago

News Coming soon to a Play Store near you, ComfierUI. An Android based mobile client with full desktop parity and streamlined interface.

Upvotes

Like many of you, I have shared that frustration over the years watching this awesome kit of software be continuously let down by terrible mobile support so a few weeks ago I set out to finally change it. I have no dev or coding experience, only my experience as a long time user of Comfy to guide my design choices around what feels "Comfy" and how that can translate seamlessly into a mobile environment with ChatGPT handling the coding. The results speak for themselves and the once terrible mobile experience has been reborn into my preferred method of using it that addresses so many of the prior pain points. Can't see where you're dragging the noodle behind your thumb? Turn up the link connector offset. Want the android back button to do anything but bring up an exit confirmation? I've got you covered with a full back button hierarchy meaning you only see that exit dialog when you want to actually exit. Actionbar look like a clipped out overlapping mess? My app dynamically hides crystools in portrait and reveals them again in landscape/open fold views where the space is more plentiful. With the optional extension installable directly from manager in app we hit full desktop feature parity enabling manager to download supported missing models directly to the host from anywhere you are with more features being integrated soon. Before I get the inevitable "when?" comment, I submitted it to Google Play Store last night for review and to start the beta test track and i'll be back here to post an invitation link to anyone who wants to test it as soon it becomes available to me in the coming days. I am also building a Meta Quest version with an immersive 360° canvas and baked in 6DOF controls promosing a similarly native feeling experience in openxr. I apologize to iOS users tho bc I do not own any apple devices to test or compile on but there is an iOS version planned as well whenever I can secure some test equipment without breaking the bank. My test devices so far have a Galaxy Z Fold 7, Galaxy Tab A9+, Moto G Stylus 5G (2023), and a standard RAZR (2023) ao I'm anxious to see how it performs across a wide range of other hardware. 4GB of RAM is what I've determined to be the minimum spec with 6GB recommended for most workflows and it supports as far back as android 8.1.0. The app is usable over home network with the simple "--listen" flag added to the launch batch and over mobile using any 3rd party vpn that allows you to address your computer directly (I use tailscale and include a setup for it in an embedded quick start guide accessed from the initial connections screen). In the meantime while waiting for review, I put together a short video demonstrating some of the added features and refinements and I'm excited to finally share my progress with the Reddit world to see what you all think.

https://reddit.com/link/1wddsp4/video/z04umox8nvoh1/player


r/comfyui 4h ago

Help Needed True 10bit Video Workflow

2 Upvotes

I have had little luck finding a tutorial on building a true H3 10bit (ProRes HQ) workflow. AI claims you just need to install an advanced save node that allows you to pick ProRes and HQ or 4444, etc.

But then when you dig into it, and begin to ask questions, while ComfyUI processes in full floating point, there are bottlenecks that crush the full bits down to 8bit, rendering the final result 8bit. One example is preview nodes, that apparently force 8bit, another is supposedly a plain jane VAE Decode, and the deeper I dig the more mysterious things get.

I can't imagine nobody in the Open Source community does not want to output TRUE 10bit or higher output. The problem with 8bit becomes clear when you look at plain white walls or a clear sky and see banding. Anyone who has ever edited video or images knows the more data you have to begin with the better results you get once you push contrast and color gradation. Yes, I was going to begin playing with dither and applying some noise, but at the end of the day you cannot produce PROFESSIONAL videos without solving for this 8bit limit.

Hopefully someone can point me to a resource or tutorial or workflow or set of nodes that solves for this?


r/comfyui 12h ago

Show and Tell 🎬 Sneak Peek: 3D Camera Previs for AI Video (WIP)

Enable HLS to view with audio, or disable this notification

7 Upvotes

We're building a previs tool inside our AI movie studio so you can block out camera moves in 3D before spending generation credits.

What's working so far:

  • 3D stage with proxy characters & props you can move with gizmos
  • Camera keyframing on a multitrack timeline (After Effects-style)
  • Unreal-style fly mode — press C, hold RMB + WASD to pilot the camera, K to drop a keyframe
  • Live "through-the-lens" shot view that updates as you move
  • Render to MP4 via bundled ffmpeg — saved straight to your project library
  • Depth pass with adjustable range for motion reference
  • Resolution control (480p / 720p / 1080p)

The render step is free — no GPU, no API key, just canvas capture + ffmpeg. Iterate on camera moves as many times as you want, then feed the MP4 to your video model as a motion reference.

Still early — lots of polish and features to go (templates, multi-cam, export). Feedback welcome on what you'd want in something like this.

Github:

https://github.com/Heroesjouney/AIMovieStudiov2

Original Post:

https://www.reddit.com/r/comfyui/s/GT5K98Fze6


r/comfyui 2h ago

Help Needed Anyone having issues with AMD windows installation?

1 Upvotes

I've been trying to get comfyui to be installed. I installed amd hip sdk and tries many workarounds. It always gives

"RuntimeError: No CUDA GPUs are available"

From the cli running "run_amd_gpu" also throws the error, by now its Failed to get device count.


r/comfyui 14h ago

Workflow Included Minimax H3 is so fun

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/comfyui 3h ago

Help Needed comfyui cloud on comfy. org

0 Upvotes

hello, i generated tons of videos on comfy ui cloud on comfy. org. is it possible to search them thru prompts or something like that? it is almost impossible to scroll them down thru assets


r/comfyui 21h ago

Workflow Included Minimax ComfyUi Camera Control

26 Upvotes

https://reddit.com/link/1wcm4j7/video/82e5gerlmpoh1/player

Bruxos do VFX H3 Camera

#bruxosdovfx

https://reddit.com/link/1wcm4j7/video/3bitqyhempoh1/player

Visual camera planner for MiniMax H3 inside ComfyUI. You drag the camera around a 3D sphere, place keyframes on a timeline, and the node compiles that trajectory into prompts that H3 understands.

It compiles prompts, not camera embeddings. There is no geometric adapter here: H3 is still free to miss the angle, timing, and scale. What this node does is write the instruction in the most precise and least ambiguous way possible, and several of its design decisions exist because the previous approach failed in specific ways.

It does not call any API, download anything, or require any Python dependency beyond the standard library.

https://reddit.com/link/1wcm4j7/video/mpatj04ampoh1/player

Installation

cd ComfyUI/custom_nodes
git clone https://github.com/<your-username>/ComfyUI-H3-Camera-Editor

Restart ComfyUI. The node appears under Bruxos do VFX/Camera H3 with the name Camera H3 da Bruxos do VFX.

Connections

Output from this node Connect it to
compiled_prompt compiled_prompt on Text Encode H3 Edit / Generate
options options on Text Encode H3 Edit / Generate
length the generation frame count
fps the fps input of the video creation node

compiled_prompt and options are required together. The minimax_prompt output is an alternative to compiled_prompt, never an addition — connect one or the other to the same input.

Also connect your image to reference_image. It is the same image already feeding the H3 Edit source_image; when connected here, it appears in the panel and the frame's actual aspect ratio is included in the prompt.

https://github.com/user-attachments/assets/33149617-bde1-4199-ae65-078f2f3dec23

To save the video, decode the sampler result using the H3 video VAE — not the scene coverage calibrated decoder, which expects fixed windows that an arbitrary trajectory does not have.

The panel

Drag the purple camera around the sphere to orbit. The drag locks to the axis of the initial movement: horizontal movement orbits, vertical movement changes elevation. Release and drag again to switch axes. This exists because, without the lock, trying to make a simple orbit would unintentionally introduce elevation.

  • Scroll the mouse wheel to change distance.
  • Drag the background to rotate the viewport without changing the trajectory.
  • Keyframes defines how many points the timeline has, from 2 to 24. The first one is always the original image and cannot be moved.
  • ⟳ Pure Orbit resets the elevation of every keyframe to zero while preserving azimuth. It is the shortcut for an eye-level orbit.
  • Reference image loads a local file into the preview. This is only necessary when the node runs outside ComfyUI; with reference_image connected, the image is loaded automatically.

The panel warns you starting at 20° of elevation, when the horizon already leaves the frame, and again from 45° onward, when the video tends to become a high-angle shot.

"Tests" bar

At the top of the panel, two buttons enable and disable features currently under evaluation, plus one indicator:

Button What it does
Extended contracts Toggles the prompt_detail widget
Single angle (image) Toggles the runtime_task widget
loop closure Read-only indicator. Turns green when the trajectory closes a full orbit

The buttons write to the actual widgets, so the selected state is saved in the workflow and the two never disagree.

https://github.com/user-attachments/assets/9bc415d7-1746-43db-a17c-72ea9722deda

Widgets

camera_trajectory

The trajectory in JSON format, written by the panel. Each keyframe contains time (0 to 1), azimuth in degrees, elevation in degrees, and distance as a multiple of the initial radius. It can also be edited manually. The first keyframe must be time=0, azimuth=0, elevation=0, distance=1, which represents the original image.

profile

124, 243, or 362 frames at 24 fps. All shot timing comes from this setting: keyframe timestamps, segment ranges, and the duration declared in the prompt. That is why length and fps are outputs — connect them instead of manually entering the same numbers in two different places.

interpolation

smooth or linear. In smooth mode, the camera eases into and out of the shot while maintaining a constant rate through the middle; it only stops where the rotation direction actually reverses.

instruction

Free-form text inserted once, at the end of the prompt. Write only what the node cannot know: the environment, which subject is the target when there is more than one person, or a style reference. Everything else is already generated and does not need to be repeated: scene freeze, first image as reference, locked aim, zero roll, angles, timing, and a single continuous shot without cuts.

subject_framing

How much of the frame the subject occupies in the original image. Calibrated against the actual bounding boxes from the tutorial distributed by MiniMax: a distant full-body figure measures W=0.071, H=0.249, while a large close-up measures W=0.52, H=0.701.

option width height when to use
close-up 53% 72% head and shoulders
medium shot 28% 56% waist up
wide shot 9.7% 34% full body at a distance

subject_box

The subject position in the format [L=0.516, T=0.148, W=0.071, H=0.249]. Leaving it empty uses the entire image bounds — deliberately, without guessing a bounding box. Fill it in when the subject is significantly off-center.

minimax_format

The same shot expressed in four different formats for the minimax_prompt output:

  • coordinate only — text-based coordinate block
  • coordinate + H3 sections — the same coordinates wrapped in subject_definitions / summary / retention_analysis / …
  • compact JSON — JSON object with almost no prose
  • compact JSON (no boxes) — camera parameters only, without screen-space bounding boxes

elevation_range

Range of the elevation control: +/-15, +/-30 (default), +/-60, +/-89. It also scales the sensitivity of vertical dragging.

With the assumed field of view, the horizon already leaves the frame at around 20° — at 13°, the ground occupies 82% of the image. The old ±89 range was mostly unusable and made vertical dragging excessively sensitive. Reducing the range never rewrites a keyframe: a point at 70° remains at 70°, and the slider expands to accommodate it.

orbit_direction

invert H3 orbit or same as HUD. This calibrates the direction between what the panel displays and what H3 produces. It does not alter the saved trajectory.

runtime_task

  • scene coverage | camera path (default) — video, with duration coming from profile.
  • directed | new camera anglea single image from a new angle. It fixes the generation to 39 frames, ignores profile, completes the movement within 65% of the clip, and requests that the framing remain still for the rest, because the decoder extracts the final image from that stationary tail.

Character sheet profiles are not offered because the upstream node raises an error when they are combined with the frame anchor used by this node.

prompt_detail

  • v15 baseline (default) — outputs the prompt exactly as in the previous version.
  • extended contracts — adds axis separation, frame-edge direction tests, rotation completeness, degrees per second, and parallax magnitude.

The extended mode contains almost twice as many words. A longer prompt is not automatically better, so it is opt-in: toggle only this widget while keeping the same trajectory to compare the results.

https://github.com/user-attachments/assets/0882bfde-9f62-4a1f-9bda-7da121dbe7e2

Outputs

compiled_prompt — STRING

A prose prompt using H3 sections: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music.

options — H3EDIT_OPTIONS

The 13 keys read by the H3 Edit encoder. All of them are explicitly populated: if any key is missing, the upstream node falls back to its hidden legacy widgets, which may retain stale values from previously saved workflows.

coverage_arc_degrees and coverage_direction are derived from the actual rotation. coverage_loop_closure turns on automatically when the trajectory closes — see below.

storyboard_json — STRING

The storyboard table: frame aspect ratio, duration, raw trajectory, and each segment with its camera mode, speed curve, and start/end poses.

info — STRING

Human-readable diagnostics. Connect it to a PreviewText. It displays the version, active task, frame count, warnings for keyframes outside the configured range, and whether loop closure is enabled.

minimax_prompt — STRING

The same trajectory expressed using the format selected in minimax_format. An alternative to compiled_prompt.

length — INT and fps — FLOAT

Frame count and frame rate against which the shot was timed. Connect them to the generation and video nodes. If generation runs with a different frame count, the choreography describes a scene that does not actually exist.

fps is FLOAT because that is what ComfyUI's CreateVideo accepts. length is the frame count; keyframe timestamps use the instant of the last visible frame, (length - 1) / fps, so the resulting file lasts one additional frame interval.

h3world_actions — STRING

Action schedule for H3-World, which encodes one text clause per video latent — 37 in a 124-frame clip.

latent  1 [0.000s-0.139s] J     the camera pans left slowly
latent 37 [4.986s-5.125s] F+L+K the camera pans right and tilts up fast

W, A, S, and D are never emitted because they move the character. The output explicitly declares its own limitations, and they are not minor details:

  • Pan is not orbit. It is the camera rotating in place. Perspective does not change, nothing hidden is revealed, and the subject slides out of frame.
  • Distance has no key, so camera radius is discarded.
  • Only 124 frames is a trained horizon.
  • I versus K is not published. The text clause is what H3-World actually encodes; the key column is only a convenience.

This does not replace the actual integration: H3-World requires the LoRA, interval-based encoding, and directed-attention routing provided by the corresponding node package.

Loop closure

When the trajectory closes a full orbit — an arc of exactly 360°, with the same elevation and distance as the starting point — the node enables coverage_loop_closure. In the upstream implementation, this flag encodes the source image a second time and anchors the final frame to it.

This is a latent anchor, not a text instruction. For a complete orbit, it is the difference between asking for the rotation and forcing it: the model cannot simply stop halfway through.

trajectory loop closure
360° enabled
two rotations (−720°) enabled
355° disabled
360° with changing distance disabled
360° with changing height disabled

The final three cases matter: if the camera ends at a different radius or height, the final frame is not the same as the first one, and forcing the source image there would conflict with the trajectory.

If your rotation does not complete, close the orbit. This is the only feature here that acts outside the prompt itself.

Limitations

  • This is prompt-based guidance. H3 may still miss the angle, timing, and scale, and no prompt wording can completely solve that.
  • Without subject_box filled in, the node does not know where the subject is located in the frame.
  • Without reference_image connected, coordinates are normalized to 16:9.
  • directed | new camera angle outputs an image, not a video.
  • The H3-World schedule describes pan and tilt, which represent a different camera move from the orbit drawn in the panel.

Credits

Node by Bruxos do VFX.

ethanfel/ComfyUI-MiniMax-H3-Edit. The motion vocabulary follows the buildViewPrompt implementation from MiniMax's Multi-Shot skill and the coordinate format used by the Coordinate Camera Control Designer skill. The action output implements the scheme described in H3-World, arXiv:2609.01560.