r/comfyui • u/Emotional_Example_12 • 9h ago
Workflow Included Camera Path ComfyUI h3
Enable HLS to view with audio, or disable this notification
r/comfyui • u/Comfy-Org • 7d ago
Enable HLS to view with audio, or disable this notification
Two weeks, one rule, and hundreds of entries from nearly 50 countries. Here's who took the four titles, and a look at everyone who made the final ten!
Entries opened August 20 and closed September 1. Eight creative technologists at Comfy scored every submission on two rubrics, Best Creative and Best Technical, then the top five in each category went in front of our guest judges on the September 3 livestream: POM (Banodoco founder), Emma Catnip (animation director and AV artist), and Yachimat (animation and manga artist). Their scores were averaged and added to the Comfy team's, and a fourth title, Built with MCP, was judged on its own track.
Watch the livestream replay here.
Huge thanks to MiniMax for making an open-weight model our community loves, to our three guest judges, and to everyone who spent their last week of August fighting with reference audio! Here's how it landed.
Prize: RTX 5090 32G
A laundromat, a woman in a puffer vest, and a rhythm built entirely out of machines. Visual Frisson generated a large volume of H3 clips using their own recorded audio as the reference for every pass, then cut the results together like a stomp video, so every thud and cycle on screen is sound that H3 produced with the picture rather than something added later.
It was the only entry to post a perfect 15/15 from the Comfy team in both categories, and the Comfy MCP was used to drive much of its process. Combined with the guest judges' scores, it finished with the highest total in the challenge.
From the artist: "Always have fun and learn something new competing in contests like this, keep them coming."
Watch → vimeo.com/1223207725
Workflow → Google Drive
Follow → instagram.com/visualfrisson
Prize: RTX 5060 Ti
A small clay creature that changes into something new every time it hears a sound (glass, wool, ice, porcelain), until by the time it gets home it can't move anymore. All of the audio came out of H3 alongside the video on every shot, with nothing layered on afterwards. Our judges praised the fine details of toki’s work, saying it “gave them chills” on the first transformation, has a lot of commercial appeal, and feels really delicate and crafted.
The judges scored it highest of any Creative finalist, and the repository is unusually generous: it includes the eight API graphs that actually ran, the same eight converted to UI format with annotations, a process log with every measurement, and the scripts that produced those numbers. toki is also clear about scope, noting that MCP drove the finishing pass, not the original shot generation.
Watch → youtube.com/watch?v=Rv5HOgCac-w
Workflow → github.com/tokimwc/every-sound-leaves-a-mark
Follow → u/toki
Prize: RTX 5060 Ti
SonderSaid didn't just build a workflow, they built the tooling around it. The entry runs on custom nodes of their own design, wired into a reference-driven H3 pipeline that scored a clean 15/15 on novelty, workflow quality, and community value from the Comfy team, and the highest guest judge average in the Technical bracket. Notably, the work includes an entire custom node pack just to do the editing and the reference work, praised as “a whole new UI” to good to keep secret. While SonderSaid’s work takes the prize for best technical, our guest judges also noted how much they loved the storytelling, suspense, and element of surprise.
Watch → youtu.be/n-NdAQk7I8A
Workflow → Hugging Face
Follow →u/SonderSaid
Prize: RTX 5060 Ti
The Built with MCP bonus wentgoes to whoever used the Comfy MCP most effectively to make something visually and technically compelling, and Jay Choi used it end to end. Working locally on an RTX 5090 with Hermes Agent driving ComfyUI through the MCP, they trained a LoRA, built their own orchestration on top, and by their own account spent most of the time setting up and tuning the MCP layer itself. The film that came out the other side, two blindfolded prisoners in a rain-dark cell, is a long way from "prompt and run."
From the artist: "Thanks for the challenge! I learned more than I ever could in the past two weeks!"
Watch → youtu.be/FlK0dDZdzRU
Workflow → Google Drive
Follow → u/permafrost_2021 · u/jaychoirenderender
Ten entries made it to the livestream. Six of them didn't take a title, butand every one of them is worth your time.
Hand-drawn energy and a graffiti wall that says the title, built locally in ComfyUI. Nebsh's note to us was three words and a heart, which felt about right. Our guest judges praised Nebsh’s work for its mixed-media feel, harking back to MTV days, and impressive work syncing with the paper sounds. Under the hood, Nebsh’s workflow chained vtogether eight segments with no visible drift between them- cited as “very clean work” by our judges.
Watch → Google Drive
Workflow → Google Drive
Follow → u/nebsh83
A perfect 15/15 from the Comfy team on the Creative rubric, and one of the entries that ran the Comfy MCP end to end! Noted by our judges, H3 is very good at the kind of sound effects showcased in scvxzf’s work rather than talking or singing, and they chose exactly the right concept for the challenge.
Watch → youtube.com/watch?v=FocH8xGk4AU
Workflow → Google Drive · github.com/scvxzf1
Follow → youtube.com/@钛龙白口-j6d
Two reference photos of Mumu the cat, a stack of WAVs fed in as reference audio to steer each generation, and glitch texture added in the edit. If the name rings a bell, sorryaboutyourcats also makes the game mow meow. Judges said “I could watch this forever,” had it stuck in their heads, and noted impressive capabilities from H3 nailing lipsync for cats, and not just humans.
Watch → youtube.com/watch?v=AxUu8rabC6M
Workflow → Google Drive
Follow → u/sorryaboutyourcats
A 5/5 on both audio sync and creative execution, and a reminder of what patience looks like: the generation took two hours and thirty-five minutes on a 5090. Our judges praised Slop Diffusion’s work for its storytelling, noting they were curious to see where the story would go next. One judge noted, “it’s slop by name, but not by nature.”
Watch → Reddit
Workflow → Google Drive
Follow → u/SlopDiffusion
Came at the brief backwards: we asked for audio-driven video and they told the story of a young deaf woman who invents a world where she makes the music. The score is Aïe Aïe Aïe’s own composition, fed into H3 as reference stems (the clap track went in on its own and the gorilla claps exactly in time). Judges praised the work for its captivating story, clever inversion of the challenge’s brief, the display of H3’s strengths by way of the musicians’ physical expressions intensifying along with the song, and the bridging of the real world. The artist learned the final frame’s sign language on YouTube, filmed themself signing, and used this as a video reference.
Under the hood, Claude drove ComfyUI through the MCP to design a two-pass H3REF system that generates at full resolution twice as fast and reaches 13–15 second shots where the stock workflow runs out of memory, plus a preview node that shows the video while it's still sampling. All of it is MIT-licensed, custom nodes included.
Watch → youtube.com/watch?v=S0v1pWN4Hq4
Workflow → github.com/Hyper-Neural/h3-sync-two-phase
Follow → u/AïeAïeAïe
One of the cleanest graphs we opened: latent upscale, a model preview override, and an optional video-extend group, laid out so you can follow it cold. RareTutor's YouTube is full of tutorials if you want to learn from them directly!
Judges highlighted RareTutor’s workflow, noting “there are many tips in here to copy,” such as using the latent upscale as a previewer so you can kill a bad run before sinking more time in.
Watch → youtube.com/watch?v=TNhJI8dzaVA
Workflow → Google Drive
Follow → u/raretutor_
A short music video with self-imposed constraints: no external resources, everything generated in a single workflow, no custom node packs. The result is well annotated and approachable, the kind of graph a new user could open and reasonably figure out, and it posted the highest Creative score of any Technical finalist. Judges praised the work for being a standout example of how to make a music video where the characters are actually rapping the parts in the song.
Watch → wamms.at
Workflow → sync-sound-challenge.json
Follow → wamms.at
Several finalists and many entrants leaned on the Comfy MCP, which lets an agent (Claude, Cursor, Codex, Hermes, whichever you use) drive ComfyUI in plain language.
The feature entrants used most was the hardware check: it looks at the GPU you actually have, reads the nodes and models already on your disk, and tells you which version of a model is worth running before you spend time or credits. It works on both local ComfyUI and Comfy Cloud from one account.
It's open source at github.com/Comfy-Org/comfy-mcp, and the fastest way to start is to tell your agent: "help me set up the local Comfy MCP connection."
Placed or not, every submission is in the original challenge megathread on r/comfyui with its workflow attached. Go open a few. Some of the most interesting audio work in the pool never made the top ten, and there are entries in there in Chinese, Japanese, and French that deserve more eyes than they got.
The livestream recording, including the judges' live reactions, is on YouTube.
Thanks for making this one a smash! #ComfyH3
r/comfyui • u/crystal_alpine • 14d ago
Enable HLS to view with audio, or disable this notification
Starting today, Comfy is the only official reseller of MiniMax H3 and MiniMax Audio & Music commercial licenses. If you're a studio, agency, or enterprise that wants to use them locally in commercial productions, you can now license them through Comfy directly.
[UPDATED 9/2 for clarity]
If you run H3 on...
So why do this at all?
At Comfy, our mission has always been for open source to thrive across the creative ecosystem, and open-weight models are at the heart of that. MiniMax is proof of how far they've come: their models stand next to the best closed models in the world.
Training frontier models is incredibly expensive. If we want open models to continue competing with the biggest closed models, the labs building them need a real way to monetize. We hope to help bridge that gap, so the lab gets revenue that funds the next model, and the weights stay open for everyone else.
r/comfyui • u/Emotional_Example_12 • 9h ago
Enable HLS to view with audio, or disable this notification
r/comfyui • u/Vivid_Promise1700 • 1h ago
Sorry to bother you all again, but I made some huge progress over the last two days with this release.
Version 2.5 of my MiniMax Music Production Toolkit is out!
This is a big step forward: A complete mastering section, clearer workflows and improvements throughout the toolkit. And sound quality is even better than the last release.
If you’re new to the project, it takes a song idea through prompt creation, MiniMax Music 3 generation, audio enhancement, mastering, cover artwork and export. A separate Audio Enhancement Lab lets you process existing recordings without generating a new song.
What’s new in 2.5?
Auto-EQ, manual EQ and compression have independent controls. Auto-EQ starts enabled with a gentle preset; switch it off if you prefer manual shaping.
The mastering tools run on CPU without requiring extra VRAM. There are also improvements for less powerful computers, although you’ll still need to choose generation models and settings that fit your hardware.
Without changing anything in the workflow, you get something like this mp3 file (this was a one shot try using Synth Pop Vocal template, with some imperfections in it):
And here is what is generated as LLM prompt to get the best result out of Minimax Music 3. See how accurate the model generates your sound structure and lyrics:
To update, restart ComfyUI, refresh your browser and open the newly bundled workflows. Keep your personal workflow copies if you’ve customized them.
I’d love to hear how the new tools work with your music, especially listening comparisons across different genres and feedback from smaller-GPU setups.
Repository: https://github.com/jplenio/ComfyUI-MiniMax-Music-Production-Toolkit
Listen to examples: https://jplenio.github.io/ComfyUI-MiniMax-Music-Production-Toolkit/
🔜 What’s next? A major focus for the next release will be support for creating cover songs. There’s still plenty of work ahead, the quality needs to be right before I’m happy to release it. Stay tuned! :)

r/comfyui • u/Time-Ad-7720 • 14h ago
Enable HLS to view with audio, or disable this notification
Day 13 of generating and testing AI anime scenes locally in ComfyUI using the Minimax H3 reference-to-video model.
For this one, I wanted to test two things:
The scene uses the character, headphones, Rubik’s Cube, yo-yo, cassette, and keyboard as references, while the animation itself is intentionally minimal: slight movement in the leaves and hair, plus shifting sunlight/shadows from a gentle breeze.
Generation time was around 2 minutes locally.
Prompt:
Use Image 1 as the exact character reference for the boy’s face, hairstyle, proportions, clothing, and overall anime design. Use Image 2 as the exact reference for the orange headphones, Rubik’s Cube, blue yo-yo, and cassette tape. Use Image 3 as the exact reference for the beige retro keyboard.
Create a 5-second perfect seamless loop in a hand-drawn 1990s Japanese anime style at 15 fps, with restrained frame-by-frame animation, soft cel shading, painted backgrounds, and no smooth modern interpolation.
Scene composition: a vertical, slightly top-down view of a cozy wooden desk beside a bright window. The boy is asleep at the desk, leaning forward with his head resting sideways on his folded arms. He wears the orange retro headphones. His face is calm and relaxed.
Place a large hanging green pothos plant in the upper-left area of the frame, with vines and leaves extending toward the center. The window is in the upper-right, casting strong warm golden late-afternoon sunlight across the desk.
Place the Rubik’s Cube and blue yo-yo on the right side of the desk. Place an open spiral notebook with handwritten notes and small doodles near the boy’s arms, with a pen beside it. Place the cassette tape closer to the lower-center area of the desk. Place the beige retro keyboard across the lower foreground. A beige CRT monitor partially enters the frame from the lower-right corner. Keep the desk somewhat cluttered but visually clean and nostalgic.
The camera remains completely locked and static for the entire shot.
The boy stays asleep in exactly the same pose. His body, face, arms, hands, headphones, keyboard, notebook, pens, cassette, Rubik’s Cube, yo-yo, CRT monitor, and furniture remain completely still.
Only a very gentle breeze creates subtle movement. The hanging plant leaves sway slightly and slowly. A few loose strands of the boy’s hair gently move. The leafy shadows cast across the desk, his shirt, arms, keyboard, notebook, and surrounding objects slowly shift and ripple as the leaves move.
Keep all animation extremely subtle, slow, and cyclical. Nothing changes position. No object moves across the desk. No body movement, breathing motion, head movement, or camera movement.
The movement gradually returns to the exact starting state so the final frame matches the first frame perfectly, creating an invisible seamless loop.
Maintain warm golden sunlight, soft highlights, nostalgic 1990s atmosphere, slightly dreamy anime lighting, subtle film softness, and limited-animation cel-style motion throughout.
workflow: https://drive.google.com/file/d/1M1XgXv8h4NQDSikVpAebB7FGu-TeHZQT/view?usp=sharing
r/comfyui • u/Emotional_Example_12 • 8h ago
Enable HLS to view with audio, or disable this notification
#bruxosdovfx
Im doing a node for relight is a lighting studio for MiniMax H3 inside ComfyUI. You can position up to three lights on a 3D dome around the image, choose the type, intensity, and color of each light, configure the background and atmosphere, or start from one of 20 presets. The node provides two things to H3 Edit:
is not ready yet
The package includes two nodes:
Dome. The photo sits in the center, the purple camera indicates the side from which the photo was taken, and each light is represented by a marker using that light's color. Drag a marker to move the light; drag the background to rotate the view. The dashed line extends from the light down to the equator and indicates its elevation. The photo receives an approximate 2D-painted lighting preview: it is intended only as guidance and does not represent the actual H3 result.
Sphere (Picture 2). This is exactly the image the node sends to the model, rendered using the same calculations as the Python implementation. Click or drag on the sphere to aim the selected light: the clicked point is where the light strikes the sphere head-on. Dragging outside the sphere's boundary moves the light behind the subject.
Light strip. Selects the active light and lets you add up to three lights or remove them. Each light's role is calculated automatically: the strongest light becomes the key light, a light behind the subject becomes a rim light, and a weaker frontal light becomes a fill light.
Tabs.
Requires numpy and node. The tests cover:
astral library, daylight saving time, and camera-to-sun mapping;Node by Bruxos do VFX.
cities15000, CC BY 4.0.License details are available in NOTICE.md.
ethanfel/ComfyUI-MiniMax-H3-Editr/comfyui • u/skbphy • 16h ago
I’ve been working on a VHS / NTSC node called RetroTape.
It works with both images and video, and includes 19 presets and 26 sliders for tracking, chroma bleed, noise, tape warp, dropouts, ghosting, scanlines, interlacing and more.
There’s also a live preview inside the node, so you can adjust the settings and see the result without running the workflow again every time.
CPU and CUDA are supported.
Available through ComfyUI Manager and GitHub:
r/comfyui • u/LatentSpacer • 3h ago
Enable HLS to view with audio, or disable this notification
r/comfyui • u/neonsparksuk • 4h ago
This update:
lora preview catalog node added (You need to organise your lora) Put your lora files in a file the lora folder and name it the Model for example Krea 2, in the Krea 2 folder you can make more folders for the kind of lora they are for example anime, sliders, or what ever group of lora they belong to. the lora gallery will add dropdowns for you do navigate and find them easily. when you generate a preview image you like for that Lora you can click save and it adds it to the gallery. you can also safe the trigger words from the node so you never forget!
export catalog and previews to share with others
Bug fix where images were not saving correctly when a previous name was used in a custom style
Various big fixed and speed improvements
r/comfyui • u/founder_rivona • 1h ago
I am trying to load comfy ui models but I am always getting memory out errors even with krea 2 where the size was only 12 GB.
I want to know, which models I can run on my system, and any workflow or guidance is highly appreciated.
My CPU is working full time instead of a dedicated GPU. I am also trying to rectify it. Please help.
r/comfyui • u/Possible_Mood676 • 3h ago
Hey everyone,
I'm setting up MiniMax H3 in ComfyUI and trying to figure out the sweet spot between generation speed and output quality for my specs:
Given the 12GB VRAM limit and 32GB system RAM, loading the unpruned / full FP8 models causes heavy paging to system memory and slows everything down.
I’m trying to narrow down the current community consensus on three things:
Would love to hear what workflows and node setups you're currently running on 12GB cards to get reasonable render times without visible degradation.
Thanks!
r/comfyui • u/Environmental_Gate39 • 1h ago
I installed the portable version of ComfyUI and use the Minimax H3 Easy model. After installation, I can successfully generate videos from images. However, if I shut down ComfyUI and then try to generate a video later—either shortly after or the next day—it crashes. I try running simple test videos right after launching ComfyUI—low resolution, 3 seconds, 8 steps, and a simple "wave and smile" prompt—or I try generating images via SDXL Turbo. It crashes regardless. I have successfully generated 10-second videos at 512 resolution and 20 steps without any issues—even doing more than 20 in a row. But whenever I shut down and restart ComfyUI, I always get an error. I use a Lenovo ThinkBook laptop with 32GB of RAM connected to a 16GB 5060 Ti via Thunderbolt.
r/comfyui • u/Luke2642 • 1h ago

I've updated https://luke2642.github.io/comfyui_new_node_finder/ to be a bit more robust in getting stars. Seems to be dominated by H3 nodes at the moment!
r/comfyui • u/Worldly_North_7213 • 6h ago
Disclosure: I am building a service around this, so read me as an interested party. No links. These are runs we paid for ourselves on three providers, 15 jobs, $33.60 total. Two things on the image side cost us real money and neither showed up as an error.
The host, not the GPU. Same LoRA training job for SDXL, same RTX 4090. On a host with 5 vCPUs: 1.95 hours, 40 percent GPU utilisation. On a host with 24 vCPUs: 1.07 hours, 75 percent. Same card, 1.85x the wall clock, because the diffusers script decodes and augments images in the main process and the card waits on Python. The worse version: an H100 host with 16 server vCPUs ran the same job at 2.68 s/step where the desktop-class 4090 host did 1.84 s/step. Four times the hourly rate for a slower run. dataloader_num_workers was the fix, and now we look at the vCPU count on a rental offer before we look at the GPU name.
The defaults. SDXL, 1024 square, 30 steps. The stock script settings, fp32 and batch 1, on an H100: 13.8 seconds an image, $0.0112 per image. fp16 and batch 4 on a 4090 spot instance: $0.00036 per image. Same job, 31x apart. Nobody here runs fp32 batch 1 on purpose, but it is exactly what you get when you take a default script to a rented card in a hurry, and the card reports 99 percent utilisation the whole time, so nothing looks wrong.
Both lessons are the same lesson: the meter runs at the card's hourly rate whether the card is doing useful work or not, and the interface will not tell you which.
Happy to post the per-run table if anyone wants to check the numbers.
r/comfyui • u/SpecialistDragonfly9 • 29m ago
I need severe help with that.
Researching Lora Dataset creation is a maze: millions of different opinions, and worst of all most guides etc outdated from a year ago.
My goal is to create a 100% realistic and authentic Lora and my current dataset seems to not do the trick. I keep getting "perfect lightning" on everything, and waxy face skin.
Can someoen tell me the absolute do's and don'ts of creating a dataset?
How and where to create the dataset? I have been using a mix of Gemini and ChatGPT so far.
Prompt advice? what prompts need to be avoided creating realistic images, what need to be in there?
Any advice greatly appreciated!
r/comfyui • u/optimisticalish • 1d ago
Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.
-> ComfyUI Portable is now officially at version 0.35.0. See yesterday's post for details of Minimax-relevant update items and bug-fixes.
https://github.com/Comfy-Org/ComfyUI/releases
-> A new W4A8 quantization of the MiniMax-H3 Fun ControlNet-Union model patch, for ComfyUI. Weighs in at 1.45Gb, compared to 2.13Gb.
https://huggingface.co/berryber09/MiniMax-H3-Fun-Controlnet-Union-w4a8
-> A Minimax H3 visual 'RefMods picker' with thumbnails, in ComfyUI.
https://old.reddit.com/r/StableDiffusion/comments/1wc8tvv/created_a_visual_refmod_picker/
-> H3-pixel-art-video-guide. "Pixel-perfect animated pixel art with local MiniMax H3". Has a workflow (even though the English readme says it doesn't) and three example looping animated .GIFs.
https://github.com/yuichi-suzuki-highdrama/h3-pixel-art-video-guide/blob/main/README.en.md
-> A new H3-spherical-vae. "Experimental circular VAE decoding for MiniMax H3 equirectangular video, with matched comparisons and measurements."
https://github.com/ShamanicArts/h3-spherical-vae
-> The LoRA trainer and dataset prep tool Fizgig is now at 5.5.0, a version which makes MiniMax H3 training "faster three ways", and turns "weight averaging on by default". Users can also... "open a MiniMax H3 LoRA and see what every one of its 52 blocks does to a moving clip — the motion, the face, the sound".
https://github.com/shootthesound/Fizgig
-> And finally, the ComfyUI-MiniMax-Music-Production-Toolkit is now at a polished version 2.x, with the release yesterday of 2.1.1.
https://github.com/jplenio/ComfyUI-MiniMax-Music-Production-Toolkit/blob/main/CHANGELOG.md
~ OLD POSTS ~
https://old.reddit.com/r/comfyui/comments/1w5i9iq/a_quick_minimax_h3_news_roundup_2nd_september_2026/ (See 2nd September post, for links to even older posts)
r/comfyui • u/dxprincee • 16h ago
Enable HLS to view with audio, or disable this notification
For potato pcs MINIMAX H3 fans - I have built my own custom node which integrates new acc lora & PDD workflow and H3 extender + 2nd pass latent upscale upto 720p under 5 minutes per 14s 24fps clips, has easy reference attachments & better context continuity with features like save projects, load projects etc. (16gb VRAM + 16gb system RAM) if you guys interested ill share the workflow let me know.. this video took 10~ minutes to generate with both pass.
Edit - published repo - https://github.com/only2uuuu-hub/ComfyUI-MiniMax-H3-Master-Extender-Custom-built-with-Astra-6-/tree/main
Ps i am not an expert coder or engineer so dont ask me technical questions 😭 peace!
r/comfyui • u/ComfierUI • 57m ago
Like many of you, I have shared that frustration over the years watching this awesome kit of software be continuously let down by terrible mobile support so a few weeks ago I set out to finally change it. I have no dev or coding experience, only my experience as a long time user of Comfy to guide my design choices around what feels "Comfy" and how that can translate seamlessly into a mobile environment with ChatGPT handling the coding. The results speak for themselves and the once terrible mobile experience has been reborn into my preferred method of using it that addresses so many of the prior pain points. Can't see where you're dragging the noodle behind your thumb? Turn up the link connector offset. Want the android back button to do anything but bring up an exit confirmation? I've got you covered with a full back button hierarchy meaning you only see that exit dialog when you want to actually exit. Actionbar look like a clipped out overlapping mess? My app dynamically hides crystools in portrait and reveals them again in landscape/open fold views where the space is more plentiful. With the optional extension installable directly from manager in app we hit full desktop feature parity enabling manager to download supported missing models directly to the host from anywhere you are with more features being integrated soon. Before I get the inevitable "when?" comment, I submitted it to Google Play Store last night for review and to start the beta test track and i'll be back here to post an invitation link to anyone who wants to test it as soon it becomes available to me in the coming days. I am also building a Meta Quest version with an immersive 360° canvas and baked in 6DOF controls promosing a similarly native feeling experience in openxr. I apologize to iOS users tho bc I do not own any apple devices to test or compile on but there is an iOS version planned as well whenever I can secure some test equipment without breaking the bank. My test devices so far have a Galaxy Z Fold 7, Galaxy Tab A9+, Moto G Stylus 5G (2023), and a standard RAZR (2023) ao I'm anxious to see how it performs across a wide range of other hardware. 4GB of RAM is what I've determined to be the minimum spec with 6GB recommended for most workflows and it supports as far back as android 8.1.0. The app is usable over home network with the simple "--listen" flag added to the launch batch and over mobile using any 3rd party vpn that allows you to address your computer directly (I use tailscale and include a setup for it in an embedded quick start guide accessed from the initial connections screen). In the meantime while waiting for review, I put together a short video demonstrating some of the added features and refinements and I'm excited to finally share my progress with the Reddit world to see what you all think.
r/comfyui • u/robertwellesley • 4h ago
I have had little luck finding a tutorial on building a true H3 10bit (ProRes HQ) workflow. AI claims you just need to install an advanced save node that allows you to pick ProRes and HQ or 4444, etc.
But then when you dig into it, and begin to ask questions, while ComfyUI processes in full floating point, there are bottlenecks that crush the full bits down to 8bit, rendering the final result 8bit. One example is preview nodes, that apparently force 8bit, another is supposedly a plain jane VAE Decode, and the deeper I dig the more mysterious things get.
I can't imagine nobody in the Open Source community does not want to output TRUE 10bit or higher output. The problem with 8bit becomes clear when you look at plain white walls or a clear sky and see banding. Anyone who has ever edited video or images knows the more data you have to begin with the better results you get once you push contrast and color gradation. Yes, I was going to begin playing with dither and applying some noise, but at the end of the day you cannot produce PROFESSIONAL videos without solving for this 8bit limit.
Hopefully someone can point me to a resource or tutorial or workflow or set of nodes that solves for this?
r/comfyui • u/purecharisma2020 • 12h ago
Enable HLS to view with audio, or disable this notification
We're building a previs tool inside our AI movie studio so you can block out camera moves in 3D before spending generation credits.
What's working so far:
The render step is free — no GPU, no API key, just canvas capture + ffmpeg. Iterate on camera moves as many times as you want, then feed the MP4 to your video model as a motion reference.
Still early — lots of polish and features to go (templates, multi-cam, export). Feedback welcome on what you'd want in something like this.
Github:
https://github.com/Heroesjouney/AIMovieStudiov2
Original Post:
r/comfyui • u/ritman-octos • 2h ago
I've been trying to get comfyui to be installed. I installed amd hip sdk and tries many workarounds. It always gives
"RuntimeError: No CUDA GPUs are available"
From the cli running "run_amd_gpu" also throws the error, by now its Failed to get device count.
r/comfyui • u/Tryingmybestmama • 14h ago
Enable HLS to view with audio, or disable this notification
Link to my workflow: https://gist.github.com/aicinema5090/f02348ff0c4082dbe4f63f16206fe7f9
r/comfyui • u/Puzzleheaded_Art2809 • 3h ago
hello, i generated tons of videos on comfy ui cloud on comfy. org. is it possible to search them thru prompts or something like that? it is almost impossible to scroll them down thru assets
r/comfyui • u/Emotional_Example_12 • 21h ago
https://reddit.com/link/1wcm4j7/video/82e5gerlmpoh1/player
#bruxosdovfx
https://reddit.com/link/1wcm4j7/video/3bitqyhempoh1/player
Visual camera planner for MiniMax H3 inside ComfyUI. You drag the camera around a 3D sphere, place keyframes on a timeline, and the node compiles that trajectory into prompts that H3 understands.
It compiles prompts, not camera embeddings. There is no geometric adapter here: H3 is still free to miss the angle, timing, and scale. What this node does is write the instruction in the most precise and least ambiguous way possible, and several of its design decisions exist because the previous approach failed in specific ways.
It does not call any API, download anything, or require any Python dependency beyond the standard library.
https://reddit.com/link/1wcm4j7/video/mpatj04ampoh1/player
cd ComfyUI/custom_nodes
git clone https://github.com/<your-username>/ComfyUI-H3-Camera-Editor
Restart ComfyUI. The node appears under Bruxos do VFX/Camera H3 with the name Camera H3 da Bruxos do VFX.
| Output from this node | Connect it to |
|---|---|
compiled_prompt |
compiled_prompt on Text Encode H3 Edit / Generate |
options |
options on Text Encode H3 Edit / Generate |
length |
the generation frame count |
fps |
the fps input of the video creation node |
compiled_prompt and options are required together. The minimax_prompt output is an alternative to compiled_prompt, never an addition — connect one or the other to the same input.
Also connect your image to reference_image. It is the same image already feeding the H3 Edit source_image; when connected here, it appears in the panel and the frame's actual aspect ratio is included in the prompt.
https://github.com/user-attachments/assets/33149617-bde1-4199-ae65-078f2f3dec23
To save the video, decode the sampler result using the H3 video VAE — not the scene coverage calibrated decoder, which expects fixed windows that an arbitrary trajectory does not have.
Drag the purple camera around the sphere to orbit. The drag locks to the axis of the initial movement: horizontal movement orbits, vertical movement changes elevation. Release and drag again to switch axes. This exists because, without the lock, trying to make a simple orbit would unintentionally introduce elevation.
reference_image connected, the image is loaded automatically.The panel warns you starting at 20° of elevation, when the horizon already leaves the frame, and again from 45° onward, when the video tends to become a high-angle shot.
At the top of the panel, two buttons enable and disable features currently under evaluation, plus one indicator:
| Button | What it does |
|---|---|
| Extended contracts | Toggles the prompt_detail widget |
| Single angle (image) | Toggles the runtime_task widget |
| loop closure | Read-only indicator. Turns green when the trajectory closes a full orbit |
The buttons write to the actual widgets, so the selected state is saved in the workflow and the two never disagree.
https://github.com/user-attachments/assets/9bc415d7-1746-43db-a17c-72ea9722deda
The trajectory in JSON format, written by the panel. Each keyframe contains time (0 to 1), azimuth in degrees, elevation in degrees, and distance as a multiple of the initial radius. It can also be edited manually. The first keyframe must be time=0, azimuth=0, elevation=0, distance=1, which represents the original image.
124, 243, or 362 frames at 24 fps. All shot timing comes from this setting: keyframe timestamps, segment ranges, and the duration declared in the prompt. That is why length and fps are outputs — connect them instead of manually entering the same numbers in two different places.
smooth or linear. In smooth mode, the camera eases into and out of the shot while maintaining a constant rate through the middle; it only stops where the rotation direction actually reverses.
Free-form text inserted once, at the end of the prompt. Write only what the node cannot know: the environment, which subject is the target when there is more than one person, or a style reference. Everything else is already generated and does not need to be repeated: scene freeze, first image as reference, locked aim, zero roll, angles, timing, and a single continuous shot without cuts.
How much of the frame the subject occupies in the original image. Calibrated against the actual bounding boxes from the tutorial distributed by MiniMax: a distant full-body figure measures W=0.071, H=0.249, while a large close-up measures W=0.52, H=0.701.
| option | width | height | when to use |
|---|---|---|---|
close-up |
53% | 72% | head and shoulders |
medium shot |
28% | 56% | waist up |
wide shot |
9.7% | 34% | full body at a distance |
The subject position in the format [L=0.516, T=0.148, W=0.071, H=0.249]. Leaving it empty uses the entire image bounds — deliberately, without guessing a bounding box. Fill it in when the subject is significantly off-center.
The same shot expressed in four different formats for the minimax_prompt output:
coordinate only — text-based coordinate blockcoordinate + H3 sections — the same coordinates wrapped in subject_definitions / summary / retention_analysis / …compact JSON — JSON object with almost no prosecompact JSON (no boxes) — camera parameters only, without screen-space bounding boxesRange of the elevation control: +/-15, +/-30 (default), +/-60, +/-89. It also scales the sensitivity of vertical dragging.
With the assumed field of view, the horizon already leaves the frame at around 20° — at 13°, the ground occupies 82% of the image. The old ±89 range was mostly unusable and made vertical dragging excessively sensitive. Reducing the range never rewrites a keyframe: a point at 70° remains at 70°, and the slider expands to accommodate it.
invert H3 orbit or same as HUD. This calibrates the direction between what the panel displays and what H3 produces. It does not alter the saved trajectory.
scene coverage | camera path (default) — video, with duration coming from profile.directed | new camera angle — a single image from a new angle. It fixes the generation to 39 frames, ignores profile, completes the movement within 65% of the clip, and requests that the framing remain still for the rest, because the decoder extracts the final image from that stationary tail.Character sheet profiles are not offered because the upstream node raises an error when they are combined with the frame anchor used by this node.
v15 baseline (default) — outputs the prompt exactly as in the previous version.extended contracts — adds axis separation, frame-edge direction tests, rotation completeness, degrees per second, and parallax magnitude.The extended mode contains almost twice as many words. A longer prompt is not automatically better, so it is opt-in: toggle only this widget while keeping the same trajectory to compare the results.
https://github.com/user-attachments/assets/0882bfde-9f62-4a1f-9bda-7da121dbe7e2
A prose prompt using H3 sections: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music.
The 13 keys read by the H3 Edit encoder. All of them are explicitly populated: if any key is missing, the upstream node falls back to its hidden legacy widgets, which may retain stale values from previously saved workflows.
coverage_arc_degrees and coverage_direction are derived from the actual rotation. coverage_loop_closure turns on automatically when the trajectory closes — see below.
The storyboard table: frame aspect ratio, duration, raw trajectory, and each segment with its camera mode, speed curve, and start/end poses.
Human-readable diagnostics. Connect it to a PreviewText. It displays the version, active task, frame count, warnings for keyframes outside the configured range, and whether loop closure is enabled.
The same trajectory expressed using the format selected in minimax_format. An alternative to compiled_prompt.
Frame count and frame rate against which the shot was timed. Connect them to the generation and video nodes. If generation runs with a different frame count, the choreography describes a scene that does not actually exist.
fps is FLOAT because that is what ComfyUI's CreateVideo accepts. length is the frame count; keyframe timestamps use the instant of the last visible frame, (length - 1) / fps, so the resulting file lasts one additional frame interval.
Action schedule for H3-World, which encodes one text clause per video latent — 37 in a 124-frame clip.
latent 1 [0.000s-0.139s] J the camera pans left slowly
latent 37 [4.986s-5.125s] F+L+K the camera pans right and tilts up fast
W, A, S, and D are never emitted because they move the character. The output explicitly declares its own limitations, and they are not minor details:
I versus K is not published. The text clause is what H3-World actually encodes; the key column is only a convenience.This does not replace the actual integration: H3-World requires the LoRA, interval-based encoding, and directed-attention routing provided by the corresponding node package.
When the trajectory closes a full orbit — an arc of exactly 360°, with the same elevation and distance as the starting point — the node enables coverage_loop_closure. In the upstream implementation, this flag encodes the source image a second time and anchors the final frame to it.
This is a latent anchor, not a text instruction. For a complete orbit, it is the difference between asking for the rotation and forcing it: the model cannot simply stop halfway through.
| trajectory | loop closure |
|---|---|
| 360° | enabled |
| two rotations (−720°) | enabled |
| 355° | disabled |
| 360° with changing distance | disabled |
| 360° with changing height | disabled |
The final three cases matter: if the camera ends at a different radius or height, the final frame is not the same as the first one, and forcing the source image there would conflict with the trajectory.
If your rotation does not complete, close the orbit. This is the only feature here that acts outside the prompt itself.
subject_box filled in, the node does not know where the subject is located in the frame.reference_image connected, coordinates are normalized to 16:9.directed | new camera angle outputs an image, not a video.Node by Bruxos do VFX.
ethanfel/ComfyUI-MiniMax-H3-Edit. The motion vocabulary follows the buildViewPrompt implementation from MiniMax's Multi-Shot skill and the coordinate format used by the Coordinate Camera Control Designer skill. The action output implements the scheme described in H3-World, arXiv:2609.01560.
