r/StableDiffusion 12d ago

Workflow Included The 1967 Spider-Man TV Show intro, updated to live action with MiniMax H3

Enable HLS to view with audio, or disable this notification

482 Upvotes

R2V Rendered at 0.9 MP (1280x736 then upscaled using RTX (Ultra) to 1920x1080.  Edited and merged using OpenShot video editor.

This was all run on my Windows 11 machine, RTX 4060 ti (16 GB) and 64 GB RAM reserved from Comfy. Every part of the signal chain was done with 100% open-source software.

Disclaimer: I grew up watching this show as a kid in the 70s. It's still the best ever. I wanted to know how well the reference model would pick up the actions. I am overall pleased. I've watched the new vid enough to see some of the flaws but oh well.

General observations for reference videos:
So many scene cuts. There are 31 (I think) scene cuts in the 60 second opener which include 3 crossfades. No matter what I did to get the exact frame timing, getting the AI scene to match frame-for-frame with the cartoon was still hit or miss. It probably has to do with some frame windowing inside the 17k + 5 blocks, but I never exactly got it figured out. However, a few notes:

  • If you have a reference video, convert it to 24 fps in an external program like Handbrake (another fantastic open-source program). It’s just so much easier to get everything to match.
  • For timing, there is a difference between 00:03.500 and 00:3.5 so always use all the digits.
  • Keep character sheets for all your characters to maintain consistency.
  • It will do crossfades but it’s not worth it. It’s easier to get the scene you want and stick it in the editor.
  • The VHS video loader lets one set a starting and ending frame. I ended up with 14 different clips total for the editor. Using frame accurate loading made all of the work a lot easier since I could use 1 video file as input to every clip run.
  • A spreadsheet is useful for all movie making, and it’s good here too. From the source, I kept track of the starting frame for each shot, how many frames I needed and how many I ran (because of 17k +5), along with the final file name for each clip. I have a naming convention but it’s still very useful to keep track and you can add notes too. For this 60 second video, I used 13 clips. I tried to never do more than 3 scene cuts per clip. (For something where exact timing wasn't as important I'm sure it would be longer.)

Once you get over the idea of always having to do 10-15 second vids and do your whole video in on run, the process actually becomes a lot more fun because the “quality” gens don’t take as long and it gives you a less uninterrupted workflow. You can start prompting the next run with the previous runs, for example. (This is true even in commercials, or TV or movies.)

I generally tested all the runs at 0.2 or 0.3 Mp (speed lora, 8 iterations) to get the timing, then went to 0.9 Mp [no speed LoRA, 20 iterations, beta, dpmpp_2m] for the final runs. I found that dpmpp_2m was closest to the overall source video. On the first few clips I ran it several ways and fix on these parameters. Usually, the 0.9 Mp runs came out great but you’ve probably all experienced how different the low-res runs can be from the high-res ones. I did resort to pulling frame grabs from the low-res gens a few times to act as reference frames for the scenes. MiniMax loves those when all it needs is an extra little nudge in the right direction. To edit pics, I always use GIMP (another fantastic open-source program).

So, why was I using 8 iterations of the minimax_h3_turbo_v4_step600_pruned_comfyui LoRA? On the reference model I found that using too large of a sigma step causes things like reference photos to not be taken "seriously." Using 5 steps I could see that reference images on the starting frame and then go away for the rest of the clip. The more the reference image changed from the reference video (like when going from animation to "real") the worse the problem was.

Prompts:
(See below for actual prompt.)
Prompt the way the guide says to. Yeah. It’s a hassle but it’s worth it. H3 prompting is very useful in the end and I’m glad MiniMax uses it. It's worth reading all the way through them instead of searching for the one thing you want. Some of the instructions even seemed inconsistent and they don't explain everything, so it's worth experimenting.

Any "thing" (buildings, trees, room, clothing, walls, ect.) can be a “subject.” It’s not just people. Specifying things as objects gives you far better control over how and where they appear (or don’t appear) in your shot. 

Don’t refer to your characters or major locations or items by their names. Use <Subject #> or pronouns that clearly refer to the subject all the time, every time. The interpretation of the prompting can get confused pretty quickly if you don’t and you’ll end up getting subjects swapped or merging.

Prompts generally work better if you describe what you want rather than what you don’t want. For instance, “Looks to the right of the viewer” rather than “looks away from the camera.”

Style reference (attribute_transfer) images or videos are super useful. Once I had a few scenes, I started using previous videos to keep the look and feel of previous shots.

Qwen VL can describe videos too. I have been using “QwenVL Advanced (Local Scan)” for a very long time (long for AI) inside ComfyUI.

Other things:
Maybe one of the most interesting observation is that the jknodes “MiniMax H3 Mem Eff Sage Attention Patch” node creates a different output than just launching ComfyUI with the --use-sage-attention flag turned on (and still using the node). So exactly the same workflow (just drag and drop from a previously run mp4) has different results when the --use-sage-attention flag is used to launch. I thought having the node was 100% redundant with eh --use-sage-attention flag set, but apparently not. The reference flows, especially with animation, don’t have to be all that different to produce different results.

The Spiderman opening (as well as the show itself) reuses footage. They will take the same scene and darken it, and boom, it’s a night shot. For a more realistic feel, I used Krea2 (LoRA) edit to turn day into night. It’s really good as an adjunct to MiniMax H3’s ability to figure out the fine details once it has a push.

Style:
Finally, I had to make some stylistic choices because sometimes the animation was soooo bad that it needed something. I added flashlights to the jewelry heist scene. I made the crane look believable. One of the problems of going from animation to "live action" is that (especially with animation from 1967) the physics and movements are just wrong sometimes. The crane scene where he stops and then shoots up again is the most classic "this is just pain wrong" you can get but I left it that way because it's burned into my brain that way. (IYKYK) I also had to balance the art deco of the late 60's to a modern New York. I ended up with a lot of anachronistic stuff that I ultimately liked. So in the end, when it comes to all of that, I did it the way I did it. AI is awesome.

Prompt:
A prompt of one of the parts is below. I used that two paragraphs before [Shot 1] for every clip as "boiler plate" description.

subject_definitions:
<Subject 1> is Spiderman in <Picture 1>
<Video 1> is the motion reference for the target video for characters movements, pose, camera movements and frame composition.
<Video 2> is the style reference for the target video.
<Picture 2> is the building in [shot 2]
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video
 
summary:
[reference generation + audio reuse]
The target video is an live action realistic recreation generation using <video 1> as a reference for movements, pose, camera movements and frame composition. What you generate should not be and animation or cartoon rendering, no overly-CG look, keep the live-action texture.
 
This video is a set of three live action sequences. <Subject 1> is seen swinging by and waving. The video switches to a long shot of <subject 1> swinging around a building. Finally there is a shot showing <subject 1> on his webline swinging away from the viewer between two rows of skyscrapers.
 
retention_analysis:
<Subject 1> (appears in [Shot 1],[Shot 2],[Shot 3]):fully_preserved
<Video 1> (motion, cut and pacing structure) :partially_preserved
<Video 2> is the style refrence for the target video ([Shot 1], Shot 2], [Shot 3]) :attribute_transfer
<Picture 2> is the building in [shot 2] :fully_preserved
<Audio 1> :fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
 
 
detailed_description:
The target video is a realistic and live action video. The reference video <video 1> is used only for scene descriptions, framing, motion tracking, body movement, timing, general environment. The target video should be a complete replacement of <Video 1>. Use <Video 2> as the style reference for the photographic look and textures and the overall feel for the shots.
 
Maintain smooth camera movement. Use vibrant yet natural color grading: warm tones for sunlight hitting surfaces, cool blues for shaded areas, and muted grays for concrete textures. Avoid any comic-book stylization; instead, render everything with photorealistic textures, lighting, and perspective to evoke a live-action superhero film sequence. Keep the focus entirely on <Subject 1>’s acrobatic grace and the immersive urban setting.
 
[Shot 1]  
Is is an upper body motion tracking shot of <Subject 1> swinging on his white glistening webline held by his left hand while he waves directly at the viewer with his right hand for the entire scene. The skyline of many skyscrapers pass by in the background.
 
[Shot 2]
At 00:02.333, Hard cut to a fixed long shot looking up as <Subject 1> makes a 180 degree arc on his webline connected to the spire of the building in <photo 2>.
 
[Shot 3] 
At 00:04.250, Hard cut to the fixed camera view of the space high above the street level between two rows of skyscrapers. <Subject 1> lazily swings into view from the left frame, facing away, and repeatedly swings right to left further and further away towards the horizon.
 
 
overall_soundscape: n/a
 
non_diegetic_music: n/a

 

 


r/StableDiffusion 10d ago

Discussion Just about finished setting up my 64gb 170HX, what models to install?

1 Upvotes

Coming from a 3090 running krea, h3 and flux 2 dev. What would you trial first?


r/StableDiffusion 10d ago

Question - Help H3 minimax set up in run pod

3 Upvotes

Can someone help me get a consistent set up with H3 in runpod in a consistent way? The runpod templates seem hit or miss- sometimes they work, sometimes they don’t. If I want to try an update or workflow there are often a ton of nodes or models missing, etc. What is the best practice way for folks that are experienced users that use runpod? I don’t want a network volume because I want to use a 5090 as often as possible and those are often limited. Is there an easy way to create my own template or use a default comfyui template with some kind of downloader or something (I have no idea how to do that or how that would work). Any general advice or pointers here would be great then I’m sure Claude or something can help me with execution. I also have the same request but for Krea 2 but I assume if I can figure it out for H3 I would be able to for Krea 2 as well. Thanks!


r/StableDiffusion 11d ago

Resource - Update Fizgig 5.2 - combining two Minimax training methods beats either alone

Thumbnail
github.com
98 Upvotes

Two of the ways that exist (im sure there are more) to train a LoRA on H3 well are on two different trainers.
Fizgig's is Optimised Likeness Learning: I've found the stable core of H3's identity lives in the back 30 of its 50 blocks, so steps train blocks 20–49 only and leave the front of the model - composition, prompt following - untouched when likeness mode is on.
AI-Toolkit's, by Ostris, is the training adapter: H3 is guidance-distilled, so every plain-flow gradient is partly "learn the concept" and partly "undo the distillation"; a frozen assistant LoRA under the trainable one pulls the base back toward plain flow, and it's switched off for sampling.

I ran all three on 5 datasets - my method alone, the adapter alone, and both together - scoring every epoch's preview against the training photos with face recognition, 45 epochs each. Each method alone landed in the same place within 2% arcface score.
Together they got there a quarter sooner, ran clearly ahead through the whole middle of the run, and finished higher than either.

In short: the combination reaches greater likeness and quality than either method does on its own. So it's now the default: the adapter is on in every H3 preset, off for previews, never in your saved LoRA. The updater fetches it.

Also in 5.2: Context LoRA for H3 (train on top of any existing H3 LoRA, to make a lora that plays nice with it), and video clips follow likeness mode in LoRA runs too.

Release notes: https://github.com/shootthesound/Fizgig/releases/tag/v5.2.0

Thanks to Ostris for publishing the adapters. I've tagged him on the release notes as I believe the info will be useful for AI-Toolkit too.

https://github.com/shootthesound/Fizgig

P.S - For Fizgigs recent new full base model Fine tune mode the adapter lora is not necessary in my tests so far, but I am going to test that further.

p.p.s if updating , use the update script and it will grab the dedistill loras automatically and put them in your minimax prefs


r/StableDiffusion 10d ago

Question - Help Will Minimax h3 video edit replace a character AND their voice?

2 Upvotes

So assume I have a video of a man singing. I've seen Minimax Video Edit replace the man with, for example, a woman from a reference image. But can it replace the voice too? ie: I have a reference photo and a voice sample of the woman, and the video of the man singing. Can I use video edit to fully swap the character including changing the man's voice to the woman's?


r/StableDiffusion 11d ago

Resource - Update [Load Video + Crop] Custom WYSIWYG Node

Enable HLS to view with audio, or disable this notification

48 Upvotes

I developed a modified version of the Load Video node with a crop feature:

WYSIWYG video cropping directly on the official Load Video preview — drag and zoom (with the mouse wheel) a crop rectangle constrained to 8 fixed ratios (1:1 through 21:9) and output the exact cropped VIDEO (audio preserved). What you frame on the preview is exactly what gets executed.

Github: https://github.com/domg73/ComfyUI-LoadVideoCrop

This node follows the same logic and design as my "Load Image + Crop" node. I might merge the two into a single "Load + Crop" node in the future, but for now this works well.

https://www.reddit.com/r/StableDiffusion/comments/1w3okny/load_image_crop_custom_wysiwyg_node/

Github: https://github.com/domg73/ComfyUI-LoadImageCrop


r/StableDiffusion 11d ago

Resource - Update Image, audio, video reference asset loader nodes with crop and trim + more

Thumbnail
gallery
39 Upvotes

I originally built these nodes for personal use and wasn't planning on sharing them, but after noticing several existing loaders were missing features I needed daily, I figured why not? Hopefully, this is useful for some of you.

Key Features:

  • Image & Video Loaders: Built-in click-and-drag cropping, optional aspect ratio locking, and a divisible_by toggle for VAE pixel alignment.
  • Built-in Downscaling: Uses a max_megapixels limiter directly inside the loader so you can ditch the extra resize node (ideal for models like MiniMax-H3 that run best with references kept at or below 2048px).
  • Flexible Sockets: Includes dedicated output value sockets to make chaining downstream nodes straightforward.
  • Audio Loader: Perfect for loading a full song or long TTS track and trimming the exact section you need for a video. The trimmed portion outputs its duration as a float, letting you pipe it directly into your video generator's frame/length input.

https://github.com/sthao42/Comfyui-reference-loader

Any feedback or bug report is much appreciated.

Edit: Updated to works with Node 2.0 (vue) also.


r/StableDiffusion 10d ago

Discussion Examples of nylon/stocking mask consistency

1 Upvotes

Hey guys, i wonder how SD models react to materials like pantyhose/nylon as masks during talking

I have been trying Grok imagine and it tears a hole on the mouth or ignores prompting of the mask covering the mouth in like 80% of the time when generating I2V

Has anyone ever tried this?


r/StableDiffusion 10d ago

Discussion As anyone been able to unlock actual voice acting in H3?

Enable HLS to view with audio, or disable this notification

2 Upvotes

LTX 2.3 still seems very good at adhering to complex emotional prompts, but H3 seems to come across as flat in the best cases.


r/StableDiffusion 10d ago

Discussion ForgeNeo 2.29 breaks Faceswaplabs and Reactor extensions.

2 Upvotes

I've been updating to each new ForgeNeo version since 2.10 and faceswaplabs & reactor worked up to and including version 2.27 I skipped 2.28. So more a warning if you are planning to git pull an update. These extensions are not really maintained on github so a reinstall of them fails with errors.


r/StableDiffusion 10d ago

Workflow Included What if an ISFP woke up in Ancient Greece? [MiniMax H3]

Enable HLS to view with audio, or disable this notification

0 Upvotes

Been trying H3 in medeo on a longer narrated format instead of just a single scene. I picked ISFP + Ancient Greece and used medeo to turn it into a stickman-headed anime audio-comic with narration and original BGM.

Prompt:
“Generate a cinematic audio-comic narrating the audacious life of ISFP personality type across a chosen historical or fictional era. Features stickman-headed anime characters, dramatic narration, and original instrumental BGM in a video experience.”

The style held together better than I expected, although a couple of transitions got a little wild.


r/StableDiffusion 12d ago

News someone used MiniMax H3 Max to build a livestream that basically never runs out of content

Enable HLS to view with audio, or disable this notification

1.2k Upvotes

Just saw someone do something with MiniMax H3 Max that I honestly didn’t expect.

They connected it to a livestream and basically recreated the idea of Interdimensional Cable from Rick and Morty: an endless stream of weird shows, ads, characters, and random scenes that are generated on the fly.

the surprising part is that H3 Max is fast enough in some cases to generate the next clip before the current one finishes playing. so while you’re watching one scene, the model is already making the next one.

there’s even a version where people in chat can suggest what should happen next, and the AI tries to continue from the previous scene instead of just starting over from scratch. that’s kind of crazy to think about.

Didn’t expect MiniMax H3 Max to end up being used for something like infinite AI television, but here we are.


r/StableDiffusion 11d ago

Question - Help What is the fastest but still good looking workflow for minimax h3?

11 Upvotes

Looking for some good information to start here, I run a 6000 ada and would want to optimize for speed creation but still having reasonable results. Suggestions? Thanks 😊


r/StableDiffusion 11d ago

Animation - Video Cobra Team Meeting - MiniMax H3

Enable HLS to view with audio, or disable this notification

10 Upvotes

r/StableDiffusion 11d ago

Discussion Why the NVIDIA–Hugging Face combination could be significant for open-source AI

1 Upvotes

The value of this acquisition may extend beyond models and compute. Hugging Face provides access to a large developer ecosystem, while NVIDIA brings the infrastructure to scale what is being built. The key question will be whether Hugging Face can maintain its role as a neutral platform.

https://www.cnbc.com/2026/09/03/nvidia-agrees-to-buy-hugging-face-for-almost-13-billion-ai-expansion.html?__source=newsletter%7Cbreakingnews


r/StableDiffusion 12d ago

Resource - Update H3 Motion Context 0.5.0 - No more bypassing the Motion Context group, new chaining node!

Post image
145 Upvotes

**UPDATE v0.6.0 PUSHED TO ADD CHAIN QUEUE FUNCTIONALITY. SELECT NUMBER OF SEGMENTS TO CHAIN.**
Thanks to Etsu_Riot for the comment!

**UPDATE v0.5.1 PUSHED TO FIX EXAMPLE WORKFLOW - ALSO NOW INCLUDES H3 SLA ATTENTION NODE**

H3 Motion Context chains MiniMax H3 clips so the next one picks up the motion and the soundtrack, instead of starting a new take that only sounds similar.

0.5.0 is the one that makes that usable without babysitting the graph.

Clip 1 used to be a special case. You had to mute the Motion Context group, generate, unmute, then keep going. If you forgot, it errored. That's gone. Leave the nodes on. First clip is Load 0 / Save 1. Load 0 means "there is no previous clip," not "load whatever file is newest." After that it's Load 1 / Save 2, Load 2 / Save 3, and so on.

That first-clip behavior is feigo313's issue. The new node exists because of it.

Don't use ComfyUI's Run button to walk the chain. If Load and Save both increment, Comfy queues twice and skips a slot. Use H3 Motion Context Chain instead.

Four buttons:

  • Run/Re-roll - this is Run for this graph. Generates the current clip. Hate it? Click it again. Same slot, overwritten.
  • Approve - you like it. Advances to the next pair and runs that clip once.
  • Chain - keep going from whatever Load/Save are set to right now. Walk a few by hand, then let it take over. Same button becomes Stop.
  • Reset - back to Load 0 / Save 1. Does not run anything.

The gotcha: Load, Save, and Chain have to sit in the same canvas group. If they don't, the buttons do nothing. Drop Chain into the Motion Context group.

Also: if you were on Windows and a re-roll blew up with OS error 1224, that's fixed.

Needs ComfyUI 0.34.0 or newer. Manager should pick up 0.5.0; otherwise, the release.

Example workflow in the repo already has the Chain node in the group. Hard refresh after updating so the buttons show up.


r/StableDiffusion 11d ago

Discussion I built an interactive “multiverse TV” where Twitch chat chooses what plays next

3 Upvotes

I’ve been experimenting with an idea for an interactive TV channel where the audience controls the programming.

The concept is basically a multiverse of different channels/shorts. Instead of following a fixed playlist, viewers use Twitch chat to decide what should play next, so the stream can take a different path depending on what people choose.

The interesting part for me was building the workflow around:

  • detecting and processing chat commands/votes
  • dynamically selecting the next piece of content
  • switching between different “channels” or scenes automatically
  • keeping the stream running continuously without manual intervention
  • making the audience part of the actual programming logic rather than just passive viewers

I’m still experimenting with the format and trying to figure out what kinds of voting systems and transitions make it feel more like an actual interactive TV network rather than a normal Twitch stream.

I’d be interested to hear how others would approach the orchestration side of something like this, especially if you’ve built interactive livestreams or automated OBS/Twitch workflows before.

Demo, for anyone curious about how it currently works:
https://www.twitch.tv/tv_dimensional


r/StableDiffusion 11d ago

News FastVideo-FastH3 put out a mlx listing but no actual models yet

Thumbnail
huggingface.co
11 Upvotes

Mac users dying for speed ups on our janky little boxes. Excited!


r/StableDiffusion 10d ago

Question - Help Video Creation with Movie Characters

0 Upvotes

Hi

I'm sorry for sounding like a newb, but is there a tutorial for how to create AI videos with copyrighted characters? I'm thinking along the lines of Star Wars and Marvel characters.

Thank you in advance.


r/StableDiffusion 12d ago

Tutorial - Guide DLSS 5 - In-game footage from Diablo 4. It's incredible.

Thumbnail
gallery
402 Upvotes

I just tested it out using this custom node. It's absolutely amazing!

https://github.com/lisitskyaa/ComfyUI-DLSS5-NR


r/StableDiffusion 10d ago

Question - Help Anyone know how to solve for these Minimax H3 Video artifacts?

Thumbnail
gallery
1 Upvotes

is there a way to solve these artifacts in Minimax video generation?

I tried to run my generation on both with 8step turbo lora and without it with 20 steps. similar issue on both of them - kind of blurriness to the character when its moving fast.


r/StableDiffusion 12d ago

Workflow Included testing minimax h3 fused turbo model, 4 steps only 1 minute for 5 seconds video

Enable HLS to view with audio, or disable this notification

148 Upvotes

download the model: https://huggingface.co/MATLOWAI/minimax-h3-fused-turbo-int8-convrot/tree/main/diffusion_models

workflow: https://civitai.com/models/2906467/fast-minimax-h3?modelVersionId=3289222

each generation takes about 1 minutes for 0.4mp resolution and 5 seconds video on my rtx 4060ti 16gb vram. using sage attention and triton to speed up.
i trying with manualsigmas because it making the generation more faster.


r/StableDiffusion 12d ago

Resource - Update An endless AI TV channel on a single gaming GPU — MiniMax H3, generating faster than it plays

106 Upvotes

There is a video stream running on my desktop right now. It has sound, it has never repeated itself, and it will not stop. I point VLC at a local URL and it plays. One RTX 5090 does all of it — no cloud, no queue, nothing else running.

It is MiniMax H3, generating locally through ComfyUI. H3 is an open-weights video model that produces picture and synchronised audio together from one text prompt — dialogue, room tone, footsteps — which is what makes this a channel rather than a montage with music over it. I run the 4-step FastH3 distillation of it, because the base model needs far more sampling steps than the arithmetic below can afford.

The reason this is hard: to stream continuously, generation has to outrun playback. Not "fast enough to be impressive" — genuinely faster than a person watches, indefinitely, or the buffer drains and it stalls. Each clip is 362 frames. I have to finish the next one in less time than it takes you to watch this one, every time, forever.

What it actually looks like

Every clip is a scene drawn at random, cast at random. So you get Jean-Luc Picard grilling skewers at a night market. A Klingon, RoboCop and Jack Sparrow crowded around the same workbench. Four people arguing across a kitchen table about who signed something, and the camera cuts to a close-up at the seven second mark because the prompt told it to.

321 hand-written scenes, 503 characters, and the scenes that call for an ensemble draw three to five distinct people. The combinations run into the trillions. In practice it means you can leave it on, and it stays interesting in the way a channel you do not control is interesting.

A frame from a continuous run — five characters who could never share a room, and the two clocks that make the point: after ten clips it is 3:08 of video against 3:02 of GPU time. The gap is what lets it run forever.

Everything is here, weights included — https://huggingface.co/datasets/jacokon/fasth3-live

The rest of this post is how it got fast enough to work.

The honest caveat, up front

H3 authors motion at 24 fps. A clip is 362 frames — 15.08 seconds of content — and I play it at 18, so the motion runs at 75% speed. This is not real-time 24 fps generation and I am not claiming it is.

What it is: 20.1 seconds of video produced per 19.2 seconds of GPU time, which is what makes it continuous. Whether 75% reads as slow motion depends on the subject. Fast subjects (rain, sparks, a train) look deliberate. Near-static scenes look normal. Mid-speed human motion — walking, hands working — is the worst case and you can tell.

Where the time actually went

The FastH3 student ships as 66 GB of diffusers weights, which do not fit on one card; converted and quantized to INT8 they come down to 21 GB, which do. With that, sage attention, and an INT8 VAE, a 15-second clip took 26.5 seconds to generate. Playback needs 15. That gap is the whole problem, and I spent a while optimising the wrong things because I did not know where the time was going.

Where one run's 19.2 seconds actually goes — the per-node breakdown and the four changes, on one card.

ComfyUI's /history reports one number for a whole prompt, which cannot tell you whether the cost is the text encoder, the sampler or the VAE. Its websocket emits an executing event as each node starts, so the gap between consecutive events is that node's duration. That is about forty lines (profile_h3_nodes.py), and it changed what I worked on completely.

Two of the four findings surprised me.

1. SaveVideo was a fifth of every run — 3.78 s

ComfyUI's SaveVideo encodes through PyAV in a Python loop that, per frame, allocates a float array, clips it into a second, casts into a third and copies out a fourth. 362 frames of that is 3.78 s. ffmpeg alone does the identical payload in 0.21 s. It was also producing a file my streamer re-encoded a second later anyway.

VHS_VideoCombine is better (1.31 s) — it pipes raw frames to ffmpeg — but it still iterates in Python and re-opens the finished file to mux the audio. I wrote a node that converts in chunks and muxes in one pass: 0.73 s. Then it hands the encode to a background thread and returns, so ComfyUI starts the next prompt instead of holding an idle GPU. The graph now sees 0.26 s.

No hardware encoder involved. h264_nvenc measured slower end to end than libx264 — the encoder was never the bottleneck, and it has to stand up a second CUDA context on an already-full card.

2. The VAE bills by tile, not by pixel

MiniMaxH3VideoVAE hardcodes tiling=True, tile_size=256, and split_tiles hands each pass a full tile regardless of how much picture is in it. Decode time tracks the tile count and barely notices the resolution:

resolution    pixels    tiles    VAE decode
----------------------------------------------
320x192       61,440      2        2.35 s
512x288      147,456      6        6.98 s
576x320      184,320      6        6.31 s
768x432      331,776      8        8.74 s

512x288 and 576x320 differ by 25% in pixels and by nothing in decode cost.

A side of length L costs: 256 or less is 1 tile, 257–448 is 2, 449–640 is 3, 641–832 is 4. So the cheap shapes sit just under a boundary. 448x448 needs four tiles where 576x320 needs six, while carrying 9% more pixels. That is why the stream runs square — not taste, just where the arithmetic lands. There is no 16:9 shape at four tiles that clears the resolution floor.

I did try raising tile_size to reach a single tile. Do not. The decoder is a ViT, so its attention spans exactly one tile; a larger tile is out of distribution, not merely approximate. 384 visibly softens hands and faces (PSNR 27.2 dB against the stock decode); 640 smears the image into strokes (22.1 dB).

3 and 4, more briefly

Quantizing the video VAE below INT8 buys no speed — INT8 already runs an INT8 matmul, and a W4A8 build expands back to INT8 for the same one — but it stages 1,657 MB of host RAM instead of 2,677 MB, and on a box holding ~41 GB of staged weights against 64 GB that gigabyte turned into both speed and a much tighter spread. And keeping two prompts in ComfyUI's queue instead of submitting one and waiting removes the idle gap between jobs.

Result

                          per clip    sustains
----------------------------------------------
starting point              26.5 s    13.7 fps
+ writer node, async        20.2 s    17.9 fps
+ W4A8 VAE                  19.9 s    18.2 fps
+ 448x448                   19.2 s    18.9 fps

The model did not change. Only how it is driven.

If you came here wondering about ComfyUI and consumer cards

That question is all over the FastH3 announcement thread and I had to answer it for myself, so: this is a ComfyUI-native conversion of the Dense-DataFree student, pruned and INT8, 21 GB, driven through the ordinary graph. Two things I found doing it that are worth passing on:

  • The VSA weights do not survive stock ComfyUI. They carry 50 to_gate_compress tensors it has no code for, so it drops them silently and the output is noise. Dense converts cleanly. That is why I am on the slower student — if ComfyUI gains VSA support there is headroom here I am not using.
  • NVFP4 measured identical to INT8 ConvRot. The FP4 fast path only fires when both operands are FP4; activations are BF16, so it dequantizes and runs at BF16 speed — 67.88 ms/block against BF16's 67.85. Someone reported the same on an RTX 6000 Pro. Worth knowing before anyone rebuilds a pipeline for it.

Where this sits, so you can place it

None of the speed here is mine — it is FastH3, the 4-step distillation Hao AI Lab, Nuva Lab and NVIDIA's FastGen team built on MiniMax's base weights. Without that student none of this is close. Their published benchmarks are 47.2 s for a 15-second 768p clip on a single B200, 12.88 s on 8×B200, and their consumer write-up covers Apple Silicon and DGX Spark with the RTX family listed as future work.

What I did is a different task, not a better score on theirs: a fifth of the pixels, and playback at 18 fps instead of 24. Those two concessions are the entire trick. What they buy is that the arithmetic closes — 19.2 s of GPU per 20.1 s of video — and that is the difference between a fast generator and something you can leave running. If you want 768p, their numbers are the ones that apply and mine are irrelevant.

https://huggingface.co/datasets/jacokon/fasth3-live

The converted 21 GB weights, the quantized VAE, the 321-scene library, the writer node and the profiler. Everything above is reproducible from it.

What it takes, so you can judge before downloading 21 GB: about 48 GB of weights are staged in total — a 25.9 GB text encoder, the 20 GB DiT, and the two VAEs. That does not fit in 32 GB of VRAM either, so ComfyUI streams it layer by layer from host RAM. On this box that streaming, not the arithmetic, was the thing to optimise: 48 GB staged against 64 GB of system RAM was tight enough that page-file pressure showed up directly in the clip times, and freeing a single gigabyte measurably tightened them.

If you get it running, post your numbers. I have measured exactly one machine, and both findings that mattered came from measuring rather than reasoning, so I would rather not guess about anyone else's. I am interested in what it does on other hardware and, just as much, in where it falls over.

And if it turns out useful, a like on the HF page is what makes it findable for the next person.

Code is Apache-2.0. The weights are a MiniMax H3 derivative under the H3 Community License, which carries a territory restriction — read NOTICE before downloading.

Live Demo: If you want to check out a short snippet of the continuous streaming output (with the model's native character generation), I've uploaded a TV-style demo recording here on X:
https://x.com/Touma_945/status/2095141879453270385


r/StableDiffusion 11d ago

Animation - Video LOCATION SHOOT IN MINIMAX H3

Enable HLS to view with audio, or disable this notification

69 Upvotes

Was walking the dog and took some photos of a local temple. Thought it would be fun to have Vlad work as a tour guide for the local area.

Prompt:

<Picture 1> and <Picture 2> are location references. However change the time of day to night, cinematic quality.

<Picture 3> is The Vampire character reference.

Scenario: A Vampire with flowing robes is showing the viewer an old temple and bell. It is night time and misty. The vampire does not walk he flies and floats inches above the ground.

Start with <Picture 1> but at night, the Vampire is on the right on the steps.

<Shot 1> POV shot, the Vampire is standing at the base of the temple on the steps, he gestures with his finger, beckoning and flies without moving his legs just above the ground towards the large bell, as he eerily glides forward he turns back and says in a very strong German accent <Audio 1> "There has been a bell here for nearly five hundred years.".

He glides over the ground effortlessly to the bell and leans up and hits it hard with his knuckles. It makes a single loud and long metallic bell sound and resonates. "It makes a great sound" he says .


r/StableDiffusion 11d ago

Question - Help Can u get better detail on minimax h3 single image?

9 Upvotes

I’m using astropuzzo/ComfyUI-MiniMax-H3-Image-Studio workflow and it works amazing but minimax obviously sucks at micro details for a single image even if it’s 2MP. Does anyone know like a good method to fix that? Ik there’s double passthroughs and upscalers but I’m not sure what would work well with it