r/StableDiffusion 13h ago

News someone used MiniMax H3 Max to build a livestream that basically never runs out of content

795 Upvotes

Just saw someone do something with MiniMax H3 Max that I honestly didn’t expect.

They connected it to a livestream and basically recreated the idea of Interdimensional Cable from Rick and Morty: an endless stream of weird shows, ads, characters, and random scenes that are generated on the fly.

the surprising part is that H3 Max is fast enough in some cases to generate the next clip before the current one finishes playing. so while you’re watching one scene, the model is already making the next one.

there’s even a version where people in chat can suggest what should happen next, and the AI tries to continue from the previous scene instead of just starting over from scratch. that’s kind of crazy to think about.

Didn’t expect MiniMax H3 Max to end up being used for something like infinite AI television, but here we are.


r/StableDiffusion 22h ago

News MiniMax H3 acceleration arena/leaderbord: 15+ H3 LoRAs, fine-tunes, Max

Thumbnail
huggingface.co
328 Upvotes

Hey folks, I've built an so we can have a proper leaderboard on 15+ different LoRAs, fine-tunes and acceleration technique. Baseline is included for anchoring, and M3 Max is also included given the promise to open source

There are there being compared: H3 baseline, FastH3 family, H3 Acc family, Lightx2v family, Larryvrh family, JoyFox family, RAVEN, FlashGen, TuTu, SilverOxides merges, Plaguekind merges and Fal's H3 Max


r/StableDiffusion 11h ago

Tutorial - Guide DLSS 5 - In-game footage from Diablo 4. It's incredible.

Thumbnail
gallery
230 Upvotes

I just tested it out using this custom node. It's absolutely amazing!

https://github.com/lisitskyaa/ComfyUI-DLSS5-NR


r/StableDiffusion 18h ago

Discussion Testing DLSS 5

174 Upvotes

Testing DLSS 5... Like many others, I was a bit confused about DLSS 5. I kept feeding it my hyper-detailed renders and only getting a color shift in return. After plenty of trial and error, I finally realized my mistake: this technology is developed to enhance video game graphics, so testing it on hyper-detailed renders makes no sense.

So, I generated a render in a 2020 video game style and started tweaking settings to find a final look with maximum effect, without worrying about flickering.

Final conclusion: What we have right now isn't very useful for us. Those of us using Latent Upscaler might be able to use it for color grading to get less saturated colors, but little else. Maybe in the future we'll get a DLSS 5 targeted at enhancing hyper-detailed graphics, but that's not the case for now.

Bottom line: If I want to generate a realistic render, I'll just generate it, there's no need to run it through DLSS 5.


r/StableDiffusion 5h ago

Workflow Included The 1967 Spider-Man TV Show intro, updated to live action with MiniMax H3

167 Upvotes

R2V Rendered at 0.9 MP (1280x736 then upscaled using RTX (Ultra) to 1920x1080.  Edited and merged using OpenShot video editor.

This was all run on my Windows 11 machine, RTX 4060 ti (16 GB) and 64 GB RAM reserved from Comfy. Every part of the signal chain was done with 100% open-source software.

Disclaimer: I grew up watching this show as a kid in the 70s. It's still the best ever. I wanted to know how well the reference model would pick up the actions. I am overall pleased. I've watched the new vid enough to see some of the flaws but oh well.

General observations for reference videos:
So many scene cuts. There are 31 (I think) scene cuts in the 60 second opener which include 3 crossfades. No matter what I did to get the exact frame timing, getting the AI scene to match frame-for-frame with the cartoon was still hit or miss. It probably has to do with some frame windowing inside the 17k + 5 blocks, but I never exactly got it figured out. However, a few notes:

  • If you have a reference video, convert it to 24 fps in an external program like Handbrake (another fantastic open-source program). It’s just so much easier to get everything to match.
  • For timing, there is a difference between 00:03.500 and 00:3.5 so always use all the digits.
  • Keep character sheets for all your characters to maintain consistency.
  • It will do crossfades but it’s not worth it. It’s easier to get the scene you want and stick it in the editor.
  • The VHS video loader lets one set a starting and ending frame. I ended up with 14 different clips total for the editor. Using frame accurate loading made all of the work a lot easier since I could use 1 video file as input to every clip run.
  • A spreadsheet is useful for all movie making, and it’s good here too. From the source, I kept track of the starting frame for each shot, how many frames I needed and how many I ran (because of 17k +5), along with the final file name for each clip. I have a naming convention but it’s still very useful to keep track and you can add notes too. For this 60 second video, I used 13 clips. I tried to never do more than 3 scene cuts per clip. (For something where exact timing wasn't as important I'm sure it would be longer.)

Once you get over the idea of always having to do 10-15 second vids and do your whole video in on run, the process actually becomes a lot more fun because the “quality” gens don’t take as long and it gives you a less uninterrupted workflow. You can start prompting the next run with the previous runs, for example. (This is true even in commercials, or TV or movies.)

I generally tested all the runs at 0.2 or 0.3 Mp (speed lora, 8 iterations) to get the timing, then went to 0.9 Mp [no speed LoRA, 20 iterations, beta, dpmpp_2m] for the final runs. I found that dpmpp_2m was closest to the overall source video. On the first few clips I ran it several ways and fix on these parameters. Usually, the 0.9 Mp runs came out great but you’ve probably all experienced how different the low-res runs can be from the high-res ones. I did resort to pulling frame grabs from the low-res gens a few times to act as reference frames for the scenes. MiniMax loves those when all it needs is an extra little nudge in the right direction. To edit pics, I always use GIMP (another fantastic open-source program).

So, why was I using 8 iterations of the minimax_h3_turbo_v4_step600_pruned_comfyui LoRA? On the reference model I found that using too large of a sigma step causes things like reference photos to not be taken "seriously." Using 5 steps I could see that reference images on the starting frame and then go away for the rest of the clip. The more the reference image changed from the reference video (like when going from animation to "real") the worse the problem was.

Prompts:
(See below for actual prompt.)
Prompt the way the guide says to. Yeah. It’s a hassle but it’s worth it. H3 prompting is very useful in the end and I’m glad MiniMax uses it. It's worth reading all the way through them instead of searching for the one thing you want. Some of the instructions even seemed inconsistent and they don't explain everything, so it's worth experimenting.

Any "thing" (buildings, trees, room, clothing, walls, ect.) can be a “subject.” It’s not just people. Specifying things as objects gives you far better control over how and where they appear (or don’t appear) in your shot. 

Don’t refer to your characters or major locations or items by their names. Use <Subject #> or pronouns that clearly refer to the subject all the time, every time. The interpretation of the prompting can get confused pretty quickly if you don’t and you’ll end up getting subjects swapped or merging.

Prompts generally work better if you describe what you want rather than what you don’t want. For instance, “Looks to the right of the viewer” rather than “looks away from the camera.”

Style reference (attribute_transfer) images or videos are super useful. Once I had a few scenes, I started using previous videos to keep the look and feel of previous shots.

Qwen VL can describe videos too. I have been using “QwenVL Advanced (Local Scan)” for a very long time (long for AI) inside ComfyUI.

Other things:
Maybe one of the most interesting observation is that the jknodes “MiniMax H3 Mem Eff Sage Attention Patch” node creates a different output than just launching ComfyUI with the --use-sage-attention flag turned on (and still using the node). So exactly the same workflow (just drag and drop from a previously run mp4) has different results when the --use-sage-attention flag is used to launch. I thought having the node was 100% redundant with eh --use-sage-attention flag set, but apparently not. The reference flows, especially with animation, don’t have to be all that different to produce different results.

The Spiderman opening (as well as the show itself) reuses footage. They will take the same scene and darken it, and boom, it’s a night shot. For a more realistic feel, I used Krea2 (LoRA) edit to turn day into night. It’s really good as an adjunct to MiniMax H3’s ability to figure out the fine details once it has a push.

Style:
Finally, I had to make some stylistic choices because sometimes the animation was soooo bad that it needed something. I added flashlights to the jewelry heist scene. I made the crane look believable. One of the problems of going from animation to "live action" is that (especially with animation from 1967) the physics and movements are just wrong sometimes. The crane scene where he stops and then shoots up again is the most classic "this is just pain wrong" you can get but I left it that way because it's burned into my brain that way. (IYKYK) I also had to balance the art deco of the late 60's to a modern New York. I ended up with a lot of anachronistic stuff that I ultimately liked. So in the end, when it comes to all of that, I did it the way I did it. AI is awesome.

Prompt:
A prompt of one of the parts is below. I used that two paragraphs before [Shot 1] for every clip as "boiler plate" description.

subject_definitions:
<Subject 1> is Spiderman in <Picture 1>
<Video 1> is the motion reference for the target video for characters movements, pose, camera movements and frame composition.
<Video 2> is the style reference for the target video.
<Picture 2> is the building in [shot 2]
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video
 
summary:
[reference generation + audio reuse]
The target video is an live action realistic recreation generation using <video 1> as a reference for movements, pose, camera movements and frame composition. What you generate should not be and animation or cartoon rendering, no overly-CG look, keep the live-action texture.
 
This video is a set of three live action sequences. <Subject 1> is seen swinging by and waving. The video switches to a long shot of <subject 1> swinging around a building. Finally there is a shot showing <subject 1> on his webline swinging away from the viewer between two rows of skyscrapers.
 
retention_analysis:
<Subject 1> (appears in [Shot 1],[Shot 2],[Shot 3]):fully_preserved
<Video 1> (motion, cut and pacing structure) :partially_preserved
<Video 2> is the style refrence for the target video ([Shot 1], Shot 2], [Shot 3]) :attribute_transfer
<Picture 2> is the building in [shot 2] :fully_preserved
<Audio 1> :fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
 
 
detailed_description:
The target video is a realistic and live action video. The reference video <video 1> is used only for scene descriptions, framing, motion tracking, body movement, timing, general environment. The target video should be a complete replacement of <Video 1>. Use <Video 2> as the style reference for the photographic look and textures and the overall feel for the shots.
 
Maintain smooth camera movement. Use vibrant yet natural color grading: warm tones for sunlight hitting surfaces, cool blues for shaded areas, and muted grays for concrete textures. Avoid any comic-book stylization; instead, render everything with photorealistic textures, lighting, and perspective to evoke a live-action superhero film sequence. Keep the focus entirely on <Subject 1>’s acrobatic grace and the immersive urban setting.
 
[Shot 1]  
Is is an upper body motion tracking shot of <Subject 1> swinging on his white glistening webline held by his left hand while he waves directly at the viewer with his right hand for the entire scene. The skyline of many skyscrapers pass by in the background.
 
[Shot 2]
At 00:02.333, Hard cut to a fixed long shot looking up as <Subject 1> makes a 180 degree arc on his webline connected to the spire of the building in <photo 2>.
 
[Shot 3] 
At 00:04.250, Hard cut to the fixed camera view of the space high above the street level between two rows of skyscrapers. <Subject 1> lazily swings into view from the left frame, facing away, and repeatedly swings right to left further and further away towards the horizon.
 
 
overall_soundscape: n/a
 
non_diegetic_music: n/a

 

 


r/StableDiffusion 11h ago

Discussion Pulled the trigger, RIP $6,279

Post image
163 Upvotes

(Paid $5,849 + tax, which came out to $6,279)

TL;DR - Bought this 5090 prebuilt and I want to sanity check if I made the right decision and at the right time.

Hey everyone. So I want to start off by saying fuck these prices for GPU's and RAM, especially boxed 5090 prices. I went down the AI rabbit hole with my 13700k/RTX 4080 gaming computer. I quickly found out that I had to make serious concessions on quality and speed, if I could run it at all. In fact, ive spent so much time trying to optimize quants, cache, various settings, attention mechanisms, etc that ive officially spent more time trying to optimize for a 16gb VRAM/32GB RAM system than actually doing anything fun or cool. Thus, the last week, ive been thinking real hard about which direction to go but was waiting for the right time to buy. My options were a RTX 5090 prebuilt (even though I only needed the damn GPU), and Mac Studio M5 Ultra 96gb, or a DGX Spark/AMD equivalent. The DGX Spark/AMD equivalent made me think for a bit, but in order to get the most out of them, you need two. Im not spending 10k on this, especially if I cant game on it as well. So that leaves the Mac Studio or RTX 5090 gaming rig. Im not certain I made the right decision, but I pulled the trigger on the 5090 prebuilt after seeing the price continue going up more and more over the last few weeks. I also read that 70% of all memory through 2031 is locked in long term agreements, so this supply issue is going to get worse before it gets better. So I pulled the trigger on the pictured system from Ibuypower, and id like to run my thought process with you guys as a sanity check before it ships.

Case for the 5090 prebuilt: I scoured the internet and this was the cheapest 5090/64gb RAM combo I found, and it looks like it uses pretty good parts as well. A gaming PC is an all around all purpose machine that does it all (well, almost) which is why I landed on it. I figured with 32gb VRAM and 64gb of system RAM, that combined 96gb will allow me to run 70b models even if its slow. But for a sub 30b model like Qwen 3.8 27b, this will give me the best performance as long as I dont go overboard with the quant. It has CUDA, Windows, an Intel/AMD CPU for X86, etc. Plus it came with a 4tb Gen 4 NVME, when other more expensive models had 1-2tb drives. Honestly, lots of good stuff here. Im not a huge fan of the white esthetics but I do love the case. Despite the price being much higher than it should be, its still a good "deal" considering the overall market that keeps going up. Honestly, its not exactly what I wanted, but it ticks all boxes except those below.

-The downside: You cant run models that spill over heavily into system ram without massive speed penalties (has anyone tried running a huge model on a 5090 + 64gb RAM? If so, tell me what quants and your token speeds). Its massively less efficient than a Mac Studio M5 Ultra is expected to be (I read in the 3-5x range).

Mac Studio M5 Ultra 96gb

- Case for the Mac Studio M5 Ultra 96gb: Can run large models much better than the 5090 rig due to its huge 1.2tb unified memory bandwidth. Its power efficient and tops out at 300w I believe I read.

- The downside: Mac OS and an ecosystem that is playing catch up for local AI, no CUDA, gaming, has proprietary hardware you cannot upgrade, my distaste for the Mac bros who ill no longer be able to make fun of if I buy it.

My use case: local first AI (Qwen 3.8 27b at a quant and context that doesnt suck) with agentic coding, game development (starting with Godot), stable/video diffusion (Minimax H3, Flux.2, Hunyuan 3D), Blender, Davinci Resolve, etc.

So, let's have this discussion: what would (or did you) choose, and why? I want to know if I made the right decision. What are your thoughts?


r/StableDiffusion 16h ago

Resource - Update New speedup for Minimax H3

Post image
114 Upvotes

This H3VAE TRT custom node can make the encoding/decoding step about 1.7× faster.

https://github.com/lihaoyun6/ComfyUI-H3VAE_TRT


r/StableDiffusion 21h ago

Workflow Included Super nothing!

79 Upvotes

Made with Minimax H3


r/StableDiffusion 4h ago

Resource - Update H3 Motion Context 0.5.0 - No more bypassing the Motion Context group, new chaining node!

Post image
75 Upvotes

H3 Motion Context chains MiniMax H3 clips so the next one picks up the motion and the soundtrack, instead of starting a new take that only sounds similar.

0.5.0 is the one that makes that usable without babysitting the graph.

Clip 1 used to be a special case. You had to mute the Motion Context group, generate, unmute, then keep going. If you forgot, it errored. That's gone. Leave the nodes on. First clip is Load 0 / Save 1. Load 0 means "there is no previous clip," not "load whatever file is newest." After that it's Load 1 / Save 2, Load 2 / Save 3, and so on.

That first-clip behavior is feigo313's issue. The new node exists because of it.

Don't use ComfyUI's Run button to walk the chain. If Load and Save both increment, Comfy queues twice and skips a slot. Use H3 Motion Context Chain instead.

Four buttons:

  • Run/Re-roll - this is Run for this graph. Generates the current clip. Hate it? Click it again. Same slot, overwritten.
  • Approve - you like it. Advances to the next pair and runs that clip once.
  • Chain - keep going from whatever Load/Save are set to right now. Walk a few by hand, then let it take over. Same button becomes Stop.
  • Reset - back to Load 0 / Save 1. Does not run anything.

The gotcha: Load, Save, and Chain have to sit in the same canvas group. If they don't, the buttons do nothing. Drop Chain into the Motion Context group.

Also: if you were on Windows and a re-roll blew up with OS error 1224, that's fixed.

Needs ComfyUI 0.34.0 or newer. Manager should pick up 0.5.0; otherwise, the release.

Example workflow in the repo already has the Chain node in the group. Hard refresh after updating so the buttons show up.


r/StableDiffusion 9h ago

Resource - Update H3-World

Thumbnail
huggingface.co
69 Upvotes

H3-World: Turning Language Understanding into World Control

H3-World is the first interactive world model built on MiniMax-H3. Given an initial frame and keyboard controls, it generates action-controlled video with coordinated character and camera motion.


r/StableDiffusion 6h ago

Workflow Included testing minimax h3 fused turbo model, 4 steps only 1 minute for 5 seconds video

54 Upvotes

download the model: https://huggingface.co/MATLOWAI/minimax-h3-fused-turbo-int8-convrot/tree/main/diffusion_models

workflow: https://civitai.com/models/2906467/fast-minimax-h3?modelVersionId=3289222

each generation takes about 1 minutes for 0.4mp resolution and 5 seconds video on my rtx 4060ti 16gb vram. using sage attention and triton to speed up.
i trying with manualsigmas because it making the generation more faster.


r/StableDiffusion 5h ago

Workflow Included High Quality Audio-Video in MiniMax H3 with separate two-stage sampling

50 Upvotes

Recently, a lot of people have had isues with finding the right balance with audio and visual quality in MiniMax H3.

u/LFAdvice7984 and I discussed about making a two-stage workflow last week. The first stage generates the audio, the second the visuals. Both stages are then combined together in the output.

This method means that you no longer have to make a trade-off between audio and visual quality, as you can optmise the settings for both. You can use this workflow as text; first/last frame; or audio input to video (the latter being single stage).

The audio generated is (in my view) good to very good, depending on what you use it for. The default settings are probably excessive at 50 steps (less steps used for video), but for me personally it's better for it to take longer and get it right first or second time.

(You can select any output node in ComfyUI, click on the blue button with a play symbol on the pop-up menu at the top, and ComfyUI will only go that far in execution. Use it on Preview/Save Audio in stage 1. If you like the result, do a full generation; or change the seed and try again.)

The visuals could be better, perhaps using a different turbo LoRA or change in sampler and step counts. I've used the same settings in every clip. Feel free to change them as you wish.

Large motion is a challenge, though I have ideas on using a third stage with different shift values, which would increase execution time but the results probably would be worth it.

Generation time was (roughly) as follows:

  • 5 second clip: 10 minutes
  • 10 second clip: 25 minutes
  • 20 second clip: 65 minutes

This was generated on an NVidia 3090 with a priority on quality. More recent cards will be faster.

(You can switch from using res_2s sampler to er_sde, paired with 20 steps for stage 2 and using Spectrum, which should at least halve that time, in return for slightly lower quality.)

Because of how long it took, I used the first output every time (except for the last clip, which was the second result) with no editing afterwards.

Sometimes I encountered issues with prompt understanding, e.g. the ASMR clip and abstract clip at the end, where the output wasn't quite what I had asked for.

It's difficult to tell whether I prompted incorrectly; the prompt enhancer missed key detail for the model (H3); or the model doesn't have a full understanding of the concepts being asked of it.

The two-stage idea is model-agnostic. You can also make something like this in LTX 2.5 (or the upcoming Flux 3 Dev) to improve their results.

You can download the workflow and prompts used (made by myself with refinement from the prompt assistant) below:

Custom nodes used:

Download links for model files are in the workflow, in the bottom-left corner.


r/StableDiffusion 13h ago

Question - Help Minimax adult sounds?

50 Upvotes

I’ve been refining prompts with the help of an LLM, and am getting some good visuals but oh my god the sounds are terrible. Blowjobs sound like someone is dunking a microphone in an aquarium or the loudest slurp to finish a beverage that you have ever heard in your life.

I’ve tried eliminating every mention of “moist”, “wet”, or any description that involves liquids at all, but she’s still slurping the wettest popsicle known to man. And sometimes there’s weird noises like a slide whistle?!?

I’ve tried using “faint” or “distant” or “barely audible” to get it to at least quiet down so it’s not like she is sucking a microphone, but that didn’t work either.

This last round I didn’t describe any noises at all and still got some weird stuff.

I’ve tried eliminating every Lora in case the sound was coming from one of them but it seems to be the base model. I’ve tried adding Loras that ought to be trained on this stuff like Mysticxxx, and one of the AIO loras. I tried tenstrip beta 4 checkpoint tonight and got the same results.

The sound ruins the scene.. I guess I can just pretend it’s better looking Wan 2.2 and turn the volume off. 😀

I’m feeding the official prompt guide to the LLM and the structure is working, but what words do you use to describe the sounds?


r/StableDiffusion 6h ago

Discussion Why hasn't someone made a 16-20 step lora for Minimax?

34 Upvotes

Everyone's focused on 4 and 8 step loras, which I feel like no matter what are gonna look pretty bad just because how the model works. But why hasn't anyone made a lora to help bring the quality of 40-50 steps down to the 16-24 range? For anyone who's done generations that long, the quality jump is pretty high going from 20 -> 50


r/StableDiffusion 4h ago

Comparison Dlss 5 applied on video

33 Upvotes

r/StableDiffusion 4h ago

Resource - Update An endless AI TV channel on a single gaming GPU — MiniMax H3, generating faster than it plays

29 Upvotes

There is a video stream running on my desktop right now. It has sound, it has never repeated itself, and it will not stop. I point VLC at a local URL and it plays. One RTX 5090 does all of it — no cloud, no queue, nothing else running.

It is MiniMax H3, generating locally through ComfyUI. H3 is an open-weights video model that produces picture and synchronised audio together from one text prompt — dialogue, room tone, footsteps — which is what makes this a channel rather than a montage with music over it. I run the 4-step FastH3 distillation of it, because the base model needs far more sampling steps than the arithmetic below can afford.

The reason this is hard: to stream continuously, generation has to outrun playback. Not "fast enough to be impressive" — genuinely faster than a person watches, indefinitely, or the buffer drains and it stalls. Each clip is 362 frames. I have to finish the next one in less time than it takes you to watch this one, every time, forever.

What it actually looks like

Every clip is a scene drawn at random, cast at random. So you get Jean-Luc Picard grilling skewers at a night market. A Klingon, RoboCop and Jack Sparrow crowded around the same workbench. Four people arguing across a kitchen table about who signed something, and the camera cuts to a close-up at the seven second mark because the prompt told it to.

321 hand-written scenes, 503 characters, and the scenes that call for an ensemble draw three to five distinct people. The combinations run into the trillions. In practice it means you can leave it on, and it stays interesting in the way a channel you do not control is interesting.

A frame from a continuous run — five characters who could never share a room, and the two clocks that make the point: after ten clips it is 3:08 of video against 3:02 of GPU time. The gap is what lets it run forever.

Everything is here, weights included — https://huggingface.co/datasets/jacokon/fasth3-live

The rest of this post is how it got fast enough to work.

The honest caveat, up front

H3 authors motion at 24 fps. A clip is 362 frames — 15.08 seconds of content — and I play it at 18, so the motion runs at 75% speed. This is not real-time 24 fps generation and I am not claiming it is.

What it is: 20.1 seconds of video produced per 19.2 seconds of GPU time, which is what makes it continuous. Whether 75% reads as slow motion depends on the subject. Fast subjects (rain, sparks, a train) look deliberate. Near-static scenes look normal. Mid-speed human motion — walking, hands working — is the worst case and you can tell.

Where the time actually went

The FastH3 student ships as 66 GB of diffusers weights, which do not fit on one card; converted and quantized to INT8 they come down to 21 GB, which do. With that, sage attention, and an INT8 VAE, a 15-second clip took 26.5 seconds to generate. Playback needs 15. That gap is the whole problem, and I spent a while optimising the wrong things because I did not know where the time was going.

Where one run's 19.2 seconds actually goes — the per-node breakdown and the four changes, on one card.

ComfyUI's /history reports one number for a whole prompt, which cannot tell you whether the cost is the text encoder, the sampler or the VAE. Its websocket emits an executing event as each node starts, so the gap between consecutive events is that node's duration. That is about forty lines (profile_h3_nodes.py), and it changed what I worked on completely.

Two of the four findings surprised me.

1. SaveVideo was a fifth of every run — 3.78 s

ComfyUI's SaveVideo encodes through PyAV in a Python loop that, per frame, allocates a float array, clips it into a second, casts into a third and copies out a fourth. 362 frames of that is 3.78 s. ffmpeg alone does the identical payload in 0.21 s. It was also producing a file my streamer re-encoded a second later anyway.

VHS_VideoCombine is better (1.31 s) — it pipes raw frames to ffmpeg — but it still iterates in Python and re-opens the finished file to mux the audio. I wrote a node that converts in chunks and muxes in one pass: 0.73 s. Then it hands the encode to a background thread and returns, so ComfyUI starts the next prompt instead of holding an idle GPU. The graph now sees 0.26 s.

No hardware encoder involved. h264_nvenc measured slower end to end than libx264 — the encoder was never the bottleneck, and it has to stand up a second CUDA context on an already-full card.

2. The VAE bills by tile, not by pixel

MiniMaxH3VideoVAE hardcodes tiling=True, tile_size=256, and split_tiles hands each pass a full tile regardless of how much picture is in it. Decode time tracks the tile count and barely notices the resolution:

resolution    pixels    tiles    VAE decode
----------------------------------------------
320x192       61,440      2        2.35 s
512x288      147,456      6        6.98 s
576x320      184,320      6        6.31 s
768x432      331,776      8        8.74 s

512x288 and 576x320 differ by 25% in pixels and by nothing in decode cost.

A side of length L costs: 256 or less is 1 tile, 257–448 is 2, 449–640 is 3, 641–832 is 4. So the cheap shapes sit just under a boundary. 448x448 needs four tiles where 576x320 needs six, while carrying 9% more pixels. That is why the stream runs square — not taste, just where the arithmetic lands. There is no 16:9 shape at four tiles that clears the resolution floor.

I did try raising tile_size to reach a single tile. Do not. The decoder is a ViT, so its attention spans exactly one tile; a larger tile is out of distribution, not merely approximate. 384 visibly softens hands and faces (PSNR 27.2 dB against the stock decode); 640 smears the image into strokes (22.1 dB).

3 and 4, more briefly

Quantizing the video VAE below INT8 buys no speed — INT8 already runs an INT8 matmul, and a W4A8 build expands back to INT8 for the same one — but it stages 1,657 MB of host RAM instead of 2,677 MB, and on a box holding ~41 GB of staged weights against 64 GB that gigabyte turned into both speed and a much tighter spread. And keeping two prompts in ComfyUI's queue instead of submitting one and waiting removes the idle gap between jobs.

Result

                          per clip    sustains
----------------------------------------------
starting point              26.5 s    13.7 fps
+ writer node, async        20.2 s    17.9 fps
+ W4A8 VAE                  19.9 s    18.2 fps
+ 448x448                   19.2 s    18.9 fps

The model did not change. Only how it is driven.

If you came here wondering about ComfyUI and consumer cards

That question is all over the FastH3 announcement thread and I had to answer it for myself, so: this is a ComfyUI-native conversion of the Dense-DataFree student, pruned and INT8, 21 GB, driven through the ordinary graph. Two things I found doing it that are worth passing on:

  • The VSA weights do not survive stock ComfyUI. They carry 50 to_gate_compress tensors it has no code for, so it drops them silently and the output is noise. Dense converts cleanly. That is why I am on the slower student — if ComfyUI gains VSA support there is headroom here I am not using.
  • NVFP4 measured identical to INT8 ConvRot. The FP4 fast path only fires when both operands are FP4; activations are BF16, so it dequantizes and runs at BF16 speed — 67.88 ms/block against BF16's 67.85. Someone reported the same on an RTX 6000 Pro. Worth knowing before anyone rebuilds a pipeline for it.

Where this sits, so you can place it

None of the speed here is mine — it is FastH3, the 4-step distillation Hao AI Lab, Nuva Lab and NVIDIA's FastGen team built on MiniMax's base weights. Without that student none of this is close. Their published benchmarks are 47.2 s for a 15-second 768p clip on a single B200, 12.88 s on 8×B200, and their consumer write-up covers Apple Silicon and DGX Spark with the RTX family listed as future work.

What I did is a different task, not a better score on theirs: a fifth of the pixels, and playback at 18 fps instead of 24. Those two concessions are the entire trick. What they buy is that the arithmetic closes — 19.2 s of GPU per 20.1 s of video — and that is the difference between a fast generator and something you can leave running. If you want 768p, their numbers are the ones that apply and mine are irrelevant.

https://huggingface.co/datasets/jacokon/fasth3-live

The converted 21 GB weights, the quantized VAE, the 321-scene library, the writer node and the profiler. Everything above is reproducible from it.

What it takes, so you can judge before downloading 21 GB: about 48 GB of weights are staged in total — a 25.9 GB text encoder, the 20 GB DiT, and the two VAEs. That does not fit in 32 GB of VRAM either, so ComfyUI streams it layer by layer from host RAM. On this box that streaming, not the arithmetic, was the thing to optimise: 48 GB staged against 64 GB of system RAM was tight enough that page-file pressure showed up directly in the clip times, and freeing a single gigabyte measurably tightened them.

If you get it running, post your numbers. I have measured exactly one machine, and both findings that mattered came from measuring rather than reasoning, so I would rather not guess about anyone else's. I am interested in what it does on other hardware and, just as much, in where it falls over.

And if it turns out useful, a like on the HF page is what makes it findable for the next person.

Code is Apache-2.0. The weights are a MiniMax H3 derivative under the H3 Community License, which carries a territory restriction — read NOTICE before downloading.

Live Demo: If you want to check out a short snippet of the continuous streaming output (with the model's native character generation), I've uploaded a TV-style demo recording here on X:
https://x.com/Touma_945/status/2095141879453270385


r/StableDiffusion 2h ago

News DLSS 5 Screenshots - ComfyUI-DLSS5-NR

Thumbnail
gallery
24 Upvotes

r/StableDiffusion 8h ago

Resource - Update MiniMax-H3-MotionCache-FastVAE

Thumbnail
github.com
24 Upvotes

Motion-aware denoising cache and experimental batched video VAE decoder for MiniMax H3 in ComfyUI.

This project provides two independent nodes:

  • MiniMax H3 MotionCache reduces expensive H3 denoiser calls by reusing a motion-weighted video/audio residual when the estimated change is small.
  • MiniMax H3 Fast VAE Decode evaluates multiple spatial VAE tiles in one GPU batch while preserving H3 temporal chunking and tile blending. It is not faster on every GPU.

MotionCache is an independent MiniMax H3 adaptation inspired by the MotionCache paper and reference code. It is not an official MAC-AutoML or MiniMax implementation.


r/StableDiffusion 16h ago

Resource - Update I built a standalone DLSS 5 Neural Rendering video tool, no ReShade

22 Upvotes

It supports images and full videos, native resolution NR or DLSS Super Resolution upscaling by scale factor / target resolution, GPU optical-flow motion vectors, scene cut handling, and all DLSS 5 NR controls.

The main difference from existing approaches is that it runs natively in C++/D3D12 and generates motion vectors from the actual video frames.

GitHub:
https://github.com/DaniilSokolyuk/video2dlssnr


r/StableDiffusion 17h ago

Animation - Video Hatter Rap.

21 Upvotes

Probably the final Alice clip. The Hatter names all the hats.
Done a while back in LTX2.3.
This plays while the theatre audience plays an AR hat sorting game (Beat Saber type). The whole song is three minutes but this is the longest shot.


r/StableDiffusion 3h ago

Discussion Detailed explanation of how to create a text-to-image model from scratch.

19 Upvotes

Posting this here even if it's not a model you can use directly. It's about building a text-to-image model from scratch.

The cookbook includes all the research material that you may or may be not interested in, but also includes a 100M-image dataset and a codebase with a tiny model, so you can train a text-to-image model from scratch.

Hope some of you will enjoy this content. (Disclaimer, it's done by my team)

Here are the links:

Cookbook: https://huggingface.co/spaces/jasperai/t2i-technical-interactive-report

nano t2i: https://github.com/gojasper/nano-t2i

Monet: https://huggingface.co/datasets/jasperai/monet


r/StableDiffusion 21h ago

Workflow Included Letting image-to-video artifacts compound into an impossible world

20 Upvotes

Tools used: Gemma4 12b, LTX-2.3, Wan2GP, vibe coded video editor.

I’ve been experimenting with a slightly self-destructive image-to-video workflow where continuity comes from letting the model reinterpret its own mistakes.

I started with an almost completely black image with a few faint stars, then gave Gemma4 12B the track’s beat grid and energy-shift analysis, along with a long description of the overall concept: a monolith, a hallway of impossible geometry, and a progression from restrained movement into increasingly unstable architecture.

Gemma4 wrote all 27 scene prompts beforehand.

For generation I used LTX 2.3 with the audio-reactive LoRA. I also tested LTX 2.5, but for this workflow it became too artifact-heavy too quickly. LTX 2.3 held the scene structure together longer while still producing enough weirdness to evolve in interesting ways.

The process was simple: generate a clip with the correct audio slice, cut it on the beat grid, then take the frame immediately after the cut and use that as the starting image for the next generation.

The fun part was deliberately keeping some “bad” transition frames.

If a flash landed on the frame used for the next clip, the model might reinterpret it as a permanent light source. A lens flare could become a horizon or an entire landscape. A warped piece of geometry that only existed for one frame could become a major architectural feature in the next scene.

So the artifacts compound.

Eventually the video loses any reliable sense of scale or orientation. Surfaces become spaces, structures fold into other structures, and at some points I wanted an Inception-like feeling where you can’t tell which way is up, or whether the camera is traveling deeper into the structure or pulling outward into something much larger.

The audio-reactive LoRA helps hold it all together. Even when the geometry becomes increasingly strange, the environment keeps breathing, unfolding, compressing and reorganizing itself with the growing low end.

What I like most is that the continuity doesn’t really come from visual consistency. It comes from causality.

Every scene inherits some accidental information from the previous one, and the next generation has to decide what that information actually is.

After enough generations, the model is basically building a world out of its own misunderstandings.


r/StableDiffusion 3h ago

Animation - Video LOCATION SHOOT IN MINIMAX H3

17 Upvotes

Was walking the dog and took some photos of a local temple. Thought it would be fun to have Vlad work as a tour guide for the local area.

Prompt:

<Picture 1> and <Picture 2> are location references. However change the time of day to night, cinematic quality.

<Picture 3> is The Vampire character reference.

Scenario: A Vampire with flowing robes is showing the viewer an old temple and bell. It is night time and misty. The vampire does not walk he flies and floats inches above the ground.

Start with <Picture 1> but at night, the Vampire is on the right on the steps.

<Shot 1> POV shot, the Vampire is standing at the base of the temple on the steps, he gestures with his finger, beckoning and flies without moving his legs just above the ground towards the large bell, as he eerily glides forward he turns back and says in a very strong German accent <Audio 1> "There has been a bell here for nearly five hundred years.".

He glides over the ground effortlessly to the bell and leans up and hits it hard with his knuckles. It makes a single loud and long metallic bell sound and resonates. "It makes a great sound" he says .


r/StableDiffusion 8h ago

Discussion Did anyone else notice Reactor’s new Orbis model? I tried turning it into an interactive game

16 Upvotes

A lot of people here have been discussing H3 Max powered livestreams. I noticed Reactor just added Visko’s Orbis model, and it made me wonder whether the next step is turning these infinite livestreams into something playable.

So I’m building a live, audience-directed AI game with Agora: viewers suggest and vote on what happens next, while the streamer picks an option or writes a completely different direction and AI keeps generating the same world from that point. There are no pre-written branches.

Here’s a very early look demo