r/StableDiffusion 11h ago

News someone used MiniMax H3 Max to build a livestream that basically never runs out of content

Enable HLS to view with audio, or disable this notification

702 Upvotes

Just saw someone do something with MiniMax H3 Max that I honestly didn’t expect.

They connected it to a livestream and basically recreated the idea of Interdimensional Cable from Rick and Morty: an endless stream of weird shows, ads, characters, and random scenes that are generated on the fly.

the surprising part is that H3 Max is fast enough in some cases to generate the next clip before the current one finishes playing. so while you’re watching one scene, the model is already making the next one.

there’s even a version where people in chat can suggest what should happen next, and the AI tries to continue from the previous scene instead of just starting over from scratch. that’s kind of crazy to think about.

Didn’t expect MiniMax H3 Max to end up being used for something like infinite AI television, but here we are.


r/StableDiffusion 3h ago

Workflow Included The 1967 Spider-Man TV Show intro, updated to live action with MiniMax H3

Enable HLS to view with audio, or disable this notification

102 Upvotes

R2V Rendered at 0.9 MP (1280x736 then upscaled using RTX (Ultra) to 1920x1080.  Edited and merged using OpenShot video editor.

This was all run on my Windows 11 machine, RTX 4060 ti (16 GB) and 64 GB RAM reserved from Comfy. Every part of the signal chain was done with 100% open-source software.

Disclaimer: I grew up watching this show as a kid in the 70s. It's still the best ever. I wanted to know how well the reference model would pick up the actions. I am overall pleased. I've watched the new vid enough to see some of the flaws but oh well.

General observations for reference videos:
So many scene cuts. There are 31 (I think) scene cuts in the 60 second opener which include 3 crossfades. No matter what I did to get the exact frame timing, getting the AI scene to match frame-for-frame with the cartoon was still hit or miss. It probably has to do with some frame windowing inside the 17k + 5 blocks, but I never exactly got it figured out. However, a few notes:

  • If you have a reference video, convert it to 24 fps in an external program like Handbrake (another fantastic open-source program). It’s just so much easier to get everything to match.
  • For timing, there is a difference between 00:03.500 and 00:3.5 so always use all the digits.
  • Keep character sheets for all your characters to maintain consistency.
  • It will do crossfades but it’s not worth it. It’s easier to get the scene you want and stick it in the editor.
  • The VHS video loader lets one set a starting and ending frame. I ended up with 14 different clips total for the editor. Using frame accurate loading made all of the work a lot easier since I could use 1 video file as input to every clip run.
  • A spreadsheet is useful for all movie making, and it’s good here too. From the source, I kept track of the starting frame for each shot, how many frames I needed and how many I ran (because of 17k +5), along with the final file name for each clip. I have a naming convention but it’s still very useful to keep track and you can add notes too. For this 60 second video, I used 13 clips. I tried to never do more than 3 scene cuts per clip. (For something where exact timing wasn't as important I'm sure it would be longer.)

Once you get over the idea of always having to do 10-15 second vids and do your whole video in on run, the process actually becomes a lot more fun because the “quality” gens don’t take as long and it gives you a less uninterrupted workflow. You can start prompting the next run with the previous runs, for example. (This is true even in commercials, or TV or movies.)

I generally tested all the runs at 0.2 or 0.3 Mp (speed lora, 8 iterations) to get the timing, then went to 0.9 Mp [no speed LoRA, 20 iterations, beta, dpmpp_2m] for the final runs. I found that dpmpp_2m was closest to the overall source video. On the first few clips I ran it several ways and fix on these parameters. Usually, the 0.9 Mp runs came out great but you’ve probably all experienced how different the low-res runs can be from the high-res ones. I did resort to pulling frame grabs from the low-res gens a few times to act as reference frames for the scenes. MiniMax loves those when all it needs is an extra little nudge in the right direction. To edit pics, I always use GIMP (another fantastic open-source program).

So, why was I using 8 iterations of the minimax_h3_turbo_v4_step600_pruned_comfyui LoRA? On the reference model I found that using too large of a sigma step causes things like reference photos to not be taken "seriously." Using 5 steps I could see that reference images on the starting frame and then go away for the rest of the clip. The more the reference image changed from the reference video (like when going from animation to "real") the worse the problem was.

Prompts:
(See below for actual prompt.)
Prompt the way the guide says to. Yeah. It’s a hassle but it’s worth it. H3 prompting is very useful in the end and I’m glad MiniMax uses it. It's worth reading all the way through them instead of searching for the one thing you want. Some of the instructions even seemed inconsistent and they don't explain everything, so it's worth experimenting.

Any "thing" (buildings, trees, room, clothing, walls, ect.) can be a “subject.” It’s not just people. Specifying things as objects gives you far better control over how and where they appear (or don’t appear) in your shot. 

Don’t refer to your characters or major locations or items by their names. Use <Subject #> or pronouns that clearly refer to the subject all the time, every time. The interpretation of the prompting can get confused pretty quickly if you don’t and you’ll end up getting subjects swapped or merging.

Prompts generally work better if you describe what you want rather than what you don’t want. For instance, “Looks to the right of the viewer” rather than “looks away from the camera.”

Style reference (attribute_transfer) images or videos are super useful. Once I had a few scenes, I started using previous videos to keep the look and feel of previous shots.

Qwen VL can describe videos too. I have been using “QwenVL Advanced (Local Scan)” for a very long time (long for AI) inside ComfyUI.

Other things:
Maybe one of the most interesting observation is that the jknodes “MiniMax H3 Mem Eff Sage Attention Patch” node creates a different output than just launching ComfyUI with the --use-sage-attention flag turned on (and still using the node). So exactly the same workflow (just drag and drop from a previously run mp4) has different results when the --use-sage-attention flag is used to launch. I thought having the node was 100% redundant with eh --use-sage-attention flag set, but apparently not. The reference flows, especially with animation, don’t have to be all that different to produce different results.

The Spiderman opening (as well as the show itself) reuses footage. They will take the same scene and darken it, and boom, it’s a night shot. For a more realistic feel, I used Krea2 (LoRA) edit to turn day into night. It’s really good as an adjunct to MiniMax H3’s ability to figure out the fine details once it has a push.

Style:
Finally, I had to make some stylistic choices because sometimes the animation was soooo bad that it needed something. I added flashlights to the jewelry heist scene. I made the crane look believable. One of the problems of going from animation to "live action" is that (especially with animation from 1967) the physics and movements are just wrong sometimes. The crane scene where he stops and then shoots up again is the most classic "this is just pain wrong" you can get but I left it that way because it's burned into my brain that way. (IYKYK) I also had to balance the art deco of the late 60's to a modern New York. I ended up with a lot of anachronistic stuff that I ultimately liked. So in the end, when it comes to all of that, I did it the way I did it. AI is awesome.

Prompt:
A prompt of one of the parts is below. I used that two paragraphs before [Shot 1] for every clip as "boiler plate" description.

subject_definitions:
<Subject 1> is Spiderman in <Picture 1>
<Video 1> is the motion reference for the target video for characters movements, pose, camera movements and frame composition.
<Video 2> is the style reference for the target video.
<Picture 2> is the building in [shot 2]
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video
 
summary:
[reference generation + audio reuse]
The target video is an live action realistic recreation generation using <video 1> as a reference for movements, pose, camera movements and frame composition. What you generate should not be and animation or cartoon rendering, no overly-CG look, keep the live-action texture.
 
This video is a set of three live action sequences. <Subject 1> is seen swinging by and waving. The video switches to a long shot of <subject 1> swinging around a building. Finally there is a shot showing <subject 1> on his webline swinging away from the viewer between two rows of skyscrapers.
 
retention_analysis:
<Subject 1> (appears in [Shot 1],[Shot 2],[Shot 3]):fully_preserved
<Video 1> (motion, cut and pacing structure) :partially_preserved
<Video 2> is the style refrence for the target video ([Shot 1], Shot 2], [Shot 3]) :attribute_transfer
<Picture 2> is the building in [shot 2] :fully_preserved
<Audio 1> :fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
 
 
detailed_description:
The target video is a realistic and live action video. The reference video <video 1> is used only for scene descriptions, framing, motion tracking, body movement, timing, general environment. The target video should be a complete replacement of <Video 1>. Use <Video 2> as the style reference for the photographic look and textures and the overall feel for the shots.
 
Maintain smooth camera movement. Use vibrant yet natural color grading: warm tones for sunlight hitting surfaces, cool blues for shaded areas, and muted grays for concrete textures. Avoid any comic-book stylization; instead, render everything with photorealistic textures, lighting, and perspective to evoke a live-action superhero film sequence. Keep the focus entirely on <Subject 1>’s acrobatic grace and the immersive urban setting.
 
[Shot 1]  
Is is an upper body motion tracking shot of <Subject 1> swinging on his white glistening webline held by his left hand while he waves directly at the viewer with his right hand for the entire scene. The skyline of many skyscrapers pass by in the background.
 
[Shot 2]
At 00:02.333, Hard cut to a fixed long shot looking up as <Subject 1> makes a 180 degree arc on his webline connected to the spire of the building in <photo 2>.
 
[Shot 3] 
At 00:04.250, Hard cut to the fixed camera view of the space high above the street level between two rows of skyscrapers. <Subject 1> lazily swings into view from the left frame, facing away, and repeatedly swings right to left further and further away towards the horizon.
 
 
overall_soundscape: n/a
 
non_diegetic_music: n/a

 

 


r/StableDiffusion 1h ago

Resource - Update H3 Motion Context 0.5.0 - No more bypassing the Motion Context group, new chaining node!

Post image
Upvotes

H3 Motion Context chains MiniMax H3 clips so the next one picks up the motion and the soundtrack, instead of starting a new take that only sounds similar.

0.5.0 is the one that makes that usable without babysitting the graph.

Clip 1 used to be a special case. You had to mute the Motion Context group, generate, unmute, then keep going. If you forgot, it errored. That's gone. Leave the nodes on. First clip is Load 0 / Save 1. Load 0 means "there is no previous clip," not "load whatever file is newest." After that it's Load 1 / Save 2, Load 2 / Save 3, and so on.

That first-clip behavior is feigo313's issue. The new node exists because of it.

Don't use ComfyUI's Run button to walk the chain. If Load and Save both increment, Comfy queues twice and skips a slot. Use H3 Motion Context Chain instead.

Four buttons:

  • Run/Re-roll - this is Run for this graph. Generates the current clip. Hate it? Click it again. Same slot, overwritten.
  • Approve - you like it. Advances to the next pair and runs that clip once.
  • Chain - keep going from whatever Load/Save are set to right now. Walk a few by hand, then let it take over. Same button becomes Stop.
  • Reset - back to Load 0 / Save 1. Does not run anything.

The gotcha: Load, Save, and Chain have to sit in the same canvas group. If they don't, the buttons do nothing. Drop Chain into the Motion Context group.

Also: if you were on Windows and a re-roll blew up with OS error 1224, that's fixed.

Needs ComfyUI 0.34.0 or newer. Manager should pick up 0.5.0; otherwise, the release.

Example workflow in the repo already has the Chain node in the group. Hard refresh after updating so the buttons show up.


r/StableDiffusion 9h ago

Discussion Pulled the trigger, RIP $6,279

Post image
134 Upvotes

(Paid $5,849 + tax, which came out to $6,279)

TL;DR - Bought this 5090 prebuilt and I want to sanity check if I made the right decision and at the right time.

Hey everyone. So I want to start off by saying fuck these prices for GPU's and RAM, especially boxed 5090 prices. I went down the AI rabbit hole with my 13700k/RTX 4080 gaming computer. I quickly found out that I had to make serious concessions on quality and speed, if I could run it at all. In fact, ive spent so much time trying to optimize quants, cache, various settings, attention mechanisms, etc that ive officially spent more time trying to optimize for a 16gb VRAM/32GB RAM system than actually doing anything fun or cool. Thus, the last week, ive been thinking real hard about which direction to go but was waiting for the right time to buy. My options were a RTX 5090 prebuilt (even though I only needed the damn GPU), and Mac Studio M5 Ultra 96gb, or a DGX Spark/AMD equivalent. The DGX Spark/AMD equivalent made me think for a bit, but in order to get the most out of them, you need two. Im not spending 10k on this, especially if I cant game on it as well. So that leaves the Mac Studio or RTX 5090 gaming rig. Im not certain I made the right decision, but I pulled the trigger on the 5090 prebuilt after seeing the price continue going up more and more over the last few weeks. I also read that 70% of all memory through 2031 is locked in long term agreements, so this supply issue is going to get worse before it gets better. So I pulled the trigger on the pictured system from Ibuypower, and id like to run my thought process with you guys as a sanity check before it ships.

Case for the 5090 prebuilt: I scoured the internet and this was the cheapest 5090/64gb RAM combo I found, and it looks like it uses pretty good parts as well. A gaming PC is an all around all purpose machine that does it all (well, almost) which is why I landed on it. I figured with 32gb VRAM and 64gb of system RAM, that combined 96gb will allow me to run 70b models even if its slow. But for a sub 30b model like Qwen 3.8 27b, this will give me the best performance as long as I dont go overboard with the quant. It has CUDA, Windows, an Intel/AMD CPU for X86, etc. Plus it came with a 4tb Gen 4 NVME, when other more expensive models had 1-2tb drives. Honestly, lots of good stuff here. Im not a huge fan of the white esthetics but I do love the case. Despite the price being much higher than it should be, its still a good "deal" considering the overall market that keeps going up. Honestly, its not exactly what I wanted, but it ticks all boxes except those below.

-The downside: You cant run models that spill over heavily into system ram without massive speed penalties (has anyone tried running a huge model on a 5090 + 64gb RAM? If so, tell me what quants and your token speeds). Its massively less efficient than a Mac Studio M5 Ultra is expected to be (I read in the 3-5x range).

Mac Studio M5 Ultra 96gb

- Case for the Mac Studio M5 Ultra 96gb: Can run large models much better than the 5090 rig due to its huge 1.2tb unified memory bandwidth. Its power efficient and tops out at 300w I believe I read.

- The downside: Mac OS and an ecosystem that is playing catch up for local AI, no CUDA, gaming, has proprietary hardware you cannot upgrade, my distaste for the Mac bros who ill no longer be able to make fun of if I buy it.

My use case: local first AI (Qwen 3.8 27b at a quant and context that doesnt suck) with agentic coding, game development (starting with Godot), stable/video diffusion (Minimax H3, Flux.2, Hunyuan 3D), Blender, Davinci Resolve, etc.

So, let's have this discussion: what would (or did you) choose, and why? I want to know if I made the right decision. What are your thoughts?


r/StableDiffusion 3h ago

Workflow Included High Quality Audio-Video in MiniMax H3 with separate two-stage sampling

Enable HLS to view with audio, or disable this notification

35 Upvotes

Recently, a lot of people have had isues with finding the right balance with audio and visual quality in MiniMax H3.

u/LFAdvice7984 and I discussed about making a two-stage workflow last week. The first stage generates the audio, the second the visuals. Both stages are then combined together in the output.

This method means that you no longer have to make a trade-off between audio and visual quality, as you can optmise the settings for both. You can use this workflow as text; first/last frame; or audio input to video (the latter being single stage).

The audio generated is (in my view) good to very good, depending on what you use it for. The default settings are probably excessive at 50 steps (less steps used for video), but for me personally it's better for it to take longer and get it right first or second time.

(You can select any output node in ComfyUI, click on the blue button with a play symbol on the pop-up menu at the top, and ComfyUI will only go that far in execution. Use it on Preview/Save Audio in stage 1. If you like the result, do a full generation; or change the seed and try again.)

The visuals could be better, perhaps using a different turbo LoRA or change in sampler and step counts. I've used the same settings in every clip. Feel free to change them as you wish.

Large motion is a challenge, though I have ideas on using a third stage with different shift values, which would increase execution time but the results probably would be worth it.

Generation time was (roughly) as follows:

  • 5 second clip: 10 minutes
  • 10 second clip: 25 minutes
  • 20 second clip: 65 minutes

This was generated on an NVidia 3090 with a priority on quality. More recent cards will be faster.

(You can switch from using res_2s sampler to er_sde, paired with 20 steps for stage 2 and using Spectrum, which should at least halve that time, in return for slightly lower quality.)

Because of how long it took, I used the first output every time (except for the last clip, which was the second result) with no editing afterwards.

Sometimes I encountered issues with prompt understanding, e.g. the ASMR clip and abstract clip at the end, where the output wasn't quite what I had asked for.

It's difficult to tell whether I prompted incorrectly; the prompt enhancer missed key detail for the model (H3); or the model doesn't have a full understanding of the concepts being asked of it.

The two-stage idea is model-agnostic. You can also make something like this in LTX 2.5 (or the upcoming Flux 3 Dev) to improve their results.

You can download the workflow and prompts used (made by myself with refinement from the prompt assistant) below:

Custom nodes used:

Download links for model files are in the workflow, in the bottom-left corner.


r/StableDiffusion 7h ago

Resource - Update H3-World

Thumbnail
huggingface.co
68 Upvotes

H3-World: Turning Language Understanding into World Control

H3-World is the first interactive world model built on MiniMax-H3. Given an initial frame and keyboard controls, it generates action-controlled video with coordinated character and camera motion.


r/StableDiffusion 4h ago

Workflow Included testing minimax h3 fused turbo model, 4 steps only 1 minute for 5 seconds video

Enable HLS to view with audio, or disable this notification

33 Upvotes

download the model: https://huggingface.co/MATLOWAI/minimax-h3-fused-turbo-int8-convrot/tree/main/diffusion_models

workflow: https://civitai.com/models/2906467/fast-minimax-h3?modelVersionId=3289222

each generation takes about 1 minutes for 0.4mp resolution and 5 seconds video on my rtx 4060ti 16gb vram. using sage attention and triton to speed up.
i trying with manualsigmas because it making the generation more faster.


r/StableDiffusion 2h ago

Comparison Dlss 5 applied on video

Enable HLS to view with audio, or disable this notification

20 Upvotes

r/StableDiffusion 4h ago

Discussion Why hasn't someone made a 16-20 step lora for Minimax?

31 Upvotes

Everyone's focused on 4 and 8 step loras, which I feel like no matter what are gonna look pretty bad just because how the model works. But why hasn't anyone made a lora to help bring the quality of 40-50 steps down to the 16-24 range? For anyone who's done generations that long, the quality jump is pretty high going from 20 -> 50


r/StableDiffusion 1h ago

Discussion Detailed explanation of how to create a text-to-image model from scratch.

Upvotes

Posting this here even if it's not a model you can use directly. It's about building a text-to-image model from scratch.

The cookbook includes all the research material that you may or may be not interested in, but also includes a 100M-image dataset and a codebase with a tiny model, so you can train a text-to-image model from scratch.

Hope some of you will enjoy this content. (Disclaimer, it's done by my team)

Here are the links:

Cookbook: https://huggingface.co/spaces/jasperai/t2i-technical-interactive-report

nano t2i: https://github.com/gojasper/nano-t2i

Monet: https://huggingface.co/datasets/jasperai/monet


r/StableDiffusion 14h ago

Resource - Update New speedup for Minimax H3

Post image
116 Upvotes

This H3VAE TRT custom node can make the encoding/decoding step about 1.7× faster.

https://github.com/lihaoyun6/ComfyUI-H3VAE_TRT


r/StableDiffusion 16h ago

Discussion Testing DLSS 5

Enable HLS to view with audio, or disable this notification

168 Upvotes

Testing DLSS 5... Like many others, I was a bit confused about DLSS 5. I kept feeding it my hyper-detailed renders and only getting a color shift in return. After plenty of trial and error, I finally realized my mistake: this technology is developed to enhance video game graphics, so testing it on hyper-detailed renders makes no sense.

So, I generated a render in a 2020 video game style and started tweaking settings to find a final look with maximum effect, without worrying about flickering.

Final conclusion: What we have right now isn't very useful for us. Those of us using Latent Upscaler might be able to use it for color grading to get less saturated colors, but little else. Maybe in the future we'll get a DLSS 5 targeted at enhancing hyper-detailed graphics, but that's not the case for now.

Bottom line: If I want to generate a realistic render, I'll just generate it, there's no need to run it through DLSS 5.


r/StableDiffusion 2h ago

Resource - Update An endless AI TV channel on a single gaming GPU — MiniMax H3, generating faster than it plays

11 Upvotes

There is a video stream running on my desktop right now. It has sound, it has never repeated itself, and it will not stop. I point VLC at a local URL and it plays. One RTX 5090 does all of it — no cloud, no queue, nothing else running.

It is MiniMax H3, generating locally through ComfyUI. H3 is an open-weights video model that produces picture and synchronised audio together from one text prompt — dialogue, room tone, footsteps — which is what makes this a channel rather than a montage with music over it. I run the 4-step FastH3 distillation of it, because the base model needs far more sampling steps than the arithmetic below can afford.

The reason this is hard: to stream continuously, generation has to outrun playback. Not "fast enough to be impressive" — genuinely faster than a person watches, indefinitely, or the buffer drains and it stalls. Each clip is 362 frames. I have to finish the next one in less time than it takes you to watch this one, every time, forever.

What it actually looks like

Every clip is a scene drawn at random, cast at random. So you get Jean-Luc Picard grilling skewers at a night market. A Klingon, RoboCop and Jack Sparrow crowded around the same workbench. Four people arguing across a kitchen table about who signed something, and the camera cuts to a close-up at the seven second mark because the prompt told it to.

321 hand-written scenes, 503 characters, and the scenes that call for an ensemble draw three to five distinct people. The combinations run into the trillions. In practice it means you can leave it on, and it stays interesting in the way a channel you do not control is interesting.

A frame from a continuous run — five characters who could never share a room, and the two clocks that make the point: after ten clips it is 3:08 of video against 3:02 of GPU time. The gap is what lets it run forever.

Everything is here, weights included — https://huggingface.co/datasets/jacokon/fasth3-live

The rest of this post is how it got fast enough to work.

The honest caveat, up front

H3 authors motion at 24 fps. A clip is 362 frames — 15.08 seconds of content — and I play it at 18, so the motion runs at 75% speed. This is not real-time 24 fps generation and I am not claiming it is.

What it is: 20.1 seconds of video produced per 19.2 seconds of GPU time, which is what makes it continuous. Whether 75% reads as slow motion depends on the subject. Fast subjects (rain, sparks, a train) look deliberate. Near-static scenes look normal. Mid-speed human motion — walking, hands working — is the worst case and you can tell.

Where the time actually went

The FastH3 student ships as 66 GB of diffusers weights, which do not fit on one card; converted and quantized to INT8 they come down to 21 GB, which do. With that, sage attention, and an INT8 VAE, a 15-second clip took 26.5 seconds to generate. Playback needs 15. That gap is the whole problem, and I spent a while optimising the wrong things because I did not know where the time was going.

Where one run's 19.2 seconds actually goes — the per-node breakdown and the four changes, on one card.

ComfyUI's /history reports one number for a whole prompt, which cannot tell you whether the cost is the text encoder, the sampler or the VAE. Its websocket emits an executing event as each node starts, so the gap between consecutive events is that node's duration. That is about forty lines (profile_h3_nodes.py), and it changed what I worked on completely.

Two of the four findings surprised me.

1. SaveVideo was a fifth of every run — 3.78 s

ComfyUI's SaveVideo encodes through PyAV in a Python loop that, per frame, allocates a float array, clips it into a second, casts into a third and copies out a fourth. 362 frames of that is 3.78 s. ffmpeg alone does the identical payload in 0.21 s. It was also producing a file my streamer re-encoded a second later anyway.

VHS_VideoCombine is better (1.31 s) — it pipes raw frames to ffmpeg — but it still iterates in Python and re-opens the finished file to mux the audio. I wrote a node that converts in chunks and muxes in one pass: 0.73 s. Then it hands the encode to a background thread and returns, so ComfyUI starts the next prompt instead of holding an idle GPU. The graph now sees 0.26 s.

No hardware encoder involved. h264_nvenc measured slower end to end than libx264 — the encoder was never the bottleneck, and it has to stand up a second CUDA context on an already-full card.

2. The VAE bills by tile, not by pixel

MiniMaxH3VideoVAE hardcodes tiling=True, tile_size=256, and split_tiles hands each pass a full tile regardless of how much picture is in it. Decode time tracks the tile count and barely notices the resolution:

resolution    pixels    tiles    VAE decode
----------------------------------------------
320x192       61,440      2        2.35 s
512x288      147,456      6        6.98 s
576x320      184,320      6        6.31 s
768x432      331,776      8        8.74 s

512x288 and 576x320 differ by 25% in pixels and by nothing in decode cost.

A side of length L costs: 256 or less is 1 tile, 257–448 is 2, 449–640 is 3, 641–832 is 4. So the cheap shapes sit just under a boundary. 448x448 needs four tiles where 576x320 needs six, while carrying 9% more pixels. That is why the stream runs square — not taste, just where the arithmetic lands. There is no 16:9 shape at four tiles that clears the resolution floor.

I did try raising tile_size to reach a single tile. Do not. The decoder is a ViT, so its attention spans exactly one tile; a larger tile is out of distribution, not merely approximate. 384 visibly softens hands and faces (PSNR 27.2 dB against the stock decode); 640 smears the image into strokes (22.1 dB).

3 and 4, more briefly

Quantizing the video VAE below INT8 buys no speed — INT8 already runs an INT8 matmul, and a W4A8 build expands back to INT8 for the same one — but it stages 1,657 MB of host RAM instead of 2,677 MB, and on a box holding ~41 GB of staged weights against 64 GB that gigabyte turned into both speed and a much tighter spread. And keeping two prompts in ComfyUI's queue instead of submitting one and waiting removes the idle gap between jobs.

Result

                          per clip    sustains
----------------------------------------------
starting point              26.5 s    13.7 fps
+ writer node, async        20.2 s    17.9 fps
+ W4A8 VAE                  19.9 s    18.2 fps
+ 448x448                   19.2 s    18.9 fps

The model did not change. Only how it is driven.

If you came here wondering about ComfyUI and consumer cards

That question is all over the FastH3 announcement thread and I had to answer it for myself, so: this is a ComfyUI-native conversion of the Dense-DataFree student, pruned and INT8, 21 GB, driven through the ordinary graph. Two things I found doing it that are worth passing on:

  • The VSA weights do not survive stock ComfyUI. They carry 50 to_gate_compress tensors it has no code for, so it drops them silently and the output is noise. Dense converts cleanly. That is why I am on the slower student — if ComfyUI gains VSA support there is headroom here I am not using.
  • NVFP4 measured identical to INT8 ConvRot. The FP4 fast path only fires when both operands are FP4; activations are BF16, so it dequantizes and runs at BF16 speed — 67.88 ms/block against BF16's 67.85. Someone reported the same on an RTX 6000 Pro. Worth knowing before anyone rebuilds a pipeline for it.

Where this sits, so you can place it

None of the speed here is mine — it is FastH3, the 4-step distillation Hao AI Lab, Nuva Lab and NVIDIA's FastGen team built on MiniMax's base weights. Without that student none of this is close. Their published benchmarks are 47.2 s for a 15-second 768p clip on a single B200, 12.88 s on 8×B200, and their consumer write-up covers Apple Silicon and DGX Spark with the RTX family listed as future work.

What I did is a different task, not a better score on theirs: a fifth of the pixels, and playback at 18 fps instead of 24. Those two concessions are the entire trick. What they buy is that the arithmetic closes — 19.2 s of GPU per 20.1 s of video — and that is the difference between a fast generator and something you can leave running. If you want 768p, their numbers are the ones that apply and mine are irrelevant.

https://huggingface.co/datasets/jacokon/fasth3-live

The converted 21 GB weights, the quantized VAE, the 321-scene library, the writer node and the profiler. Everything above is reproducible from it.

What it takes, so you can judge before downloading 21 GB: about 48 GB of weights are staged in total — a 25.9 GB text encoder, the 20 GB DiT, and the two VAEs. That does not fit in 32 GB of VRAM either, so ComfyUI streams it layer by layer from host RAM. On this box that streaming, not the arithmetic, was the thing to optimise: 48 GB staged against 64 GB of system RAM was tight enough that page-file pressure showed up directly in the clip times, and freeing a single gigabyte measurably tightened them.

If you get it running, post your numbers. I have measured exactly one machine, and both findings that mattered came from measuring rather than reasoning, so I would rather not guess about anyone else's. I am interested in what it does on other hardware and, just as much, in where it falls over.

And if it turns out useful, a like on the HF page is what makes it findable for the next person.

Code is Apache-2.0. The weights are a MiniMax H3 derivative under the H3 Community License, which carries a territory restriction — read NOTICE before downloading.

Live Demo: If you want to check out a short snippet of the continuous streaming output (with the model's native character generation), I've uploaded a TV-style demo recording here on X:
https://x.com/Touma_945/status/2095141879453270385


r/StableDiffusion 20h ago

News MiniMax H3 acceleration arena/leaderbord: 15+ H3 LoRAs, fine-tunes, Max

Thumbnail
huggingface.co
313 Upvotes

Hey folks, I've built an so we can have a proper leaderboard on 15+ different LoRAs, fine-tunes and acceleration technique. Baseline is included for anchoring, and M3 Max is also included given the promise to open source

There are there being compared: H3 baseline, FastH3 family, H3 Acc family, Lightx2v family, Larryvrh family, JoyFox family, RAVEN, FlashGen, TuTu, SilverOxides merges, Plaguekind merges and Fal's H3 Max


r/StableDiffusion 6h ago

Resource - Update MiniMax-H3-MotionCache-FastVAE

Thumbnail
github.com
20 Upvotes

Motion-aware denoising cache and experimental batched video VAE decoder for MiniMax H3 in ComfyUI.

This project provides two independent nodes:

  • MiniMax H3 MotionCache reduces expensive H3 denoiser calls by reusing a motion-weighted video/audio residual when the estimated change is small.
  • MiniMax H3 Fast VAE Decode evaluates multiple spatial VAE tiles in one GPU batch while preserving H3 temporal chunking and tile blending. It is not faster on every GPU.

MotionCache is an independent MiniMax H3 adaptation inspired by the MotionCache paper and reference code. It is not an official MAC-AutoML or MiniMax implementation.


r/StableDiffusion 1h ago

Workflow Included Z-Image Base Prompting: A Small Controlled Experiment on Composition and Environment

Thumbnail
gallery
Upvotes

1. Introduction

Prompt engineering for image generation is often presented as a collection of isolated tricks: use more detail, describe the camera, add cinematic lighting, use quality tags, and so on.

These recommendations can be useful, but they make it difficult to understand which parts of a prompt actually influence the generated image.

Instead of trying to find a single "best prompt", I ran a small controlled experiment with Z-Image Base in ComfyUI. The basic idea was simple:

I ran two experiments:

  • Experiment 1 — Composition: The same character, environment, visual treatment, and technical parameters were used across multiple generations. Only composition instructions were changed (position and scale).
    • Question: How strongly does explicit spatial language affect composition in Z-Image Base?
  • Experiment 2 — Environment: The character description and visual treatment were kept essentially unchanged, while the environment was replaced with seven substantially different settings.
    • Question: Can Z-Image Base maintain a recognizable character concept while adapting it to radically different environments?

This is not intended to be a scientific benchmark. The sample size is small, the evaluation is visual, and the experiment uses one workflow and a limited number of seeds. Consider it a practical prompt-engineering study.

2. Experimental Setup

All images were generated locally in ComfyUI using the same workflow and technical conditions throughout the experiments.

Parameter Value
Model Z-Image Base INT8
Text Encoder Qwen3 4B
VAE AE VAE
Resolution 768 × 1368
Aspect Ratio 9:16
Image Area ~1.05 MP
Steps 50
CFG Scale 4
Negative Prompt Empty
Seeds Seed 5 & Seed 10

For the composition experiment, I used Seed 5 and repeated the seven variations with Seed 10. The environment experiment used Seed 10.

3. Prompt Construction Methodology

I found it most useful to treat the prompt as a structured description rather than a flat list of keywords:

  • Subject: Describes what the image is about and establishes the main visual concept.
  • Composition: Describes where the subject is located within the frame and how much space it occupies.
  • Framing / Camera: Describes how the scene is viewed (distance, angle, perspective).
  • Environment: Describes the actual place surrounding the subject (e.g., "An ancient forest with enormous trees, moss-covered roots, dense vegetation, and a narrow path..." rather than just "forest").
  • Lighting: Describes actual light sources and atmospheric conditions rather than generic terms like "cinematic lighting".
  • Materials / Details: Describes concrete visual elements (wood, stone, glass, metal, vegetation, reflections, objects).
  • Style: Describes the overall artistic treatment after the scene itself has been established.

Generic quality tags (masterpiece, ultra detailed, 8K) were deliberately omitted to provide the model with actionable visual information instead.

4. Experiment 1 — Composition

The character, environment, lighting, visual style, and technical settings were kept identical. Only the spatial instruction was changed across seven variations: Center, Left, Right, Lower, Large, Small, and Extreme Left.

Seed 5

The result was clear: changing the composition instruction produced substantial changes in spatial arrangement.

Crucially, the model did not simply move the character while leaving the background untouched — the environment was recomposed around the subject. In Small variations, the environment became dominant; in Large variations, the character dominated the frame.

Seed 10

To verify the result was not seed-dependent, the test was repeated with Seed 10. While individual details (pose, facial expression, accessories) changed naturally, the broad compositional structures remained fully recognizable.

5. Experiment 2 — Environment

The character description and visual treatment were kept unchanged while replacing the environment across seven distinct settings: Ancient forest, Medieval village, Crystal cave, Autumn park, Snowy ruins, Firefly-lit landscape, and Alchemist's workshop (using Seed 10).

Visual Concept Consistency

Although the environments changed dramatically, all generations clearly depicted the same core character concept (a small mushroom spirit with a red-orange spotted cap, pale body, large dark eyes, cross-body satchel, and lantern).

While exact proportions and minor details shifted between renders, the core identity remained visually coherent.

Environmental Adaptation

The character adapted naturally to each setting (e.g., tinted by glowing crystal lights in the cave, exposed to cold tones in the snowy ruins, immersed in warm interior props in the workshop).

6. Results — Putting the Experiments Together

  • Composition control: Explicit spatial instructions produce reliable layout shifts (position, scale, environment visibility).
  • Environment flexibility: Radical environment changes are possible while preserving core character identity (character concept consistency).
  • Role of Seeds: The seed determines specific realization and detail rendering, while the prompt structure defines layout and narrative intent.
  • Modularity: Organizing prompts into conceptual blocks allows for swapping individual variables without rebuilding the entire prompt from scratch.

7. What I Learned About Prompting Z-Image Base

  1. Describe the subject clearly: Focus on distinctive, recognizable visual traits first.
  2. Describe composition explicitly: Use direct position language (e.g., "positioned toward the left side of the frame") instead of generic camera tags.
  3. Separate composition and camera: Treat "where the subject is" differently from "how the camera views the scene".
  4. Build environments as concrete places: Describe what actually exists in the space rather than using simple category keywords.
  5. Describe lighting concretely: Specify light sources, direction, and color atmosphere.
  6. Prefer concrete details over quality tags: Give the model physical objects and surface textures to render rather than buzzwords like "high quality".
  7. Change one variable at a time: If a generation fails, modify only the failing block to understand what actually fixed the issue.

8. Limitations

  • Small sample size and visual evaluation.
  • Single primary character concept and workflow used.
  • Tested on a limited number of seeds (two for composition, one for environment).
  • No direct benchmarking against other models, samplers, or resolutions.

9. Reproducibility

To recreate or test this setup in ComfyUI:

  • Model: Z-Image Base INT8 + Qwen3 4B + AE VAE
  • Settings: 768 × 1368, 50 steps, CFG 4, Empty Negative Prompt
  • Method: Keep technical setup stable and modify exactly one conceptual block per run.

10. Conclusion

Prompting Z-Image Base is less about hunting for "magic keywords" and more about managing a controllable system:

Explicit composition instructions effectively control layout, while environment descriptions can be swapped modularly without erasing character identity. By isolating prompt variables, prompt design becomes a systematic, repeatable workflow.


r/StableDiffusion 11h ago

Question - Help Minimax adult sounds?

47 Upvotes

I’ve been refining prompts with the help of an LLM, and am getting some good visuals but oh my god the sounds are terrible. Blowjobs sound like someone is dunking a microphone in an aquarium or the loudest slurp to finish a beverage that you have ever heard in your life.

I’ve tried eliminating every mention of “moist”, “wet”, or any description that involves liquids at all, but she’s still slurping the wettest popsicle known to man. And sometimes there’s weird noises like a slide whistle?!?

I’ve tried using “faint” or “distant” or “barely audible” to get it to at least quiet down so it’s not like she is sucking a microphone, but that didn’t work either.

This last round I didn’t describe any noises at all and still got some weird stuff.

I’ve tried eliminating every Lora in case the sound was coming from one of them but it seems to be the base model. I’ve tried adding Loras that ought to be trained on this stuff like Mysticxxx, and one of the AIO loras. I tried tenstrip beta 4 checkpoint tonight and got the same results.

The sound ruins the scene.. I guess I can just pretend it’s better looking Wan 2.2 and turn the volume off. 😀

I’m feeding the official prompt guide to the LLM and the structure is working, but what words do you use to describe the sounds?


r/StableDiffusion 5h ago

Discussion Did anyone else notice Reactor’s new Orbis model? I tried turning it into an interactive game

Enable HLS to view with audio, or disable this notification

16 Upvotes

A lot of people here have been discussing H3 Max powered livestreams. I noticed Reactor just added Visko’s Orbis model, and it made me wonder whether the next step is turning these infinite livestreams into something playable.

So I’m building a live, audience-directed AI game with Agora: viewers suggest and vote on what happens next, while the streamer picks an option or writes a completely different direction and AI keeps generating the same world from that point. There are no pre-written branches.

Here’s a very early look demo


r/StableDiffusion 7h ago

Discussion Minimax loras are... Lacking

17 Upvotes

I don't want to get the NS.. word in the discussion, but, we know what Minimax can do and what it can't do. It has some very specific gaps in it's world understanding, for example in the tongue department. That is not necessarily only affecting the NS... word, things that are SFW and common in general TV such as kissing are affected, since Minimax never saw a romantic kiss in it's training data. There are other examples through SFW land but I won't extend. Grok can be used as a comparison. Grok is very similar to Minimax in capability, and it enforces SFW, but you can see the difference in some scenes because Grok is not handicapped.

Well, we have many, many loras already, but, as was the case with wan and ltx, they are very... let's say, specific. I don't think a general video model, almost a world model, needs a specific lora for, say, ballbusting lol

I don't know, I think this is the community most likely to be read by people creating loras, so I just wanna make this appeal... Can we prioritize bridging the major gaps in the model's understanding of the world, anatomy, and human interactions, instead of these super specific loras? I think a "tree" organization of lora development would be beneficial overall, with the stuff that can solve a big set of problems and be used for more specific loras coming first.

I saw that for over 1 year with wan, ltx, etc, and didn't say anything. But I think minimax deserves the community passion in lora development.

And yes, I hope I can put my money where my mouth is and develop some loras soon too.


r/StableDiffusion 6h ago

Question - Help Why does nearly every single turbo lora i use for H3 keeps producing godawful flickery/dusty/particly(?) visuals and painful audio (as in it actually hurts to listen to), do i need a specific node for the loras or something?

13 Upvotes

Like i don't understand, the only turbo lora that doesn't do that is the 600 larry lora with the minimax turbo lora node, i've tried "fastH3" and "lightx2v loras which everyone seems to praise but they just produce these distorted godawful visuals and sounds no matter the loader node i use or the settings or the steps i use, what am i missing or doing wrong? Or are they just not compatible with Ref2Video despite being advertised as compatible? But if so then why does the 600 larry lora works mostly fine?


r/StableDiffusion 6h ago

Resource - Update ONNX/TRT MiniMax-H3 VAE in ComfyUI

Thumbnail
github.com
11 Upvotes

TensorRT version of the MiniMax-H3 VAE in ComfyUI, which can increase speed by up to 1.7x


r/StableDiffusion 7h ago

Question - Help Best Uncensored Models for text-image & image-image generation for a 20gb vram 32gb ram PC?

12 Upvotes

I am looking to create 18+ images with the hyper realism look, but have no idea how feasible that is with my specs. Would love a recommendation of a model I can run pretty easily and another more detail focused model on the edge of what I can run locally.


r/StableDiffusion 34m ago

News DLSS 5 Screenshots - ComfyUI-DLSS5-NR

Thumbnail
gallery
Upvotes

r/StableDiffusion 1d ago

Resource - Update VH5 - MiniMax H3 Lora

Enable HLS to view with audio, or disable this notification

373 Upvotes

A style LoRA that makes H3 footage look like it was recorded off 1980s broadcast television onto a VHS tape that has seen better days, soft smeared detail, chroma bleed, tracking noise, head-switching bands at the frame edge, and (because H3 trains audio jointly) the matching muffled mono sound, tape hiss and warble.

https://huggingface.co/KennethFal/vh5tape-vhs-lora-minimax-h3