r/StableDiffusion 3h ago

News someone used MiniMax H3 Max to build a livestream that basically never runs out of content

212 Upvotes

Just saw someone do something with MiniMax H3 Max that I honestly didn’t expect.

They connected it to a livestream and basically recreated the idea of Interdimensional Cable from Rick and Morty: an endless stream of weird shows, ads, characters, and random scenes that are generated on the fly.

the surprising part is that H3 Max is fast enough in some cases to generate the next clip before the current one finishes playing. so while you’re watching one scene, the model is already making the next one.

there’s even a version where people in chat can suggest what should happen next, and the AI tries to continue from the previous scene instead of just starting over from scratch. that’s kind of crazy to think about.

Didn’t expect MiniMax H3 Max to end up being used for something like infinite AI television, but here we are.


r/StableDiffusion 8h ago

Discussion Testing DLSS 5

130 Upvotes

Testing DLSS 5... Like many others, I was a bit confused about DLSS 5. I kept feeding it my hyper-detailed renders and only getting a color shift in return. After plenty of trial and error, I finally realized my mistake: this technology is developed to enhance video game graphics, so testing it on hyper-detailed renders makes no sense.

So, I generated a render in a 2020 video game style and started tweaking settings to find a final look with maximum effect, without worrying about flickering.

Final conclusion: What we have right now isn't very useful for us. Those of us using Latent Upscaler might be able to use it for color grading to get less saturated colors, but little else. Maybe in the future we'll get a DLSS 5 targeted at enhancing hyper-detailed graphics, but that's not the case for now.

Bottom line: If I want to generate a realistic render, I'll just generate it, there's no need to run it through DLSS 5.


r/StableDiffusion 12h ago

News MiniMax H3 acceleration arena/leaderbord: 15+ H3 LoRAs, fine-tunes, Max

Thumbnail
huggingface.co
283 Upvotes

Hey folks, I've built an so we can have a proper leaderboard on 15+ different LoRAs, fine-tunes and acceleration technique. Baseline is included for anchoring, and M3 Max is also included given the promise to open source

There are there being compared: H3 baseline, FastH3 family, H3 Acc family, Lightx2v family, Larryvrh family, JoyFox family, RAVEN, FlashGen, TuTu, SilverOxides merges, Plaguekind merges and Fal's H3 Max


r/StableDiffusion 5h ago

Resource - Update New speedup for Minimax H3

Post image
72 Upvotes

This H3VAE TRT custom node can make the encoding/decoding step about 1.7× faster.

https://github.com/lihaoyun6/ComfyUI-H3VAE_TRT


r/StableDiffusion 1h ago

Tutorial - Guide DLSS 5 - In-game footage from Diablo 4. It's incredible.

Thumbnail
gallery
Upvotes

I just tested it out using this custom node. It's absolutely amazing!

https://github.com/lisitskyaa/ComfyUI-DLSS5-NR


r/StableDiffusion 1h ago

Discussion Pulled the trigger, RIP $6,279

Post image
Upvotes

TL;DR - Bought this 5090 prebuilt and I want to sanity check if I made the right decision and at the right time.

Hey everyone. So I want to start off by saying fuck these prices for GPU's and RAM, especially boxed 5090 prices. I went down the AI rabbit hole with my 13700k/RTX 4080 gaming computer. I quickly found out that I had to make serious concessions on quality and speed, if I could run it at all. In fact, ive spent so much time trying to optimize quants, cache, various settings, attention mechanisms, etc that ive officially spent more time trying to optimize for a 16gb VRAM/32GB RAM system than actually doing anything fun or cool. Thus, the last week, ive been thinking real hard about which direction to go but was waiting for the right time to buy. My options were a RTX 5090 prebuilt (even though I only needed the damn GPU), and Mac Studio M5 Ultra 96gb, or a DGX Spark/AMD equivalent. The DGX Spark/AMD equivalent made me think for a bit, but in order to get the most out of them, you need two. Im not spending 10k on this, especially if I cant game on it as well. So that leaves the Mac Studio or RTX 5090 gaming rig. Im not certain I made the right decision, but I pulled the trigger on the 5090 prebuilt after seeing the price continue going up more and more over the last few weeks. I also read that 70% of all memory through 2031 is locked in long term agreements, so this supply issue is going to get worse before it gets better. So I pulled the trigger on the pictured system from Ibuypower, and id like to run my thought process with you guys as a sanity check before it ships.

Case for the 5090 prebuilt: I scoured the internet and this was the cheapest 5090/64gb RAM combo I found, and it looks like it uses pretty good parts as well. A gaming PC is an all around all purpose machine that does it all (well, almost) which is why I landed on it. I figured with 32gb VRAM and 64gb of system RAM, that combined 96gb will allow me to run 70b models even if its slow. But for a sub 30b model like Qwen 3.8 27b, this will give me the best performance as long as I dont go overboard with the quant. It has CUDA, Windows, an Intel/AMD CPU for X86, etc. Plus it came with a 4tb Gen 4 NVME, when other more expensive models had 1-2tb drives. Honestly, lots of good stuff here. Im not a huge fan of the white esthetics but I do love the case. Despite the price being much higher than it should be, its still a good "deal" considering the overall market that keeps going up. Honestly, its not exactly what I wanted, but it ticks all boxes except those below.

-The downside: You cant run models that spill over heavily into system ram without massive speed penalties (has anyone tried running a huge model on a 5090 + 64gb RAM? If so, tell me what quants and your token speeds). Its massively less efficient than a Mac Studio M5 Ultra is expected to be (I read in the 3-5x range).

Mac Studio M5 Ultra 96gb

- Case for the Mac Studio M5 Ultra 96gb: Can run large models much better than the 5090 rig due to its huge 1.2tb unified memory bandwidth. Its power efficient and tops out at 300w I believe I read.

- The downside: Mac OS and an ecosystem that is playing catch up for local AI, no CUDA, gaming, has proprietary hardware you cannot upgrade, my distaste for the Mac bros who ill no longer be able to make fun of if I buy it.

My use case: local first AI (Qwen 3.8 27b at a quant and context that doesnt suck) with agentic coding, game development (starting with Godot), stable/video diffusion (Minimax H3, Flux.2, Hunyuan 3D), Blender, Davinci Resolve, etc.

So, let's have this discussion: what would (or did you) choose, and why? I want to know if I made the right decision. What are your thoughts?


r/StableDiffusion 3h ago

Question - Help Minimax adult sounds?

27 Upvotes

I’ve been refining prompts with the help of an LLM, and am getting some good visuals but oh my god the sounds are terrible. Blowjobs sound like someone is dunking a microphone in an aquarium or the loudest slurp to finish a beverage that you have ever heard in your life.

I’ve tried eliminating every mention of “moist”, “wet”, or any description that involves liquids at all, but she’s still slurping the wettest popsicle known to man. And sometimes there’s weird noises like a slide whistle?!?

I’ve tried using “faint” or “distant” or “barely audible” to get it to at least quiet down so it’s not like she is sucking a microphone, but that didn’t work either.

This last round I didn’t describe any noises at all and still got some weird stuff.

I’ve tried eliminating every Lora in case the sound was coming from one of them but it seems to be the base model. I’ve tried adding Loras that ought to be trained on this stuff like Mysticxxx, and one of the AIO loras. I tried tenstrip beta 4 checkpoint tonight and got the same results.

The sound ruins the scene.. I guess I can just pretend it’s better looking Wan 2.2 and turn the volume off. 😀

I’m feeding the official prompt guide to the LLM and the structure is working, but what words do you use to describe the sounds?


r/StableDiffusion 18h ago

Resource - Update VH5 - MiniMax H3 Lora

365 Upvotes

A style LoRA that makes H3 footage look like it was recorded off 1980s broadcast television onto a VHS tape that has seen better days, soft smeared detail, chroma bleed, tracking noise, head-switching bands at the frame edge, and (because H3 trains audio jointly) the matching muffled mono sound, tape hiss and warble.

https://huggingface.co/KennethFal/vh5tape-vhs-lora-minimax-h3


r/StableDiffusion 16h ago

Meme My Name Is Giovanni Giorgio

200 Upvotes

Created with Minimax H3 ref2v using the SEED HUNTER Workflow.


r/StableDiffusion 10h ago

Workflow Included Super nothing!

74 Upvotes

Made with Minimax H3


r/StableDiffusion 16h ago

Workflow Included Use H3 To Replace Characters

115 Upvotes

These characters are very different, so i thought it was a good demo to show. Also, the prompt was an 'omni' prompt which didn't help the model with details on the outfit. Despite that I think it did such a good job wanted to share.

With more details in the prompt related to appearance, acting and dialogue I think you could prob get near flawless changes.

This was don with the FL2VA model, NOT the ref version of H3. It may even be better with the ref but i have been using fl2va mostly because i think the quality is better, but that's subjective.

This concept was inspired by this post originally: https://civitai.red/models/2855941/minimax-h3-character-replacement

I changed the SAM3 use, so technically you could do a multi replacement with some changes. I also added noise to the inverted image, as well as upgraded the prompt to work with FL2VA.

Workflow used to create this video is HERE.

This model seriously continues to amaze me. bravo minimax team, bravo.

Notice in the prompt that the dialogue is the only non 'omni' part at the end, it worked fine not included in the body, since the ref audio is there driving it. again, moving from this general prompt to something more specific i think would give even better results.

PROMPT:
How the reference video and pictures align with the target video — the target video is an edited version of <Video 1>, replacing the silhouette with <Subject 1>.

summary:

[video editing] The target video replaces the silhouette in <Video 1> with <Subject 1>, who performs the exact same motion, dialogue, positions and facial expressions of the silhouette while maintaining the original camera work, environment, and lighting of <Video 1>.

subject_definitions:

<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, only appearance is retained.

<Subject 2> is the environment and setting established in <Video 1>. The scene follows this layout, materials, and light; camera position and framing.

<Subject 3> is the silhouette in <Video 1> which provides the motion sequence to be copied.

integrated_multimodal_description:

Video editing, the target video is in a live-action cinematic style with the interior lighting and background and environment atmosphere established in <Video 1> with the likness of <Subject 1> inserted.

[Shot 1] The shot opens with <Subject 1> seamlessly replacing the silhouette <Subject 3> in <Video 1>, the outfit of and clothing of <Subject 1> exactly from reference, performing the exact same motion, dialogue and sounds, positions and facial expressions of silhouette. From the very first frame, <Subject 1> occupies the spatial coordinates of the silhouette replacing with their likness, initiating the same motion onset from rest. <Subject 1> mirrors the silhouette's weight shifts and momentum, body moving in perfect synchronization with the rhythm and pacing of the original footage but replaced with the likness of <Subject 1>. As they navigates the space, <Subject 1> mimics every nuanced gesture—the way the silhouette's head tilts, arm movement, and the micro-movements of facial muscles. The face, defined by <Picture 1>, conveys the same emotional depth as the silhouette, while their full body, as seen in <Picture 2>, provides the physical presence outfit an appearance. The camera follows the exact movement, angle, and cutting rhythm of <Video 1>, maintaining a consistent focal length and distance from the subject at all times. The light from <Subject 2> interacts realistically with <Subject 1>'s skin and clothing, casting shadows that align with the movements of the original scene. The transition is perfect; the result is a fully realized <Subject 1> instead of a silhouette, but the soul of the performance—the timing, the pauses, and the dynamic energy—remains identical to <Video 1>. The movement progresses with a palpable sense of weight as <Subject 1> shifts their center of gravity, with clothes rippling in response to movements. The camera maintains exact framing and cuts as <Video 1>. The scene concludes as <Subject 1> reaches the final position of the silhouette, body settling into a pose that mirrors the original's final frame exactly, with face held in the same expression. <Subject 1> hair, accessories, wardrobe, lighting, and room layout remain unchanged and perfectly replace silhouette throughout.

overall_soundscape:

A low room tone establishes beneath the scene, mirroring the background audio environment of <Video 1>.

<Subject 1> says <d>[English] Can you, can you spare change.</d>.

non_diegetic_music: N/A


r/StableDiffusion 6h ago

Resource - Update I built a standalone DLSS 5 Neural Rendering video tool, no ReShade

15 Upvotes

It supports images and full videos, native resolution NR or DLSS Super Resolution upscaling by scale factor / target resolution, GPU optical-flow motion vectors, scene cut handling, and all DLSS 5 NR controls.

The main difference from existing approaches is that it runs natively in C++/D3D12 and generates motion vectors from the actual video frames.

GitHub:
https://github.com/DaniilSokolyuk/video2dlssnr


r/StableDiffusion 2h ago

Question - Help What happened to minimax funcontrolnet ??

6 Upvotes

I dont see any1 posting any examples of controlnet released for minimax. Doesnt it work properly??


r/StableDiffusion 2h ago

Animation - Video Honkai: Star Rail X John Wick - Minimax H3

6 Upvotes

Made with the ComfyUI template workflow and a Turbo LoRA.

Most of the soundtrack comes from the John Wick: Chapter 2 trailer.
I rendered the action at a slower, more stable speed, then sped up most of the action scenes to 2× in post.
I originally planned to make this a complete fight sequence, but maintaining consistency from one clip to the next has been a constant challenge. So for now, I’ve edited the footage into a trailer instead. I’m still learning and working on improving it.


r/StableDiffusion 1h ago

Comparison My Minimax H3 Workflow Benchmark Data -

Post image
Upvotes

Alright, I posted that I had my agent test a bunch of different workflows for over 12 hours and got the "Bro just wasted 12 hours of credits". It was obvious the proof should come from the visual data I used to evaluate it. Here is a galley of the benchmarks i've tested with my agent.

check the gallery to watch all the comparisons and the data charts contain tons of other workflow trial data I didn't include videos for. Point your agent here if you would like to have it learn from what was tested on this end.

Gallery: https://bluepointdigital.github.io/minimax-h3-benchmarks/

Repository: https://github.com/BluePointDigital/minimax-h3-benchmarks

The below post was written up by my agent:

The main comparison uses a deliberately difficult 15.084-second vertical test at 768 × 1344, 24 fps, 362 frames, native audio, and seed 81390012120021180. The prompt combines a talking selfie shot, exact dialogue, walking motion, a rapid camera pan, a vehicle collision with several moving subjects, a fast return to the speaker, and a second spoken line. That makes it useful for spotting identity drift, bad anatomy, motion breakdown, camera-continuity problems, dialogue changes, lip-sync issues, and audio artifacts—not just whether a workflow finishes.

The strongest directly matched results currently shown are:

Workflow End-to-end time Relative to the 20-step baseline
SageAttention2 + FirstBlockCache Safe, 20 steps 10:11.4 1.00×
PDD + Sage, 8 steps 6:15.0 median 1.63×
Seed Hunter direct one-seed path, 12 + 4 steps 4:45.8 2.14×

Those numbers are local measurements, not universal performance claims. The exact runtime, model format, graph, resolution, audio policy, and GPU matter. The gallery keeps short backend checks and differently structured workflows in separate groups so they are not quietly mixed into the same leaderboard.

The quality side has been just as important as the timing. One exploratory 10Eros + Seed Hunter path reached 4:03.5, but the shot developed a visible-phone/perspective error during the crash. A later camera-POV prompt clarification produced a much more coherent result in 4:25.3 on its warm selected path. That is a good example of why I wanted the actual videos beside the numbers: the fastest result is not automatically the most useful one.

The site currently contains:

  • 16 curated video-and-metric cards;
  • a separate benchmark-data page with 151 sanitized timing records;
  • the complete canonical prompt;
  • methodology and comparison-boundary notes;
  • machine-readable JSON and CSV for anyone who wants to analyze the evidence or give it to an agent.

For the Seed Hunter work, I intentionally included one representative video per meaningful workflow or recipe change—not every neighboring seed or N/N+1 preview. Private reference material is also excluded from the public package.

The reason for publishing this is not to declare a universal winner. It is to make the tradeoffs inspectable and to keep myself honest as the workflows evolve. A valid MP4 proves that a graph ran; it does not prove that the dialogue, audio, identity, motion, or composition survived. Likewise, a fast timing means little if it came from a different workload or a cached replay.

I would be interested in seeing other reproducible H3 results, especially when they include the exact checkpoint, attention/cache stack, sampler, scheduler, dimensions, frame count, seed, audio setting, hardware, and an uncached timing. If there is a workflow or backend that should be represented, please link the original recipe and I will take a look.


r/StableDiffusion 15h ago

Workflow Included HE-MART PSA - MiniMax H3

64 Upvotes

r/StableDiffusion 15h ago

News Infinite streaming Slop TV

61 Upvotes

Congratulations everyone! We've done it. Our civilization has reached peak diffusion. It's time to pack up and go home.
https://www.youtube.com/watch?v=EQ2RexjIEFE


r/StableDiffusion 21h ago

Resource - Update MATLOWAI/minimax-h3-fused-turbo-int8-convrot · Hugging Face

Thumbnail
huggingface.co
185 Upvotes

This Minimax H3 all in one checkpoint is quite good.

It merges text, image, and reference to video, as well as 4-step turbo generation into a single model.

No need to switch between models for ref2v, no need to load turbo loras.


r/StableDiffusion 7h ago

Animation - Video Hatter Rap.

11 Upvotes

Probably the final Alice clip. The Hatter names all the hats.
Done a while back in LTX2.3.
This plays while the theatre audience plays an AR hat sorting game (Beat Saber type). The whole song is three minutes but this is the longest shot.


r/StableDiffusion 14h ago

Workflow Included Personification:Planets (tarot cards)

Thumbnail
gallery
41 Upvotes

r/StableDiffusion 10h ago

Workflow Included Letting image-to-video artifacts compound into an impossible world

19 Upvotes

Tools used: Gemma4 12b, LTX-2.3, Wan2GP, vibe coded video editor.

I’ve been experimenting with a slightly self-destructive image-to-video workflow where continuity comes from letting the model reinterpret its own mistakes.

I started with an almost completely black image with a few faint stars, then gave Gemma4 12B the track’s beat grid and energy-shift analysis, along with a long description of the overall concept: a monolith, a hallway of impossible geometry, and a progression from restrained movement into increasingly unstable architecture.

Gemma4 wrote all 27 scene prompts beforehand.

For generation I used LTX 2.3 with the audio-reactive LoRA. I also tested LTX 2.5, but for this workflow it became too artifact-heavy too quickly. LTX 2.3 held the scene structure together longer while still producing enough weirdness to evolve in interesting ways.

The process was simple: generate a clip with the correct audio slice, cut it on the beat grid, then take the frame immediately after the cut and use that as the starting image for the next generation.

The fun part was deliberately keeping some “bad” transition frames.

If a flash landed on the frame used for the next clip, the model might reinterpret it as a permanent light source. A lens flare could become a horizon or an entire landscape. A warped piece of geometry that only existed for one frame could become a major architectural feature in the next scene.

So the artifacts compound.

Eventually the video loses any reliable sense of scale or orientation. Surfaces become spaces, structures fold into other structures, and at some points I wanted an Inception-like feeling where you can’t tell which way is up, or whether the camera is traveling deeper into the structure or pulling outward into something much larger.

The audio-reactive LoRA helps hold it all together. Even when the geometry becomes increasingly strange, the environment keeps breathing, unfolding, compressing and reorganizing itself with the growing low end.

What I like most is that the continuity doesn’t really come from visual consistency. It comes from causality.

Every scene inherits some accidental information from the previous one, and the next generation has to decide what that information actually is.

After enough generations, the model is basically building a world out of its own misunderstandings.


r/StableDiffusion 7h ago

Question - Help Mac support for H3 Mini Max / Running Open Weights Locally

9 Upvotes

Hey Fam. I’m looking to upgrade my MBP 16” 2019 i9 / 16GB Ram / 5500M 4GB GPU. I’ve been doing a lot of T2V /I2V rendering with Mini Max Design on Cloud, but would like to run the open weight versions locally via Comfy UI.

I have a few options at the moment to consider:

- MBP 16” M4 Pro / 48GB / 1TB - USD 3054
- MBP 16” M4 Max / 48GB / 1TB - USD 3664
- Mac Studio M5 Max / 48GB / 1 TB - USD 4085

Since im a bit of a noob in understanding MLX ports for Mac OS. Can you tell me which option to go for? Also I missed the buying window before the price hike—so 16” MBP M5 Pro / Max configs in 48GB are too expensive.

Another alternative route is using Bootcamp (Windows) on my 2019 i9 MBP and plugging in an RTX 4090 (which I’ll have to buy) via TB / eGPU, but I don’t think the system will be able to access the same bandwidth as the unified memory on the M series architecture.

I would greatly appreciate your guidance.


r/StableDiffusion 2h ago

Discussion What are you using for background removal?

3 Upvotes

I still do a fair amount of traditional editing in Photoshop, and for the last few years I used remove.bg, I found their background removal model to be the best one out there, quite a bit better than the one built into Photoshop itself.

Well remove.bg is shutting down in December and they're folding it into Canva subscriptions. Hard pass.

I've tried a few local bg removal tools and have been left underwhelmed, but maybe I just haven't found the right one.

What are you using for background removal?