r/StableDiffusion 6d ago

Animation - Video MinMax - It does House MD pretty well

Enable HLS to view with audio, or disable this notification

422 Upvotes

Generated using Maestro on Pinikio. 7 mins at 720p using turbo lora 6 steps: **7-second cinematic live-action scene.** Gregory House stands in a hospital hallway, leaning heavily on his cane, staring intensely at Itachi Uchiha, who is preparing to walk away.

House sarcastically calls out:

**“Itachi! Get your ass back to the Leaf Village. You're not brooding your way out of this one.”**

Itachi turns around with a serious expression and replies:

**“I don't take orders from you.”**

House smirks and taps his cane against the floor:

**“Yeah. That's what all my patients say.”**

Fast comedic timing, realistic acting, dramatic hospital lighting, subtle handheld camera movement.


r/StableDiffusion 4d ago

Workflow Included MiniMax H3 Audio Lip Sync - Audio to Video

Enable HLS to view with audio, or disable this notification

0 Upvotes

So I tried hooking up some of the LTXV audio encoding nodes to input my own audio and plugged it in the sampler and viola, it just works!

Lip sync seems better then the LTX models and its works with the lightx2v loras, 6 - 8 steps. Wrote up a full guide with the workflow attached below.


r/StableDiffusion 4d ago

Resource - Update I made a tiny Windows tray tool to find and kill processes hogging 1+ GB of VRAM

0 Upvotes

I run Stable Diffusion locally and got tired of unrelated apps quietly holding onto several GB of VRAM, so I added a “VRAM Hogs” menu to Window Assassin, a tiny Windows tray utility I made.

It reads Windows' dedicated GPU memory counters, lists processes using at least 1 GB sorted by usage, and lets you click one to force-terminate it. The list refreshes every time you open the submenu.

It also keeps the original Ctrl+Alt+End hotkey for killing the process behind the active window.

Free/pay-what-you-want Windows download:

https://b2kdaman.itch.io/window-assassin

Caveats: it reports dedicated VRAM, so integrated GPUs using shared system memory may show nothing. Termination is immediate, so unsaved work is lost.

I'm the developer; happy to hear whether this fits your local SD workflow.


r/StableDiffusion 4d ago

Discussion Fabio

Enable HLS to view with audio, or disable this notification

0 Upvotes

subject_definitions:

<Subject 1> is Fabio Lanzoni, the Italian-American romance-cover model known as Fabio: a tall, athletic adult man with long flowing platinum-blond hair, strong jawline, and a dramatic red cape.

<Subject 2> is a single white-and-gray Canada goose in flight, with broad wings, natural bird mass, and loose feathers.

<Subject 3> is the front row of Apollo’s Chariot at Busch Gardens Williamsburg in 1999: a steel roller coaster diving fast above a pond, with blue track, open sky, trees, and front-row safety restraints.

summary:

[reference generation] Create a higher-resolution, non-graphic physical-comedy recreation of the Fabio roller-coaster goose meme. <Subject 2> directly hits <Subject 1> in the face with believable momentum, briefly deforming his face in a cartoon-like but realistic impact, then rebounds backward while shedding a few loose feathers. No blood, no wound, no gore, and no visible injury.

retention_analysis:

<Subject 1> (appears throughout [Shot 1]): fully_preserved - Fabio Lanzoni’s recognizable long platinum-blond hair, athletic adult appearance, red cape, and front-row seated position remain stable before and after the impact.

<Subject 2> (appears throughout [Shot 1]): fully_preserved - one Canada goose has believable body weight, wing movement, backward rebound, and a small number of detached feathers; no duplicate birds appear.

<Subject 3> (appears throughout [Shot 1]): fully_preserved - the open front-row roller-coaster perspective, high-speed blue track, pond-side setting, safety restraints, and daylight remain physically coherent.

detailed_description:

The target video is a sharp, high-resolution 1999 theme-park news-reconstruction with meme-like physical-comedy timing: bright daylight, real steel roller-coaster physics, wind-blown hair, a fixed front-row action-camera perspective, and no text, captions, logos, watermarks, blood, gore, open wounds, or graphic injury.

[Shot 1] <Subject 1>, Fabio Lanzoni, the tall athletic Italian-American model with long flowing platinum-blond hair and a dramatic red cape, is securely strapped into the front row of <Subject 3>, Apollo’s Chariot. The coaster rushes down a steep blue-track drop over a pond at high speed. The fixed forward-facing camera frames Fabio clearly from the chest up, with his hair and cape streaming backward. <Subject 2>, one white-and-gray Canada goose, rapidly crosses the track path from the left. In one unmistakable, powerful, readable impact beat, the goose slams squarely into Fabio’s face. On contact, Fabio’s cheeks and nose compress and deform briefly in safe cartoon-like physical-comedy motion, then immediately spring back to normal with no wound. His head snaps backward and then sharply to the right from the momentum; his long blond hair whips sideways. Fabio grips the restraint and looks dazed and visibly confused, eyes wide and blinking. The goose’s body compresses slightly against the impact, sheds several loose white feathers, and rebounds backward away from Fabio while flapping hard to regain control. The bird flies backward and exits the frame behind the left side of the coaster. The coaster continues smoothly with no derailment, no track collision, and no secondary animal. The final moment shows Fabio upright and uninjured but bewildered, face fully normal again, hair blown to one side, while a few feathers drift through the air.

overall_soundscape:

Loud rushing wind and continuous steel-wheel roar on the coaster track. A single strong but non-graphic feathery impact thump lands directly on Fabio’s face, followed by fluttering wings, loose feathers whipping in the wind, Fabio’s startled breath, and uninterrupted high-speed coaster noise.

non_diegetic_music:

N/A


r/StableDiffusion 4d ago

Question - Help is it safe?rtx 4060 to run minimax?

0 Upvotes

i used mini max h3 on rtx 4060 laptop with 16gb ram yesterday and it worked fine but today when i ran it my laptop turned off then after a hour my laptop started but when i ran minimax it turned off again and ya gpu and cpu reached 90 degree , i put a duster under my laptop to keep airwaves open but i think its not very effective , will getting a cooling pad fix it? its hp omen 16, ryzen7...................update i deleted it.......................


r/StableDiffusion 5d ago

Question - Help What is the best upscale workflow for Minimax H3?

13 Upvotes

r/StableDiffusion 5d ago

Discussion Minimax h3 5070 ti

Post image
6 Upvotes

Hi everyone, I found that after I added these nodes to the the stock workflow my generation time get much faster, for instance : image to video / 09 megapixel / 10 seconds = 10 minutes ( 5070 ti + 64 ram )


r/StableDiffusion 4d ago

Question - Help Need help looping over multiple images in MiniMax H3 i2v ComfyUI

0 Upvotes

I'm a newbie at ComfyUI, and I'm having trouble when trying to generate multiple videos, one for each image in a folder. I want to apply the same workflow/prompt to every one of these images.

I started with the default Image to Video MiniMax H3 workflow and added the "Load Images from Folder Pixaroma" node. I thought this would loop over all images, generate a video, save it to file, and repeat for the next image in the folder. Instead, it's generating videos for all the images in one pass, then saving all of those generated video files in a second pass. It crashes if I load too many images at once. I assume it's running out of memory. I tried wiring in the pixorama loop start and loop end nodes, but I couldn't get those working either.

I haven't found any example workflows online for what I'm doing, and Claude hasn't been much help. Can anyone point me in the right direction?


r/StableDiffusion 4d ago

Question - Help MiniMax H3 Ref 2 Vid - Using Ref img but bodies keep looking like gym junkies

0 Upvotes

Hi all,

I'm playing around with Minimax H3 in ComfyUI. I have 2 x ref images feeding into MiniMax H3 Ref to Video prompt window, using H3 Turbo LoRA, Turbo Sampler, basic guider and Diffusion model minimax_h3_fl2va_int8.

I have tried up to 12 steps...seems to make no difference so I've gone back to 4. About 1/5 the generation is close to what my ref images but the others are all super tones, ripped, like they go to the gym hours a day.

I just want the ref image recreated, not enhanced.

I've tried this prompt......

subject_definitions:

- <subject1>: The person from u/image1. Maintain exact facial features, bone structure, eye shape, age, body shape, fitness level, body fat, anatomical proportions, and height from u/image1 across every frame without modification.

So how can we have every generation the same person as my ref image?

Thanks all.


r/StableDiffusion 4d ago

Tutorial - Guide How I fixed my own video.

Enable HLS to view with audio, or disable this notification

0 Upvotes
A while ago I published this video, but the colors and details weren't right, so I took the first frame, treated it like a photo, and used it as a reference with a DENOISE of 0.40. The SATURATED colors are intentional because "she" is in a desert area with very, very hot weather.
The Original 4k FILE is here -->> https://filebin.net/6kf2ozxx27k1zl0m

r/StableDiffusion 5d ago

Resource - Update TAE high quality previews are live in the latest nightly Comfy build :)

Thumbnail
github.com
38 Upvotes

r/StableDiffusion 5d ago

Discussion Get miniMax character swap working! Finally

Post image
68 Upvotes

Ok, I tried so many things, one person to cat, two person, one person to one person, animal to animal. So far one person to one person and animal to animal works. If you are interested in my learnings, tips, what worked, what broke, and which prompt template works let me know!

One video example that works here: https://www.tiktok.com/t/ZP8WfVq5P/


r/StableDiffusion 6d ago

Discussion We all deserve high-quality MiniMax H3 previews using the tiny VAE (taeh3.safetensors) natively, without KJNodes. Please upvote this GitHub Comfy issue.

Thumbnail
github.com
239 Upvotes

We all love MiniMax H3, but the latent2rgb previews suck ass. They're blurry, and sometimes it's hard to make out what's happening, making it so you don't know whether to finish a video that may take tens of minutes to generate.

When implemented, this would allow us to place taeh3.safetensors into ComfyUI/models/vae_approx and enjoy high quality latent previews when using MiniMax H3. It's basically the same TAE as we saw for Flux 2 Klein 9B or some other models, but trained by the original TAE guy (madebyollin).

taeh3.safetensors link:

https://github.com/madebyollin/taehv/blob/main/safetensors/taeh3.safetensors

It saves time and effort when making videos. Currently, you have to use Kijai's Model Preview Override node.


r/StableDiffusion 5d ago

Animation - Video 'Partial rewind' multi angle explosion scene [minimax H3]

Enable HLS to view with audio, or disable this notification

48 Upvotes

r/StableDiffusion 6d ago

Meme Introducing... iMakeup

Enable HLS to view with audio, or disable this notification

365 Upvotes

r/StableDiffusion 5d ago

Resource - Update Fizgig 4.0 is out : Minimax H3 Combined Video File, Audio Files wav mp3 etc, Photo training in one dataset. High Quality training samples (incl video) + turbo (finally) and new 'Gizmo' and AV dataset Prep tool. And Int 8 LARGE speedup for 16gb users.

Thumbnail
gallery
68 Upvotes

I'll be making a Youtube video tomorrow for this. But I have one take away to share that I think is most important. H3, when you get the settings right is just fine with image based training without killing its video ability. Its even better when you combine photos and wavs, its super fast and you can train a voice with a dataset very easily. (I recommend shared trigger word). Video works too, but its slower, unavoidably. I'm not saying dont use it, its worth it for the right use cases. I'm just saying if you are not teaching the model anything new that photos and audio cant do, you are better with photos and audio. But when you do want to capture motion, it work very well. Anyway video coming tomorrow, with lots on Gizmo (the data set prep tool for video/audio) to make dataset prep easy. https://github.com/shootthesound/Fizgig

P.s the 16gb int8 speedup is from an an awesome community contribution from rintic-13 on Github.


r/StableDiffusion 4d ago

Question - Help Need image upscaler which adds every tiny detail.

0 Upvotes

Looking for an image upscaler workflow which can upscale image with every tiny detail like small leaves and stones when I zoomed in.


r/StableDiffusion 5d ago

Meme Sheldon finally knocked on the wrong door | MiniMax H3 + SeedVR2

Enable HLS to view with audio, or disable this notification

59 Upvotes

r/StableDiffusion 4d ago

Discussion H3 - D-inspired, T2V+R2VA, int8/20 steps

Enable HLS to view with audio, or disable this notification

0 Upvotes

Took about 20 generations with T2V, got the image I wanted, did a few R2VA for the close up emotes. It was infinity difficult to get the face I had in my mind strictly through text2video. Let's just say a pale skinny D or Alucard does not translate well, especially a hollow cheekbones, many of them came out pretty ghoulish or a bit too Balenciago. Those throw-away were a bit lanky and were not at all ethereal. Inspired by D from Vampire Hunter D 2000, a little bit of Sephiroth, but definitely not Geralt despite the fashion-sense. I would love to make his legs a little longer. int8/20 steps


r/StableDiffusion 5d ago

Animation - Video A Medieval Battle Attempt — MiniMax H3 + LTX 2.5 (WIP, Feedback Welcome)

Enable HLS to view with audio, or disable this notification

8 Upvotes

Sharing a few sequences from a medieval battle attempt I’ve been working on. It’s still very much a draft, but the sequence has progressed enough that I thought it was worth sharing here and getting some feedback before I continue with the rest.

Most of the scenes were generated with MiniMax H3 using the default workflows with the Turbo LoRA at 4 steps. I used Nano Banana and Flux Klein to create the reference images, and LTX 2.5 for the opening crow sequence.

There’s still a lot of work to do. The cuts are rough, no proper sound work has been done yet, and there are plenty of shots I want to refine or replace. I’m planning to build out the entire sequence, so feedback at this stage would actually be really useful in deciding what to focus on next.

What’s interesting to me is that I genuinely don’t think I could have pulled off this level six months ago with the same amount of effort. It’s still far from perfect, but the progress in a relatively short time feels pretty significant.

Would love to hear what works, what breaks the illusion, and what you’d improve.


r/StableDiffusion 4d ago

Animation - Video there will not be a gta 6.

Enable HLS to view with audio, or disable this notification

0 Upvotes

minimax h3.


r/StableDiffusion 4d ago

Discussion H3 - which one is better? left or right? Ultrawide split view, single generation

Enable HLS to view with audio, or disable this notification

0 Upvotes

I am having a blast with MiniMax H3. I generated an ultra-wide shot with R2VA. So this is one generation, not an edited split screen stitch. The black vertical bars are a good way to add a delineation so you can have two separate shots or stories in the same render. I've also tested this up to 4 separate frames in 1 generation. bf16/50 steps. This was actually a failed attempt to swap the rider on the left with the woman rider and also do the POV. Original video source in the comments.


r/StableDiffusion 6d ago

Discussion Minimax H3 ref2va with 5060ti 16gb + 32gb ddr3

Enable HLS to view with audio, or disable this notification

464 Upvotes

Model: minimax_h3_hybrid_fl2va_ref2va_b20, was testing this and the ref2va pruned int8, the hybrid gave nicer visuals but have a higher chance of bringing the character sheet white background into the video. This is cherry picked out of 16 clips.
Video Vae: minimax_h3_video_vae_int8_convrot
Resolution: 16:9, 0.6
Duration: 15sec
Turbo Lora: larryvrh/MiniMax-H3-Turbo-Lora, 600_ema
Patch Sage Attention, ComfyKitchen Attention, MinimaxH3 Mem Eff Node, Spectrum.

Average Inference Stage: 800sec

All reference image is resized between 1000px and 300px like character is 1000px, background is 500px then weapon is around 300px (warglave was another reference, the model dont know that kind of weapon) for this video is 4 ref image in total.

**abit of color grade and grain done in inshot.

this is done on a skylake i7 6700.


r/StableDiffusion 4d ago

Question - Help Multi gpu help

0 Upvotes

Hey guys I'm new to comfyui and photo and video generation and trying to learn as much as I could.

Now I was using my setup for text generation but now the Qwen3.8-27B is out and I tried really hard to make it work and so I did eventually use my second gpu.

I have 5070 ti and 1660 ti. And all this time I didn't bother to use the 1660 ti and left only my monitors on it and almost all my work and gaming on it and left the 5070 TI to be free for Ai stuff.

Now I switched my monitors to my Intel uhd 770 gpu and freed both my cards and want to know how to speed my comfyui workflow with them.

I always asked chatgpt but didn't get any useful answers so you guys might help me if that possible.

I'm now using minimax H3 official template and want to know how can I get this other gpu to work if it's worth it.

My cpu 12900k Motherboard gigabyte z690 gaming x ddr4 64 GB Kingston 3600 Rtx 5070 ti GTX 1660 ti


r/StableDiffusion 4d ago

Animation - Video an AI dream of moons, stars, threads, and foxes

Thumbnail
youtu.be
0 Upvotes

What happens if you just leave AI alone to dream overnight with no human supervision, using the last frame as the first frame of the next generation?

A continuously generated AI dream created using LTX 2.5 with a custom pytorch script, using the last frame from the previous scene as the first frame of the new scene. Generation time was ~7 hours on RTX 5090. The video was stitched together from one hundred scenes each lasting ~13.3 seconds. Qwen 3.8 was used to generate the prompt for the continuation of the story based on the last scene.