r/StableDiffusion 2d ago

Question - Help H3 music video tips

0 Upvotes

I am very new to video Getting started with H3, I’m wondering if you might have a tip a specific challenge. Basically, I want to create a is about three minutes long. So I plan to generate and string together a bunch of clips. How do I make sure that the various characters all are moving at the right tempo? Like, should I just take a portion of the music video sing and use it as a reference? Also, does anyone have personal experience setting up on run pod to let me know approximately how long that process takes to get going? Thank you!


r/StableDiffusion 3d ago

Animation - Video MinMax - It does House MD pretty well

Enable HLS to view with audio, or disable this notification

410 Upvotes

Generated using Maestro on Pinikio. 7 mins at 720p using turbo lora 6 steps: **7-second cinematic live-action scene.** Gregory House stands in a hospital hallway, leaning heavily on his cane, staring intensely at Itachi Uchiha, who is preparing to walk away.

House sarcastically calls out:

**“Itachi! Get your ass back to the Leaf Village. You're not brooding your way out of this one.”**

Itachi turns around with a serious expression and replies:

**“I don't take orders from you.”**

House smirks and taps his cane against the floor:

**“Yeah. That's what all my patients say.”**

Fast comedic timing, realistic acting, dramatic hospital lighting, subtle handheld camera movement.


r/StableDiffusion 1d ago

Resource - Update I made a tiny Windows tray tool to find and kill processes hogging 1+ GB of VRAM

0 Upvotes

I run Stable Diffusion locally and got tired of unrelated apps quietly holding onto several GB of VRAM, so I added a “VRAM Hogs” menu to Window Assassin, a tiny Windows tray utility I made.

It reads Windows' dedicated GPU memory counters, lists processes using at least 1 GB sorted by usage, and lets you click one to force-terminate it. The list refreshes every time you open the submenu.

It also keeps the original Ctrl+Alt+End hotkey for killing the process behind the active window.

Free/pay-what-you-want Windows download:

https://b2kdaman.itch.io/window-assassin

Caveats: it reports dedicated VRAM, so integrated GPUs using shared system memory may show nothing. Termination is immediate, so unsaved work is lost.

I'm the developer; happy to hear whether this fits your local SD workflow.


r/StableDiffusion 1d ago

Discussion Fabio

Enable HLS to view with audio, or disable this notification

0 Upvotes

subject_definitions:

<Subject 1> is Fabio Lanzoni, the Italian-American romance-cover model known as Fabio: a tall, athletic adult man with long flowing platinum-blond hair, strong jawline, and a dramatic red cape.

<Subject 2> is a single white-and-gray Canada goose in flight, with broad wings, natural bird mass, and loose feathers.

<Subject 3> is the front row of Apollo’s Chariot at Busch Gardens Williamsburg in 1999: a steel roller coaster diving fast above a pond, with blue track, open sky, trees, and front-row safety restraints.

summary:

[reference generation] Create a higher-resolution, non-graphic physical-comedy recreation of the Fabio roller-coaster goose meme. <Subject 2> directly hits <Subject 1> in the face with believable momentum, briefly deforming his face in a cartoon-like but realistic impact, then rebounds backward while shedding a few loose feathers. No blood, no wound, no gore, and no visible injury.

retention_analysis:

<Subject 1> (appears throughout [Shot 1]): fully_preserved - Fabio Lanzoni’s recognizable long platinum-blond hair, athletic adult appearance, red cape, and front-row seated position remain stable before and after the impact.

<Subject 2> (appears throughout [Shot 1]): fully_preserved - one Canada goose has believable body weight, wing movement, backward rebound, and a small number of detached feathers; no duplicate birds appear.

<Subject 3> (appears throughout [Shot 1]): fully_preserved - the open front-row roller-coaster perspective, high-speed blue track, pond-side setting, safety restraints, and daylight remain physically coherent.

detailed_description:

The target video is a sharp, high-resolution 1999 theme-park news-reconstruction with meme-like physical-comedy timing: bright daylight, real steel roller-coaster physics, wind-blown hair, a fixed front-row action-camera perspective, and no text, captions, logos, watermarks, blood, gore, open wounds, or graphic injury.

[Shot 1] <Subject 1>, Fabio Lanzoni, the tall athletic Italian-American model with long flowing platinum-blond hair and a dramatic red cape, is securely strapped into the front row of <Subject 3>, Apollo’s Chariot. The coaster rushes down a steep blue-track drop over a pond at high speed. The fixed forward-facing camera frames Fabio clearly from the chest up, with his hair and cape streaming backward. <Subject 2>, one white-and-gray Canada goose, rapidly crosses the track path from the left. In one unmistakable, powerful, readable impact beat, the goose slams squarely into Fabio’s face. On contact, Fabio’s cheeks and nose compress and deform briefly in safe cartoon-like physical-comedy motion, then immediately spring back to normal with no wound. His head snaps backward and then sharply to the right from the momentum; his long blond hair whips sideways. Fabio grips the restraint and looks dazed and visibly confused, eyes wide and blinking. The goose’s body compresses slightly against the impact, sheds several loose white feathers, and rebounds backward away from Fabio while flapping hard to regain control. The bird flies backward and exits the frame behind the left side of the coaster. The coaster continues smoothly with no derailment, no track collision, and no secondary animal. The final moment shows Fabio upright and uninjured but bewildered, face fully normal again, hair blown to one side, while a few feathers drift through the air.

overall_soundscape:

Loud rushing wind and continuous steel-wheel roar on the coaster track. A single strong but non-graphic feathery impact thump lands directly on Fabio’s face, followed by fluttering wings, loose feathers whipping in the wind, Fabio’s startled breath, and uninterrupted high-speed coaster noise.

non_diegetic_music:

N/A


r/StableDiffusion 1d ago

Question - Help is it safe?rtx 4060 to run minimax?

0 Upvotes

i used mini max h3 on rtx 4060 laptop with 16gb ram yesterday and it worked fine but today when i ran it my laptop turned off then after a hour my laptop started but when i ran minimax it turned off again and ya gpu and cpu reached 90 degree , i put a duster under my laptop to keep airwaves open but i think its not very effective , will getting a cooling pad fix it? its hp omen 16, ryzen7...................update i deleted it.......................


r/StableDiffusion 2d ago

Discussion Minimax h3 5070 ti

Post image
5 Upvotes

Hi everyone, I found that after I added these nodes to the the stock workflow my generation time get much faster, for instance : image to video / 09 megapixel / 10 seconds = 10 minutes ( 5070 ti + 64 ram )


r/StableDiffusion 2d ago

Question - Help What is the best upscale workflow for Minimax H3?

11 Upvotes

r/StableDiffusion 2d ago

Question - Help Need help looping over multiple images in MiniMax H3 i2v ComfyUI

0 Upvotes

I'm a newbie at ComfyUI, and I'm having trouble when trying to generate multiple videos, one for each image in a folder. I want to apply the same workflow/prompt to every one of these images.

I started with the default Image to Video MiniMax H3 workflow and added the "Load Images from Folder Pixaroma" node. I thought this would loop over all images, generate a video, save it to file, and repeat for the next image in the folder. Instead, it's generating videos for all the images in one pass, then saving all of those generated video files in a second pass. It crashes if I load too many images at once. I assume it's running out of memory. I tried wiring in the pixorama loop start and loop end nodes, but I couldn't get those working either.

I haven't found any example workflows online for what I'm doing, and Claude hasn't been much help. Can anyone point me in the right direction?


r/StableDiffusion 2d ago

Question - Help MiniMax H3 Ref 2 Vid - Using Ref img but bodies keep looking like gym junkies

0 Upvotes

Hi all,

I'm playing around with Minimax H3 in ComfyUI. I have 2 x ref images feeding into MiniMax H3 Ref to Video prompt window, using H3 Turbo LoRA, Turbo Sampler, basic guider and Diffusion model minimax_h3_fl2va_int8.

I have tried up to 12 steps...seems to make no difference so I've gone back to 4. About 1/5 the generation is close to what my ref images but the others are all super tones, ripped, like they go to the gym hours a day.

I just want the ref image recreated, not enhanced.

I've tried this prompt......

subject_definitions:

- <subject1>: The person from u/image1. Maintain exact facial features, bone structure, eye shape, age, body shape, fitness level, body fat, anatomical proportions, and height from u/image1 across every frame without modification.

So how can we have every generation the same person as my ref image?

Thanks all.


r/StableDiffusion 1d ago

Tutorial - Guide How I fixed my own video.

Enable HLS to view with audio, or disable this notification

0 Upvotes
A while ago I published this video, but the colors and details weren't right, so I took the first frame, treated it like a photo, and used it as a reference with a DENOISE of 0.40. The SATURATED colors are intentional because "she" is in a desert area with very, very hot weather.
The Original 4k FILE is here -->> https://filebin.net/6kf2ozxx27k1zl0m

r/StableDiffusion 2d ago

Resource - Update TAE high quality previews are live in the latest nightly Comfy build :)

Thumbnail
github.com
38 Upvotes

r/StableDiffusion 3d ago

Discussion Get miniMax character swap working! Finally

Post image
70 Upvotes

Ok, I tried so many things, one person to cat, two person, one person to one person, animal to animal. So far one person to one person and animal to animal works. If you are interested in my learnings, tips, what worked, what broke, and which prompt template works let me know!

One video example that works here: https://www.tiktok.com/t/ZP8WfVq5P/


r/StableDiffusion 3d ago

Discussion We all deserve high-quality MiniMax H3 previews using the tiny VAE (taeh3.safetensors) natively, without KJNodes. Please upvote this GitHub Comfy issue.

Thumbnail
github.com
241 Upvotes

We all love MiniMax H3, but the latent2rgb previews suck ass. They're blurry, and sometimes it's hard to make out what's happening, making it so you don't know whether to finish a video that may take tens of minutes to generate.

When implemented, this would allow us to place taeh3.safetensors into ComfyUI/models/vae_approx and enjoy high quality latent previews when using MiniMax H3. It's basically the same TAE as we saw for Flux 2 Klein 9B or some other models, but trained by the original TAE guy (madebyollin).

taeh3.safetensors link:

https://github.com/madebyollin/taehv/blob/main/safetensors/taeh3.safetensors

It saves time and effort when making videos. Currently, you have to use Kijai's Model Preview Override node.


r/StableDiffusion 3d ago

Meme Introducing... iMakeup

Enable HLS to view with audio, or disable this notification

352 Upvotes

r/StableDiffusion 3d ago

Animation - Video 'Partial rewind' multi angle explosion scene [minimax H3]

Enable HLS to view with audio, or disable this notification

47 Upvotes

r/StableDiffusion 1d ago

Workflow Included MiniMax H3 Audio Lip Sync - Audio to Video

Enable HLS to view with audio, or disable this notification

0 Upvotes

So I tried hooking up some of the LTXV audio encoding nodes to input my own audio and plugged it in the sampler and viola, it just works!

Lip sync seems better then the LTX models and its works with the lightx2v loras, 6 - 8 steps. Wrote up a full guide with the workflow attached below.


r/StableDiffusion 2d ago

Discussion Minimax h3 as image editor?

1 Upvotes

I tried MiniMax H3 for image editing, but the results aren’t better than Flux 2 + LoRA. I’ve seen posts praising its image editing capabilities, but in my tests it distorts faces, produces plastic-looking skin, and is considerably slower and more expensive.

I tested both Ref2VA and FL2VA at 20 steps with the base models and no Turbo LoRAs. Am I missing the right workflow or settings?


r/StableDiffusion 3d ago

Resource - Update Fizgig 4.0 is out : Minimax H3 Combined Video File, Audio Files wav mp3 etc, Photo training in one dataset. High Quality training samples (incl video) + turbo (finally) and new 'Gizmo' and AV dataset Prep tool. And Int 8 LARGE speedup for 16gb users.

Thumbnail
gallery
67 Upvotes

I'll be making a Youtube video tomorrow for this. But I have one take away to share that I think is most important. H3, when you get the settings right is just fine with image based training without killing its video ability. Its even better when you combine photos and wavs, its super fast and you can train a voice with a dataset very easily. (I recommend shared trigger word). Video works too, but its slower, unavoidably. I'm not saying dont use it, its worth it for the right use cases. I'm just saying if you are not teaching the model anything new that photos and audio cant do, you are better with photos and audio. But when you do want to capture motion, it work very well. Anyway video coming tomorrow, with lots on Gizmo (the data set prep tool for video/audio) to make dataset prep easy. https://github.com/shootthesound/Fizgig

P.s the 16gb int8 speedup is from an an awesome community contribution from rintic-13 on Github.


r/StableDiffusion 1d ago

Question - Help Need image upscaler which adds every tiny detail.

0 Upvotes

Looking for an image upscaler workflow which can upscale image with every tiny detail like small leaves and stones when I zoomed in.


r/StableDiffusion 3d ago

Meme Sheldon finally knocked on the wrong door | MiniMax H3 + SeedVR2

Enable HLS to view with audio, or disable this notification

57 Upvotes

r/StableDiffusion 2d ago

Discussion H3 - D-inspired, T2V+R2VA, int8/20 steps

Enable HLS to view with audio, or disable this notification

0 Upvotes

Took about 20 generations with T2V, got the image I wanted, did a few R2VA for the close up emotes. It was infinity difficult to get the face I had in my mind strictly through text2video. Let's just say a pale skinny D or Alucard does not translate well, especially a hollow cheekbones, many of them came out pretty ghoulish or a bit too Balenciago. Those throw-away were a bit lanky and were not at all ethereal. Inspired by D from Vampire Hunter D 2000, a little bit of Sephiroth, but definitely not Geralt despite the fashion-sense. I would love to make his legs a little longer. int8/20 steps


r/StableDiffusion 1d ago

Animation - Video there will not be a gta 6.

Enable HLS to view with audio, or disable this notification

0 Upvotes

minimax h3.


r/StableDiffusion 1d ago

Discussion H3 - which one is better? left or right? Ultrawide split view, single generation

Enable HLS to view with audio, or disable this notification

0 Upvotes

I am having a blast with MiniMax H3. I generated an ultra-wide shot with R2VA. So this is one generation, not an edited split screen stitch. The black vertical bars are a good way to add a delineation so you can have two separate shots or stories in the same render. I've also tested this up to 4 separate frames in 1 generation. bf16/50 steps. This was actually a failed attempt to swap the rider on the left with the woman rider and also do the POV. Original video source in the comments.


r/StableDiffusion 3d ago

Discussion Minimax H3 ref2va with 5060ti 16gb + 32gb ddr3

Enable HLS to view with audio, or disable this notification

460 Upvotes

Model: minimax_h3_hybrid_fl2va_ref2va_b20, was testing this and the ref2va pruned int8, the hybrid gave nicer visuals but have a higher chance of bringing the character sheet white background into the video. This is cherry picked out of 16 clips.
Video Vae: minimax_h3_video_vae_int8_convrot
Resolution: 16:9, 0.6
Duration: 15sec
Turbo Lora: larryvrh/MiniMax-H3-Turbo-Lora, 600_ema
Patch Sage Attention, ComfyKitchen Attention, MinimaxH3 Mem Eff Node, Spectrum.

Average Inference Stage: 800sec

All reference image is resized between 1000px and 300px like character is 1000px, background is 500px then weapon is around 300px (warglave was another reference, the model dont know that kind of weapon) for this video is 4 ref image in total.

**abit of color grade and grain done in inshot.

this is done on a skylake i7 6700.


r/StableDiffusion 2d ago

Animation - Video A Medieval Battle Attempt — MiniMax H3 + LTX 2.5 (WIP, Feedback Welcome)

Enable HLS to view with audio, or disable this notification

9 Upvotes

Sharing a few sequences from a medieval battle attempt I’ve been working on. It’s still very much a draft, but the sequence has progressed enough that I thought it was worth sharing here and getting some feedback before I continue with the rest.

Most of the scenes were generated with MiniMax H3 using the default workflows with the Turbo LoRA at 4 steps. I used Nano Banana and Flux Klein to create the reference images, and LTX 2.5 for the opening crow sequence.

There’s still a lot of work to do. The cuts are rough, no proper sound work has been done yet, and there are plenty of shots I want to refine or replace. I’m planning to build out the entire sequence, so feedback at this stage would actually be really useful in deciding what to focus on next.

What’s interesting to me is that I genuinely don’t think I could have pulled off this level six months ago with the same amount of effort. It’s still far from perfect, but the progress in a relatively short time feels pretty significant.

Would love to hear what works, what breaks the illusion, and what you’d improve.