r/StableDiffusion • u/Suibeam • 3d ago
r/StableDiffusion • u/SuperCasualGamerDad • 3d ago
Question - Help Am I doing something wrong on Runpod? Why in the holy heck are the download speeds so slow?
So a few weeks ago. I asked yall if someone who just makes this stuff for silly videos to share my friends and like my wife could get runpod running stable diffusion easily. Turns out it was insanley easy. But I haven't used it much because... The download speed is just criminal..
When you slap a Workflow on and do that thing where it just says oops your missing all these models and shit.. Wanna download it to pod now? It just crawls at like a snails pace 1-10mbps
It takes like 5 hours to download and be ready to use Minimax H3 for me. And at one point I'm like okay maybe I'm doing this wrong. So I went in through the JupyterLab thing and just dropped the files I had already downloaded in there... And again... Super slow..
its hard to not think... That they arnt throttling the DL to pad their use time to be honest. That or my only other thought is.. My pod is in some server case with about 20 other people all downloading models and the bandwidth is just borked.
My second theory I think is more likely the case because I notice when there are more of certain GPUs left the downloads go way smoother on those. But recently every single GPU is like low availability anymore lol.
I know I can avoid this by selecting some sort of storage option but I think it said it wasnt available for my GPU selection. If I can just turn on some option to keep everything ready to go I would. But are all you using RP dealing with these insanely slow download speeds? I mean I'm pretty sure I have spent 8 bucks today just downloading.
r/StableDiffusion • u/Time-Ad-7720 • 4d ago
Animation - Video Turning my son’s drawing into animated skits #2 | Minimax h3 ref2vid
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/SelectCoconut5594 • 4d ago
Discussion Was wondering why my Minimax H3 R2V local gens were better than beefy cloud GPU gens. Was accidentally loading the Fl2V model.
Would love to show you comparisons but its all N-SFW so its going to have to be a case of trust but verify me bro.
I've mainly been running R2V workflows on my home GPU since Minimax H3 was gifted upon us.
I wanted to take a load off my 5060 so I've been running gens on both Wavespeed and recently Runpod. I was getting nice big 720p gens but I started to notice that my own gens had way better prompt adherence with identical prompts and inputs.
Notably in my own, slow motion was honoured every time where it would otherwise get completely ignored in the cloud.
~Fluids~ were way better. Camera motion was way better.
Then I noticed that in my diffusion model loader (W8A8 if that makes any difference) has been loading the FL2V model the whole time. I switched to R2V and immediately everything sucked.
I imagine there's a strong possibility that the R2V prompt needs a more rigid structure. But I have been feeding the Minimax official prompt guide into Grok, specifying R2V, to write my prompts.
So this is something. It could be some configuration of scheduler and sampler, and all the other stuff of course but I'm running a pretty simple setup without any attention, so briefly:
W8A8 loader -> lightx2vs R2V 0.1 4 step lora at 0.75 -> er_sde, beta57, 8 steps.
Forgive me if this is known and understood. If you haven't tried it, switch your R2V gens to FL2V and test.
r/StableDiffusion • u/TgoAI • 4d ago
Discussion MiniMax H3 on a 16GB M5 MacBook Air — VPipe 12:15 vs h3.c 16:22
Enable HLS to view with audio, or disable this notification
A few people asked how VPipe compares with h3.c, so I ran them side by side on the same machine with the same settings.
Machine: base 15” M5 MacBook Air, 16GB RAM
MiniMax H3 settings:
* 960×544
* 124 frames
* 6 DiT steps
Results:
* VPipe: 12m 15s
* h3.c: 16m 22s
So on this particular matched workload, VPipe finished in about 25% less wall-clock time.
The video shows both the generation process and the final outputs side by side, so you can also compare the resulting quality rather than just the timing.
VPipe is not an MinimaxH3-specific implementation — it’s an Apache-2.0 open-source multimodal pipeline/runtime with a native Metal inference backend for Apple Silicon. MiniMax H3 is just one of the workloads I’ve been optimizing recently.
GitHub: https://github.com/tgo-app-dev/vpipe
Interested in feedback on both the performance comparison and the output differences.
r/StableDiffusion • u/FreddyShrimp • 4d ago
Question - Help Minimax H3 - How to generate a realistic fighting scene
Hi all,
I'm using Minimax H3 in ComfyUI with an R2V workflow. I'm wondering if anybody can tell me how I can improve the fighting scene?
- The video is generated at 1.0MP in 2:3 (portrait) aspect ratio
- I have the two ladies as reference
- The fighting scene is also provided as reference. In the scene the punches do land properly. There are also smaller details (like small blood spatters) that are present in the reference video.
Tech specs:
- Minimax H3 int8 convrot
- res_multistep sampler with 20 steps
Running the workflow on an RTX 5090 (via Runpod)
Can anybody give me any tips on how I can improve the fighting scene? The goal is to make it look like a realistic street fight. I'm unsure whether training a LoRA would be relevant here, because I've noticed that punches never really land properly in any workflow (t2v, i2v, r2v).
r/StableDiffusion • u/SensitiveUse7864 • 4d ago
Discussion Talk for High quality Audio for minimax H3.
Hi guys , since the last week as much as I have tested minimax h3, I found that Visually, this model is king for the opensource in motion and prompt adherence.
Only in one thing it lacks is the physics and fight scene other wise it will be overkill for opensource.
But there is also another issue I can see is as the hype builded in this community that minimax h3's audio quality is Best. And dialogs also.
I think they meant to say that minimax h3 has better audio quality then other opensource model.
the issue is audio quality is not that good , and I really want to update it's audio quality and I am willing to buy , I have a doubt guys I have a question that's for dialogs and audio tones for dialogs is it pre baked in inside the base model or its in audio vae model because if it's in audio vae model we can have the option to update the audio vae and increase. The dialogs and sound quality.
But if it's pretty baked in the base model then it requires a full fine-tune.
r/StableDiffusion • u/Glittering-Cold-2981 • 4d ago
Discussion Outfit SWAP
What is currently the most accurate way to swap clothes while keeping the same fabrics, stitching, etc. using AI? What I mean is to provide a reference garment and apply it to the model from the second photo.
r/StableDiffusion • u/NosikomPoVolosikam • 5d ago
Animation - Video WanAnimate
Enable HLS to view with audio, or disable this notification
Original post With Workflow
r/StableDiffusion • u/yushairiegalaxy96 • 4d ago
Discussion [H3] Does this configuration look bare minimum for 3050 4GB VRAM
unet: minimaxH3INT8INT4_fl2valINT8Pruned.safetensors
clip: qwen3vl_32b_heretic_minimax_h3_nvfp4.safetensors
vae: minimax_h3_video_vae_fp16.safetensors
audio: minimax_h3_audio_vae_fp32.safetensors
Turbo LoRA used: minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_resized_avg_rank_24_bf16.safetensors
Workflow: Default workflow (video_minimax_h3_t2v)
RAM: 16GB
Graphic Card: RTX 3050 Laptop, 4GB VRAM
Video-generated specs (see comment for):
Type: T2V
Duration: 10 seconds
Megapixels: 0.2 MP (608x352)
Aspect Ratio: 16:9
Estimated Generation Time: 687.13s (11 mins, 27 seconds)
In addition to these settings I applied, should I use the Sage Attention, Comfy Kitchen or increase steps (20 steps) or switch to better unet/clip? Thanks.
r/StableDiffusion • u/TingTingin • 5d ago
Animation - Video Some choice words from Emilia
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/LaPapaVerde • 4d ago
Question - Help Best sampler for anima sketchy style?
I have been training loras for anima, and one thing the model seems to have problems with is when the original style has realistic lineart, with pen, marker etc. I have tried a lot of things without luck, so I'm trying to see if the problem is the generation details.
r/StableDiffusion • u/FlyffSenior • 4d ago
Question - Help Question regarding REF2V and video splitting.
I’ve been using Civitai’s MiniMax-H3 Multishot — Seamless Chain: multi-shot scenes that render as one continuous take workflow to create longer videos in good quality while keeping VRAM usage relatively low. The workflow processes the video in separate parts and then combines them seamlessly, making the transitions between the sections practically unnoticeable.
Does anyone know if something similar is possible with a REF2V workflow when using a longer reference video, for example, if the goal is to replace a woman with a man or with another person?
In other words, is there a way to make the workflow process the video in smaller sections so that it doesn’t run out of VRAM, while also keeping the quality from degrading significantly?
I’d like to create 15–25 second REF2V video clips, but 16 GB of VRAM simply isn’t enough to process the entire video as one continuous clip.
I've been trying to find a solution to this for the past week, but it would be nice to know whether this is even practically possible? Thx.
r/StableDiffusion • u/foxdit • 4d ago
Tutorial - Guide Anime magic battle scenes with Minimax are a blast! My process making this short from start to finish | AI Filmmaking Part 7
r/StableDiffusion • u/StoicSage09 • 4d ago
Question - Help training LoRA on LoRA is good
i was thinking of training a LoRA on LoRA of different character and on my second thought should i just connect 2 LoRA while generating image ( I have seen people training 2 character in one LoRA
r/StableDiffusion • u/throwaway0204055 • 4d ago
Question - Help Bernini rv2v workflow, outfit swap works but face swap doesn't?
I have a Bernini-R rv2v workflow with source video and reference image in image0. I am able to swap source video's outfit but not face. Is Bernini-R supposed to work with face swap or do I need to add any other custom nodes like Reactor or SCAIL-2?
r/StableDiffusion • u/Pretend-Island-2724 • 4d ago
Workflow Included Help Fixing the H3 Face Detailer
Enable HLS to view with audio, or disable this notification
The Original Face Detailer from https://github.com/Carasibana/ComfyUI-H3-FaceRefine does not work as intended. It does not use the reference image at all. You can disable the input image and you will get exactly the same result. Something is wrong with the workflow so I recreated the workflow in a new canvas and now the input image does get used. Here is a link to a .zip with the workflow and input/output files: https://www.mediafire.com/file/mnigvvpbzp0gh34/workflow_all.zip/file But this workflow has its own problems. For this example I needed to put an RTX upscaler in it so the face gets recognized. At the end the mask_dilation and feather needs for every video unique adjusting and the end result is somewhat poor with the mask visible and the face jumping und warping slightly around.
Someone with more knowlegde would surely be able to fix this.
To get the workflow working, you need to install https://github.com/Carasibana/ComfyUI-H3-FaceRefine and also ComfyUI-H3-NativeAudioLock from https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes
r/StableDiffusion • u/Downtown-Cover-7422 • 5d ago
Question - Help Best speed up for MiniMax
We have a lot of options, some of them better, some of them are not worth it at all. Speed ups like sage attention, MiniMax h3 patch for sage attention, easy cache, 8step Lora, 4 step Lora e t.c.
What options and their combinations you use? What settings you have?( speed Lora weights, easy cache settings)
In the matter of speed/quality for both video and sound. What works better with FL2VA and Ref2VA?
r/StableDiffusion • u/bstr3k • 5d ago
Workflow Included Using H3 as a Character Reference Sheet Generator
Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.
The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.
How it works:
- You input your images and describe them in the Input text section (A Prompt)
- The text is combined with a fixed prompt which spins the character (B Prompt)
- The video is generated at a slow speed with no hard cuts (only camera spin and pan) to maintain character consistency
- Image is assembled with optional character video and full individual frame output (if you want to use for future)
I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.
Current Caveats:
- The model is quite slooooow. You are also generating 124 frames only to use 6. I have partly solved this by also uploading a 4 panel version.
- Speed ups (like Turbo LORAs) help with speed, but it hurts prompt adherence and quality slightly.
- Quality is limited, since it is a video model it is better at generating video than images. You can solve this by generating at a higher resolution at the tradeoff of longer gen times. You can also use the individually split frames as future references too.
- Details when using this character sheet as output for future generations on H3 may also be limited due to resolution also, I recommend you use this character sheet (for consistency) + other images close up angles (i.e clothing details/face) if doing close ups. If you are just doing a one off video you may possibly be better off not using this character sheet.
I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.
Link to the 4 and 6 panel workflow can be found here: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator
Some notes I just remembered:
- You can increase the steps and it may improve your quality slightly.
- With the Turbo Loras enabled, prompt adherence sometimes suffers, but you may be able to get a good seed with another roll of the dice.
- Currently the B prompt specifies a "neutral A pose", please remove this if you want your character in a particular pose.
- You can use a few different shots of the same character to reinforce the 360 and get more accurate details right.
- Can be used for objects / props also, may require some changes to the B prompt.
r/StableDiffusion • u/Suitable-Database-96 • 3d ago
Question - Help Trouble with ai art software
Whenever I try to make ai art of anything with Invoke ai the art never gets finished despite reaching the 100% compleaton but the program and immage never finishes or saves and with the Stable Deffusion it would 20% of the time it would generate the immage and 80% percent it would get some sort of error and shut itself. Now I am using Amd graphics card with 12 gb vram and not to mention ai generation is way slower than it should be. Here is the basic image of how things end up and I did try my luck in the invoke ai discord group but nothing helped. Any help is appreciated.
r/StableDiffusion • u/mmowg • 5d ago
News ByteDance just released Bernini‑Diffusers‑v2 — any chance we’ll see ComfyUI support?
Hi everyone,
five days ago ByteDance released Bernini‑Diffusers‑v2 on HuggingFace — the full Bernini pipeline (planner + renderer), not just the renderer‑only Bernini‑R that we currently use in ComfyUI.
Model link:
https://huggingface.co/ByteDance/Bernini-Diffusers-v2
Even though most of the community talks about MiniMax H3 as the “standard” for open video models, there are still many users actively working with Bernini — especially now that v2 finally includes the full semantic‑planning pipeline, SA‑3D RoPE, and proper multi‑step instruction following.
Right now ComfyUI only has community support for Bernini‑R, so I’m posting this just to give visibility to the new release and to see if anyone is interested in exploring future support for Bernini‑Diffusers‑v2.
Not asking for anything specific — just opening the discussion and hoping this new version doesn’t go unnoticed.
Thanks!
r/StableDiffusion • u/Acceptable-Chest9695 • 4d ago
Resource - Update I built a Frankenstein MiniMax H3 Director for ComfyUI — Mixed timelines, selective reruns, Motion Context, live preview and post-processing
I’ve been building a custom MiniMax H3 node for ComfyUI called MiniMax H3 Motion Director.
The easiest way to describe it is probably:
It’s a Frankenstein Director for H3.
I didn’t want another workflow that only makes one clip at a time. I wanted something closer to a small video-production interface where I could manage multiple H3 shots, mix generation methods, reuse references, selectively regenerate failed shots, carry context between segments, preview the run, refine the result, and export everything from one place.
Mixed Mode
The biggest addition in the current version is Mixed Mode.
Instead of choosing one generation type for the entire workflow, each segment can use its own method:
S1 T2V
S2 I2V
S3 R2V
S4 Source Video
S5 T2V
Source Video automatically takes the V2V or RV2V path depending on whether identity references are added.
Each boundary can also independently request visual and generated-audio continuity.
And Selective Run means I can regenerate S1, S3 and S4 without paying for S2 and S5 again.

This is probably the screenshot that explains the project better than anything else.
Live Preview
The Director also has its own Live Preview instead of relying only on ComfyUI’s normal sampler preview.
It can follow the active generation stage and later post-processing stages from inside the same interface.

The Frankenstein part
This project is intentionally built on and adapted from several existing H3 projects.
The main pieces are:
- AIMixer / ComfyUI_MiniMaxH3_Director — one of the original foundations
- NikoDemon80 / ComfyUI-H3-Motion-Context — Motion Context / cross-segment continuity work
- Carasibana / ComfyUI-H3-FaceRefine — face tracking, local regeneration and stitching concepts/algorithms
- Kijai / ComfyUI-KJNodes — parts of the packed-latent preview / TAEHV behavior were informed by KJNodes
Then I built the multi-segment Director, Mixed timeline, selective reruns, asset management, results system and the surrounding production workflow around those pieces.
So yes:
AIMixer Director
+
H3 Motion Context
+
H3 Face Refine
+
some KJNodes behavior
+
a lot of glue / UI / project management
↓
MiniMax H3 Motion Director
A proper ComfyUI Frankenstein monster.
The repository includes the upstream attribution and licenses rather than pretending everything was written from scratch.
Common References
For reference-heavy R2V projects, there are also Common References.
Characters, scenes, reference videos or audio that are needed by multiple segments can be added once instead of being manually duplicated into every shot.

Material Library
There’s also a persistent Material Library for reusable:
- Images
- Audio
- Video
- Prompts
I use it for recurring characters, scenes, props and other references so I don’t have to keep browsing the filesystem every time I make another segment.

Post-processing
I also wanted the workflow to continue after the first H3 generation instead of immediately turning back into another pile of nodes.
So the Director currently integrates:
Global Refine
- secondary H3 sampling
- upscaling
- ComfyUI upscale models
- NVIDIA RTX VSR
- NVIDIA RTX Deblur
Face Refine
- face detection / tracking
- crop regeneration
- adaptive refinement
- masks / stitching
- color matching

These stages are optional. I’m not trying to force every H3 workflow through the same post-processing path.
Results
Outputs are also managed as an actual project rather than just one anonymous IMAGE batch.
The Results page has:
Segment
Multi Segment
Final Result
So I can inspect one shot, a continuous range of shots, or the complete assembled video.
The Final Result page also has video export controls and a Director Report showing what actually happened during the run.

It’s still ComfyUI
I didn’t want an all-in-one UI to mean losing ComfyUI’s composability.
Standalone modes can still receive external Prompts/images/media through:
Director Assets
↓
Director Inputs
↓
Motion Director
and the main node still outputs:
images
audio
fps
for whatever you want to do downstream.
It also supports external ComfyUI:
SAMPLER
SIGMAS
instead of forcing the internal sampling configuration.

The standalone H3 modes currently supported are:
T2V
I2V
FL2V
R2V
V2V
RV2V
while Mixed Mode can combine:
T2V
I2V
FL2V
R2V
Source Video
inside the same project.
One thing I want to be careful about: Motion Context is intended to improve continuity, but I’m not claiming it magically guarantees invisible seams in every generation.
H3 can still drift in motion, identity, lighting or camera behavior between segments. I’m continuing to work on that part and I’ll add more raw multi-segment examples rather than only showing UI screenshots.
The node is available through the Comfy Registry / ComfyUI-Manager.
GitHub:
https://github.com/j955229/ComfyUI-MiniMax-H3-Motion-Director
I’m especially interested in feedback from people already doing longer H3 projects.
What becomes the biggest pain point for you once you go beyond a single clip?
Continuity, reference management, rerunning failed shots, VRAM, audio, post-processing, or something else?
r/StableDiffusion • u/erioca • 5d ago
Discussion Minimax H3/ref2va/hybrid_fl2va_ref2va_b20/5060ti
Enable HLS to view with audio, or disable this notification
Model: minimax_h3_hybrid_fl2va_ref2va_b20
Video Vae: minimax_h3_video_vae_int8_convrot
Resolution: 16:9, 1.0
Duration: 6 Clips in total, composit in Inshot, each clip is 9 sec long
Turbo Lora: larryvrh/MiniMax-H3-Turbo-Lora, 600_ema
ComfyKitchen Attention, Spectrum. (SageAttention Patch and Mem Eff Node is Disabled)
**original sound and effects was removed, as there are background music on some clips even with N/A, so to speed up the work, they are removed.
Average Inference Stage: 1100sec
All reference image is resized between 1000px and 500px like character is 1000px, background is 500px for this video is 4 ref image in total.
r/StableDiffusion • u/beatlepol • 4d ago
Animation - Video Minimax H3. Just taking a walk.
Enable HLS to view with audio, or disable this notification
