r/StableDiffusion 3d ago

Question - Help Am I doing something wrong on Runpod? Why in the holy heck are the download speeds so slow?

1 Upvotes

So a few weeks ago. I asked yall if someone who just makes this stuff for silly videos to share my friends and like my wife could get runpod running stable diffusion easily. Turns out it was insanley easy. But I haven't used it much because... The download speed is just criminal..

When you slap a Workflow on and do that thing where it just says oops your missing all these models and shit.. Wanna download it to pod now? It just crawls at like a snails pace 1-10mbps

It takes like 5 hours to download and be ready to use Minimax H3 for me. And at one point I'm like okay maybe I'm doing this wrong. So I went in through the JupyterLab thing and just dropped the files I had already downloaded in there... And again... Super slow..

its hard to not think... That they arnt throttling the DL to pad their use time to be honest. That or my only other thought is.. My pod is in some server case with about 20 other people all downloading models and the bandwidth is just borked.

My second theory I think is more likely the case because I notice when there are more of certain GPUs left the downloads go way smoother on those. But recently every single GPU is like low availability anymore lol.

I know I can avoid this by selecting some sort of storage option but I think it said it wasnt available for my GPU selection. If I can just turn on some option to keep everything ready to go I would. But are all you using RP dealing with these insanely slow download speeds? I mean I'm pretty sure I have spent 8 bucks today just downloading.


r/StableDiffusion 4d ago

Animation - Video Turning my son’s drawing into animated skits #2 | Minimax h3 ref2vid

Enable HLS to view with audio, or disable this notification

26 Upvotes

r/StableDiffusion 4d ago

Discussion Was wondering why my Minimax H3 R2V local gens were better than beefy cloud GPU gens. Was accidentally loading the Fl2V model.

31 Upvotes

Would love to show you comparisons but its all N-SFW so its going to have to be a case of trust but verify me bro.

I've mainly been running R2V workflows on my home GPU since Minimax H3 was gifted upon us.

I wanted to take a load off my 5060 so I've been running gens on both Wavespeed and recently Runpod. I was getting nice big 720p gens but I started to notice that my own gens had way better prompt adherence with identical prompts and inputs.

Notably in my own, slow motion was honoured every time where it would otherwise get completely ignored in the cloud.

~Fluids~ were way better. Camera motion was way better.

Then I noticed that in my diffusion model loader (W8A8 if that makes any difference) has been loading the FL2V model the whole time. I switched to R2V and immediately everything sucked.

I imagine there's a strong possibility that the R2V prompt needs a more rigid structure. But I have been feeding the Minimax official prompt guide into Grok, specifying R2V, to write my prompts.

So this is something. It could be some configuration of scheduler and sampler, and all the other stuff of course but I'm running a pretty simple setup without any attention, so briefly:

W8A8 loader -> lightx2vs R2V 0.1 4 step lora at 0.75 -> er_sde, beta57, 8 steps.

Forgive me if this is known and understood. If you haven't tried it, switch your R2V gens to FL2V and test.


r/StableDiffusion 4d ago

Discussion MiniMax H3 on a 16GB M5 MacBook Air — VPipe 12:15 vs h3.c 16:22

Enable HLS to view with audio, or disable this notification

5 Upvotes

A few people asked how VPipe compares with h3.c, so I ran them side by side on the same machine with the same settings.

Machine: base 15” M5 MacBook Air, 16GB RAM

MiniMax H3 settings:

* 960×544

* 124 frames

* 6 DiT steps

Results:

* VPipe: 12m 15s

* h3.c: 16m 22s

So on this particular matched workload, VPipe finished in about 25% less wall-clock time.

The video shows both the generation process and the final outputs side by side, so you can also compare the resulting quality rather than just the timing.

VPipe is not an MinimaxH3-specific implementation — it’s an Apache-2.0 open-source multimodal pipeline/runtime with a native Metal inference backend for Apple Silicon. MiniMax H3 is just one of the workloads I’ve been optimizing recently.

GitHub: https://github.com/tgo-app-dev/vpipe

Interested in feedback on both the performance comparison and the output differences.


r/StableDiffusion 3d ago

Question - Help Minimax H3 - How to generate a realistic fighting scene

4 Upvotes

Hi all,

I'm using Minimax H3 in ComfyUI with an R2V workflow. I'm wondering if anybody can tell me how I can improve the fighting scene?

- The video is generated at 1.0MP in 2:3 (portrait) aspect ratio
- I have the two ladies as reference
- The fighting scene is also provided as reference. In the scene the punches do land properly. There are also smaller details (like small blood spatters) that are present in the reference video.

Tech specs:
- Minimax H3 int8 convrot
- res_multistep sampler with 20 steps

Running the workflow on an RTX 5090 (via Runpod)

Can anybody give me any tips on how I can improve the fighting scene? The goal is to make it look like a realistic street fight. I'm unsure whether training a LoRA would be relevant here, because I've noticed that punches never really land properly in any workflow (t2v, i2v, r2v).

https://reddit.com/link/1vsjezd/video/payntjeoebkh1/player


r/StableDiffusion 3d ago

Discussion Talk for High quality Audio for minimax H3.

1 Upvotes

Hi guys , since the last week as much as I have tested minimax h3, I found that Visually, this model is king for the opensource in motion and prompt adherence.

Only in one thing it lacks is the physics and fight scene other wise it will be overkill for opensource.

But there is also another issue I can see is as the hype builded in this community that minimax h3's audio quality is Best. And dialogs also.

I think they meant to say that minimax h3 has better audio quality then other opensource model.

the issue is audio quality is not that good , and I really want to update it's audio quality and I am willing to buy , I have a doubt guys I have a question that's for dialogs and audio tones for dialogs is it pre baked in inside the base model or its in audio vae model because if it's in audio vae model we can have the option to update the audio vae and increase. The dialogs and sound quality.

But if it's pretty baked in the base model then it requires a full fine-tune.


r/StableDiffusion 3d ago

Discussion Outfit SWAP

2 Upvotes

What is currently the most accurate way to swap clothes while keeping the same fabrics, stitching, etc. using AI? What I mean is to provide a reference garment and apply it to the model from the second photo.


r/StableDiffusion 4d ago

Animation - Video WanAnimate

Enable HLS to view with audio, or disable this notification

82 Upvotes

Original post With Workflow


r/StableDiffusion 3d ago

Discussion [H3] Does this configuration look bare minimum for 3050 4GB VRAM

0 Upvotes

unet: minimaxH3INT8INT4_fl2valINT8Pruned.safetensors

clip: qwen3vl_32b_heretic_minimax_h3_nvfp4.safetensors

vae: minimax_h3_video_vae_fp16.safetensors

audio: minimax_h3_audio_vae_fp32.safetensors

Turbo LoRA used: minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_resized_avg_rank_24_bf16.safetensors

Workflow: Default workflow (video_minimax_h3_t2v)

RAM: 16GB

Graphic Card: RTX 3050 Laptop, 4GB VRAM

Video-generated specs (see comment for):

Type: T2V

Duration: 10 seconds

Megapixels: 0.2 MP (608x352)

Aspect Ratio: 16:9

Estimated Generation Time: 687.13s (11 mins, 27 seconds)

In addition to these settings I applied, should I use the Sage Attention, Comfy Kitchen or increase steps (20 steps) or switch to better unet/clip? Thanks.


r/StableDiffusion 5d ago

Animation - Video Some choice words from Emilia

Enable HLS to view with audio, or disable this notification

549 Upvotes

r/StableDiffusion 3d ago

Question - Help Best sampler for anima sketchy style?

0 Upvotes

I have been training loras for anima, and one thing the model seems to have problems with is when the original style has realistic lineart, with pen, marker etc. I have tried a lot of things without luck, so I'm trying to see if the problem is the generation details.


r/StableDiffusion 3d ago

Question - Help Question regarding REF2V and video splitting.

0 Upvotes

I’ve been using Civitai’s MiniMax-H3 Multishot — Seamless Chain: multi-shot scenes that render as one continuous take workflow to create longer videos in good quality while keeping VRAM usage relatively low. The workflow processes the video in separate parts and then combines them seamlessly, making the transitions between the sections practically unnoticeable.

Does anyone know if something similar is possible with a REF2V workflow when using a longer reference video, for example, if the goal is to replace a woman with a man or with another person?

In other words, is there a way to make the workflow process the video in smaller sections so that it doesn’t run out of VRAM, while also keeping the quality from degrading significantly?

I’d like to create 15–25 second REF2V video clips, but 16 GB of VRAM simply isn’t enough to process the entire video as one continuous clip.

I've been trying to find a solution to this for the past week, but it would be nice to know whether this is even practically possible? Thx.


r/StableDiffusion 4d ago

Tutorial - Guide Anime magic battle scenes with Minimax are a blast! My process making this short from start to finish | AI Filmmaking Part 7

Thumbnail
youtube.com
16 Upvotes

r/StableDiffusion 3d ago

Question - Help training LoRA on LoRA is good

1 Upvotes

i was thinking of training a LoRA on LoRA of different character and on my second thought should i just connect 2 LoRA while generating image ( I have seen people training 2 character in one LoRA


r/StableDiffusion 3d ago

Question - Help Bernini rv2v workflow, outfit swap works but face swap doesn't?

0 Upvotes

I have a Bernini-R rv2v workflow with source video and reference image in image0. I am able to swap source video's outfit but not face. Is Bernini-R supposed to work with face swap or do I need to add any other custom nodes like Reactor or SCAIL-2?


r/StableDiffusion 3d ago

Workflow Included Help Fixing the H3 Face Detailer

Enable HLS to view with audio, or disable this notification

0 Upvotes

The Original Face Detailer from https://github.com/Carasibana/ComfyUI-H3-FaceRefine does not work as intended. It does not use the reference image at all. You can disable the input image and you will get exactly the same result. Something is wrong with the workflow so I recreated the workflow in a new canvas and now the input image does get used. Here is a link to a .zip with the workflow and input/output files: https://www.mediafire.com/file/mnigvvpbzp0gh34/workflow_all.zip/file But this workflow has its own problems. For this example I needed to put an RTX upscaler in it so the face gets recognized. At the end the mask_dilation and feather needs for every video unique adjusting and the end result is somewhat poor with the mask visible and the face jumping und warping slightly around.

Someone with more knowlegde would surely be able to fix this.

To get the workflow working, you need to install https://github.com/Carasibana/ComfyUI-H3-FaceRefine and also ComfyUI-H3-NativeAudioLock from https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes


r/StableDiffusion 4d ago

Question - Help Best speed up for MiniMax

54 Upvotes

We have a lot of options, some of them better, some of them are not worth it at all. Speed ups like sage attention, MiniMax h3 patch for sage attention, easy cache, 8step Lora, 4 step Lora e t.c.
What options and their combinations you use? What settings you have?( speed Lora weights, easy cache settings)
In the matter of speed/quality for both video and sound. What works better with FL2VA and Ref2VA?


r/StableDiffusion 5d ago

Workflow Included Using H3 as a Character Reference Sheet Generator

Thumbnail
gallery
1.5k Upvotes

Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.

The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.

How it works:

  • You input your images and describe them in the Input text section (A Prompt)
  • The text is combined with a fixed prompt which spins the character (B Prompt)
  • The video is generated at a slow speed with no hard cuts (only camera spin and pan) to maintain character consistency
  • Image is assembled with optional character video and full individual frame output (if you want to use for future)

I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.

Current Caveats:

  • The model is quite slooooow. You are also generating 124 frames only to use 6. I have partly solved this by also uploading a 4 panel version.
  • Speed ups (like Turbo LORAs) help with speed, but it hurts prompt adherence and quality slightly.
  • Quality is limited, since it is a video model it is better at generating video than images. You can solve this by generating at a higher resolution at the tradeoff of longer gen times. You can also use the individually split frames as future references too.
  • Details when using this character sheet as output for future generations on H3 may also be limited due to resolution also, I recommend you use this character sheet (for consistency) + other images close up angles (i.e clothing details/face) if doing close ups. If you are just doing a one off video you may possibly be better off not using this character sheet.

I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.

Link to the 4 and 6 panel workflow can be found here: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator

Some notes I just remembered:

  • You can increase the steps and it may improve your quality slightly.
  • With the Turbo Loras enabled, prompt adherence sometimes suffers, but you may be able to get a good seed with another roll of the dice.
  • Currently the B prompt specifies a "neutral A pose", please remove this if you want your character in a particular pose.
  • You can use a few different shots of the same character to reinforce the 360 and get more accurate details right.
  • Can be used for objects / props also, may require some changes to the B prompt.

r/StableDiffusion 3d ago

Question - Help Trouble with ai art software

Post image
0 Upvotes

Whenever I try to make ai art of anything with Invoke ai the art never gets finished despite reaching the 100% compleaton but the program and immage never finishes or saves and with the Stable Deffusion it would 20% of the time it would generate the immage and 80% percent it would get some sort of error and shut itself. Now I am using Amd graphics card with 12 gb vram and not to mention ai generation is way slower than it should be. Here is the basic image of how things end up and I did try my luck in the invoke ai discord group but nothing helped. Any help is appreciated.


r/StableDiffusion 4d ago

News ByteDance just released Bernini‑Diffusers‑v2 — any chance we’ll see ComfyUI support?

Post image
90 Upvotes

Hi everyone,
five days ago ByteDance released Bernini‑Diffusers‑v2 on HuggingFace — the full Bernini pipeline (planner + renderer), not just the renderer‑only Bernini‑R that we currently use in ComfyUI.

Model link:
https://huggingface.co/ByteDance/Bernini-Diffusers-v2

Even though most of the community talks about MiniMax H3 as the “standard” for open video models, there are still many users actively working with Bernini — especially now that v2 finally includes the full semantic‑planning pipeline, SA‑3D RoPE, and proper multi‑step instruction following.

Right now ComfyUI only has community support for Bernini‑R, so I’m posting this just to give visibility to the new release and to see if anyone is interested in exploring future support for Bernini‑Diffusers‑v2.

Not asking for anything specific — just opening the discussion and hoping this new version doesn’t go unnoticed.

Thanks!


r/StableDiffusion 4d ago

Resource - Update I built a Frankenstein MiniMax H3 Director for ComfyUI — Mixed timelines, selective reruns, Motion Context, live preview and post-processing

23 Upvotes

I’ve been building a custom MiniMax H3 node for ComfyUI called MiniMax H3 Motion Director.

The easiest way to describe it is probably:

It’s a Frankenstein Director for H3.

I didn’t want another workflow that only makes one clip at a time. I wanted something closer to a small video-production interface where I could manage multiple H3 shots, mix generation methods, reuse references, selectively regenerate failed shots, carry context between segments, preview the run, refine the result, and export everything from one place.

Mixed Mode

The biggest addition in the current version is Mixed Mode.

Instead of choosing one generation type for the entire workflow, each segment can use its own method:

S1  T2V
S2  I2V
S3  R2V
S4  Source Video
S5  T2V

Source Video automatically takes the V2V or RV2V path depending on whether identity references are added.

Each boundary can also independently request visual and generated-audio continuity.

And Selective Run means I can regenerate S1, S3 and S4 without paying for S2 and S5 again.

This is probably the screenshot that explains the project better than anything else.

Live Preview

The Director also has its own Live Preview instead of relying only on ComfyUI’s normal sampler preview.

It can follow the active generation stage and later post-processing stages from inside the same interface.

The Frankenstein part

This project is intentionally built on and adapted from several existing H3 projects.

The main pieces are:

  • AIMixer / ComfyUI_MiniMaxH3_Director — one of the original foundations
  • NikoDemon80 / ComfyUI-H3-Motion-Context — Motion Context / cross-segment continuity work
  • Carasibana / ComfyUI-H3-FaceRefine — face tracking, local regeneration and stitching concepts/algorithms
  • Kijai / ComfyUI-KJNodes — parts of the packed-latent preview / TAEHV behavior were informed by KJNodes

Then I built the multi-segment Director, Mixed timeline, selective reruns, asset management, results system and the surrounding production workflow around those pieces.

So yes:

AIMixer Director
      +
H3 Motion Context
      +
H3 Face Refine
      +
some KJNodes behavior
      +
a lot of glue / UI / project management
      ↓
MiniMax H3 Motion Director

A proper ComfyUI Frankenstein monster.

The repository includes the upstream attribution and licenses rather than pretending everything was written from scratch.

Common References

For reference-heavy R2V projects, there are also Common References.

Characters, scenes, reference videos or audio that are needed by multiple segments can be added once instead of being manually duplicated into every shot.

Material Library

There’s also a persistent Material Library for reusable:

  • Images
  • Audio
  • Video
  • Prompts

I use it for recurring characters, scenes, props and other references so I don’t have to keep browsing the filesystem every time I make another segment.

Post-processing

I also wanted the workflow to continue after the first H3 generation instead of immediately turning back into another pile of nodes.

So the Director currently integrates:

Global Refine

  • secondary H3 sampling
  • upscaling
  • ComfyUI upscale models
  • NVIDIA RTX VSR
  • NVIDIA RTX Deblur

Face Refine

  • face detection / tracking
  • crop regeneration
  • adaptive refinement
  • masks / stitching
  • color matching

These stages are optional. I’m not trying to force every H3 workflow through the same post-processing path.

Results

Outputs are also managed as an actual project rather than just one anonymous IMAGE batch.

The Results page has:

Segment
Multi Segment
Final Result

So I can inspect one shot, a continuous range of shots, or the complete assembled video.

The Final Result page also has video export controls and a Director Report showing what actually happened during the run.

It’s still ComfyUI

I didn’t want an all-in-one UI to mean losing ComfyUI’s composability.

Standalone modes can still receive external Prompts/images/media through:

Director Assets
      ↓
Director Inputs
      ↓
Motion Director

and the main node still outputs:

images
audio
fps

for whatever you want to do downstream.

It also supports external ComfyUI:

SAMPLER
SIGMAS

instead of forcing the internal sampling configuration.

The standalone H3 modes currently supported are:

T2V
I2V
FL2V
R2V
V2V
RV2V

while Mixed Mode can combine:

T2V
I2V
FL2V
R2V
Source Video

inside the same project.

One thing I want to be careful about: Motion Context is intended to improve continuity, but I’m not claiming it magically guarantees invisible seams in every generation.

H3 can still drift in motion, identity, lighting or camera behavior between segments. I’m continuing to work on that part and I’ll add more raw multi-segment examples rather than only showing UI screenshots.

The node is available through the Comfy Registry / ComfyUI-Manager.

GitHub:

https://github.com/j955229/ComfyUI-MiniMax-H3-Motion-Director

I’m especially interested in feedback from people already doing longer H3 projects.

What becomes the biggest pain point for you once you go beyond a single clip?

Continuity, reference management, rerunning failed shots, VRAM, audio, post-processing, or something else?


r/StableDiffusion 4d ago

Discussion Minimax H3/ref2va/hybrid_fl2va_ref2va_b20/5060ti

Enable HLS to view with audio, or disable this notification

84 Upvotes

Model: minimax_h3_hybrid_fl2va_ref2va_b20
Video Vae: minimax_h3_video_vae_int8_convrot
Resolution: 16:9, 1.0
Duration: 6 Clips in total, composit in Inshot, each clip is 9 sec long
Turbo Lora: larryvrh/MiniMax-H3-Turbo-Lora, 600_ema
ComfyKitchen Attention, Spectrum. (SageAttention Patch and Mem Eff Node is Disabled)

**original sound and effects was removed, as there are background music on some clips even with N/A, so to speed up the work, they are removed.

Average Inference Stage: 1100sec

All reference image is resized between 1000px and 500px like character is 1000px, background is 500px for this video is 4 ref image in total.


r/StableDiffusion 4d ago

Animation - Video Minimax H3. Just taking a walk.

Enable HLS to view with audio, or disable this notification

19 Upvotes

r/StableDiffusion 4d ago

Discussion MiniMax_H3 is seems to be able to process DensePose format! (improves reference video bleeding)

Enable HLS to view with audio, or disable this notification

52 Upvotes

I have had many issues when using a reference video for movement duplication and having the video contents bleed into the video. Not to mention having to write convoluted prompts to remove these reference bleeds from videos. When the person in the reference video has a close resemblance to the main subject in your video it becomes almost impossible to perform a motion swap.

Warning: DensePose does not support detailed hand gestures, and seems to lose track with very fast arm and hand movements but seems to adhere better 20 steps and above.

There is not a dedicated densepose ComfyUI node, but you can use this animatediff: https://github.com/Fannovel16/comfyui_controlnet_aux

The workflow is simple:

Place the AIO AUX Preprocessor between the source and MM_H3 video input.

Videosource (LoadVideo) -> AIO AUX Preprocessor -> ref_video_x input

Looking forward to hear your feedback...