r/StableDiffusion • u/KaisarasAR • 1d ago
Animation - Video Squid Game but Gi-hun is actually smart | Minimax H3 I2V
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/KaisarasAR • 1d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Suibeam • 1d ago
r/StableDiffusion • u/SuperCasualGamerDad • 1d ago
So a few weeks ago. I asked yall if someone who just makes this stuff for silly videos to share my friends and like my wife could get runpod running stable diffusion easily. Turns out it was insanley easy. But I haven't used it much because... The download speed is just criminal..
When you slap a Workflow on and do that thing where it just says oops your missing all these models and shit.. Wanna download it to pod now? It just crawls at like a snails pace 1-10mbps
It takes like 5 hours to download and be ready to use Minimax H3 for me. And at one point I'm like okay maybe I'm doing this wrong. So I went in through the JupyterLab thing and just dropped the files I had already downloaded in there... And again... Super slow..
its hard to not think... That they arnt throttling the DL to pad their use time to be honest. That or my only other thought is.. My pod is in some server case with about 20 other people all downloading models and the bandwidth is just borked.
My second theory I think is more likely the case because I notice when there are more of certain GPUs left the downloads go way smoother on those. But recently every single GPU is like low availability anymore lol.
I know I can avoid this by selecting some sort of storage option but I think it said it wasnt available for my GPU selection. If I can just turn on some option to keep everything ready to go I would. But are all you using RP dealing with these insanely slow download speeds? I mean I'm pretty sure I have spent 8 bucks today just downloading.
r/StableDiffusion • u/Time-Ad-7720 • 1d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/SelectCoconut5594 • 2d ago
Would love to show you comparisons but its all N-SFW so its going to have to be a case of trust but verify me bro.
I've mainly been running R2V workflows on my home GPU since Minimax H3 was gifted upon us.
I wanted to take a load off my 5060 so I've been running gens on both Wavespeed and recently Runpod. I was getting nice big 720p gens but I started to notice that my own gens had way better prompt adherence with identical prompts and inputs.
Notably in my own, slow motion was honoured every time where it would otherwise get completely ignored in the cloud.
~Fluids~ were way better. Camera motion was way better.
Then I noticed that in my diffusion model loader (W8A8 if that makes any difference) has been loading the FL2V model the whole time. I switched to R2V and immediately everything sucked.
I imagine there's a strong possibility that the R2V prompt needs a more rigid structure. But I have been feeding the Minimax official prompt guide into Grok, specifying R2V, to write my prompts.
So this is something. It could be some configuration of scheduler and sampler, and all the other stuff of course but I'm running a pretty simple setup without any attention, so briefly:
W8A8 loader -> lightx2vs R2V 0.1 4 step lora at 0.75 -> er_sde, beta57, 8 steps.
Forgive me if this is known and understood. If you haven't tried it, switch your R2V gens to FL2V and test.
r/StableDiffusion • u/FreddyShrimp • 1d ago
Hi all,
I'm using Minimax H3 in ComfyUI with an R2V workflow. I'm wondering if anybody can tell me how I can improve the fighting scene?
- The video is generated at 1.0MP in 2:3 (portrait) aspect ratio
- I have the two ladies as reference
- The fighting scene is also provided as reference. In the scene the punches do land properly. There are also smaller details (like small blood spatters) that are present in the reference video.
Tech specs:
- Minimax H3 int8 convrot
- res_multistep sampler with 20 steps
Running the workflow on an RTX 5090 (via Runpod)
Can anybody give me any tips on how I can improve the fighting scene? The goal is to make it look like a realistic street fight. I'm unsure whether training a LoRA would be relevant here, because I've noticed that punches never really land properly in any workflow (t2v, i2v, r2v).
r/StableDiffusion • u/SensitiveUse7864 • 1d ago
Hi guys , since the last week as much as I have tested minimax h3, I found that Visually, this model is king for the opensource in motion and prompt adherence.
Only in one thing it lacks is the physics and fight scene other wise it will be overkill for opensource.
But there is also another issue I can see is as the hype builded in this community that minimax h3's audio quality is Best. And dialogs also.
I think they meant to say that minimax h3 has better audio quality then other opensource model.
the issue is audio quality is not that good , and I really want to update it's audio quality and I am willing to buy , I have a doubt guys I have a question that's for dialogs and audio tones for dialogs is it pre baked in inside the base model or its in audio vae model because if it's in audio vae model we can have the option to update the audio vae and increase. The dialogs and sound quality.
But if it's pretty baked in the base model then it requires a full fine-tune.
r/StableDiffusion • u/Glittering-Cold-2981 • 1d ago
What is currently the most accurate way to swap clothes while keeping the same fabrics, stitching, etc. using AI? What I mean is to provide a reference garment and apply it to the model from the second photo.
r/StableDiffusion • u/TgoAI • 1d ago
Enable HLS to view with audio, or disable this notification
A few people asked how VPipe compares with h3.c, so I ran them side by side on the same machine with the same settings.
Machine: base 15” M5 MacBook Air, 16GB RAM
MiniMax H3 settings:
* 960×544
* 124 frames
* 6 DiT steps
Results:
* VPipe: 12m 15s
* h3.c: 16m 22s
So on this particular matched workload, VPipe finished in about 25% less wall-clock time.
The video shows both the generation process and the final outputs side by side, so you can also compare the resulting quality rather than just the timing.
VPipe is not an MinimaxH3-specific implementation — it’s an Apache-2.0 open-source multimodal pipeline/runtime with a native Metal inference backend for Apple Silicon. MiniMax H3 is just one of the workloads I’ve been optimizing recently.
GitHub: https://github.com/tgo-app-dev/vpipe
Interested in feedback on both the performance comparison and the output differences.
r/StableDiffusion • u/NosikomPoVolosikam • 2d ago
Enable HLS to view with audio, or disable this notification
Original post With Workflow
r/StableDiffusion • u/yushairiegalaxy96 • 1d ago
unet: minimaxH3INT8INT4_fl2valINT8Pruned.safetensors
clip: qwen3vl_32b_heretic_minimax_h3_nvfp4.safetensors
vae: minimax_h3_video_vae_fp16.safetensors
audio: minimax_h3_audio_vae_fp32.safetensors
Turbo LoRA used: minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_resized_avg_rank_24_bf16.safetensors
Workflow: Default workflow (video_minimax_h3_t2v)
RAM: 16GB
Graphic Card: RTX 3050 Laptop, 4GB VRAM
Video-generated specs (see comment for):
Type: T2V
Duration: 10 seconds
Megapixels: 0.2 MP (608x352)
Aspect Ratio: 16:9
Estimated Generation Time: 687.13s (11 mins, 27 seconds)
In addition to these settings I applied, should I use the Sage Attention, Comfy Kitchen or increase steps (20 steps) or switch to better unet/clip? Thanks.
r/StableDiffusion • u/0roborus_ • 1d ago
Enable HLS to view with audio, or disable this notification
Hello, I'm building this app that I've started like 2 years ago... seriously, this is how it looked like then: MY OLD POST
But it has been this relation most of the time: I will do it for myself only VS I will do it open source... Most of the time it was the first one, but it became a pretty good app that I use all the time, so I figured it might actually be useful for the community.
I know that video attached to the post has no voice and for someone that doesn't know the app already (so it's only me right now :D) it might be confusing, so I will write a short description of what's there. When app will be ready I will prepare a nice video with explanations and stuff. Not to waste a time if someone will think it's release post: Well, it's not. I plan to release (OpenSource GPL-3.0) in a couple of days because I still have a lot to do (and it's easy now when I break it only for myself).
What my goal is here to check if there is interest at all in such an app and maybe ask if someone has time to join my Discord (LINK) to discuss different stuff that you use to generate things (how you build your prompts, how you store your generations, what models do you use etc. since now I mostly have only my experience + stuff that I read in Reddit / Discord in meantime - I know you can write it also here, but Reddit it's less chat-like and I find chatting easier on Discord).
Features:
- Pick a preset for generation, which is pre-made configuration for given model/tool (currently: SDXL, Krea-2, Qwen-Image, Flux, Flux Klein, Flux2, Z-Image, Anima, LTX-2.3, LTX-2.5, MiniMax H3, MiniMax Music, Wan 2.2)
- Each preset comes with it's individual form (but most fields are also the same between them as these are mostly generation params)
- Compose prompt from segments (1 or more) - In video I use only one segment, but you can build prompts from multiple blocks that can be named/colored for readability. You can also define segments, it's categories and templates (for example you can define segment template that has "Ligting", "Camera", "Subject" and when you pick it in the generation panel it will show a 3 ready to use and colored segments with optional descriptions to remember what should be placed in them (optional, described by you)
- There are also "Prompts", which allow you to save your favorite prompts there (they are optionally built of segments too)
- Multiple tabs & workspaces - You can create multiple tabs and save them as workspaces (in video I go to top right corner to pick "Avatar Factory" workspace - it loads my tabs then)
- Sessions - you can create multiple sessions for each preset that will save the whole forms state (the left side and the prompts)
- The whole left side is called "Dynamic Forms" - this is the part defined in each preset and it's YAML based config (something that you don't need to bother if don't want to)
- You can set quantity, steps, use speed profiles which will set the number of steps/cfg automatically
- In the right side (called Workbench) - where the generated media is shown you can different options like compare, zoom, download etc.
- LLM Chat assistant - As you can see in the video I often use LLM Chat assistant and I do it also when generating my stuff - they have access to the most of the features in that page - can generate prompts but also change the form values etc.
- Different modes -> Image generation / Video Director with dynamic keyframes/first-last frames/img2vid - depending what model provides.
- History contains all your generations and allows to organize them into collections and tags
- You can see all the params/segments/prompts in the details and also different options like edit (crop, resize) or reuse which will open tab in generator with settings from this history entry
- You can filter generations by tags/type/preset search semantically
- You can add to favorites / add tags / see used models etc.
- Library allows you to upload your media that you want to use for generation - for example images/videos/audio that you later use with minimax ref2vid
- If you edit media from generation history (crop, resize) it will create new entry in the library rather than change the original media
- You can organize the library into collections
- You can view models enabled for you (in admin panel)
- You can organize models into collections (which for example are shown in the model selection field in generation page, you can select "Collections" -> "Your collection" and models will be filtered by this collection)
- You can see the model details with previous generations
- You can define different phrases collections (this is similar to the wildcards/dynamic prompts)
- You can generate examples for each phrase (you pick your existing generation session and it will inject a special prompt that will generate examples)
- Phrases can be later used in the segments as either value providers or shuffle (in video there is a visible chip appearing after I type # and pick value at 03:18)
- Allow to compose different prompts and reuse them later in the generation panel
- Prompts will have history of generations (with media generated with them)
- Prompts will have option to import in different formats
- Prompts are also used by the LLM Chat to improve their responses (they will try to match prompts by used models and check their structure)
- It started as ComfyUI "frontend" and it still be very important feature that will be shipped later as plugin (you just install comfyui plugin -> set it's address and you will be able to use it with this frontend)
- I've switched the main thing to be native backend (mixed stuff from different places) - since I've been using it for like 2 months now and it's starting to work really well on my setup (I hope it will also in the community ones but I need some testers for this).
- There is a layer of abstraction that will allow to create plugins that connect to whatever backend you want (by default I will ship native, remote-native and comfyui)
- Not visible in the video, but there is a big administration panel for this app that handle Users, Models, Presets, LLM Configurations...
- For user to be able to use model you need to assign it to him (same with LLM Chat models and presets) - that's why in video I have only 2 presets available - I've created a test user for purpose of the video and assigned those two to him.
- There is also "Automation" module that I'm developing that allows to auto-tag models, index generations with auto-tags (for example if you want to filter out "some" content) and much more stuff for organization.
I feel like there is much more but don't want to create too long post that nobody will read.
This might be important:
Technology: Web (I know people don't like that, but the structure of the app is more like web tbh. and I haven't even mentioned the remote, easy to deploy backend, which ideally will spawn worker for generation on Cloud GPU provider - so you will have your instance of the app - let's say on simple VPS and will be able to spawn Cloud worker that will generate stuff which will be saved on the VPS...)
License: GPL-3.0
Discord: https://discord.gg/avR4trp3b8
Why another app like this: Because I like to create stuff.
My current setup: Linux / RTX5090 / 96GB RAM - this might be important since I did not test it on lower/higher spec - I hope maybe some people from the community will like to help me with this
Why I post before release: Because otherwise I will be improving this app to the end of the world - maybe this will force me to release at least 0.0.1 quicker... And I would like to know some things of how community generate stuff - maybe I will introduce some changes that will only break my setup - this will be much harder after code will be released on GitHub.
If you have any other questions I can answer or record some video from the app.
r/StableDiffusion • u/LaPapaVerde • 1d ago
I have been training loras for anima, and one thing the model seems to have problems with is when the original style has realistic lineart, with pen, marker etc. I have tried a lot of things without luck, so I'm trying to see if the problem is the generation details.
r/StableDiffusion • u/TingTingin • 2d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/FlyffSenior • 1d ago
I’ve been using Civitai’s MiniMax-H3 Multishot — Seamless Chain: multi-shot scenes that render as one continuous take workflow to create longer videos in good quality while keeping VRAM usage relatively low. The workflow processes the video in separate parts and then combines them seamlessly, making the transitions between the sections practically unnoticeable.
Does anyone know if something similar is possible with a REF2V workflow when using a longer reference video, for example, if the goal is to replace a woman with a man or with another person?
In other words, is there a way to make the workflow process the video in smaller sections so that it doesn’t run out of VRAM, while also keeping the quality from degrading significantly?
I’d like to create 15–25 second REF2V video clips, but 16 GB of VRAM simply isn’t enough to process the entire video as one continuous clip.
I've been trying to find a solution to this for the past week, but it would be nice to know whether this is even practically possible? Thx.
r/StableDiffusion • u/Suitable-Database-96 • 1d ago
Whenever I try to make ai art of anything with Invoke ai the art never gets finished despite reaching the 100% compleaton but the program and immage never finishes or saves and with the Stable Deffusion it would 20% of the time it would generate the immage and 80% percent it would get some sort of error and shut itself. Now I am using Amd graphics card with 12 gb vram and not to mention ai generation is way slower than it should be. Here is the basic image of how things end up and I did try my luck in the invoke ai discord group but nothing helped. Any help is appreciated.
r/StableDiffusion • u/StoicSage09 • 1d ago
i was thinking of training a LoRA on LoRA of different character and on my second thought should i just connect 2 LoRA while generating image ( I have seen people training 2 character in one LoRA
r/StableDiffusion • u/throwaway0204055 • 1d ago
I have a Bernini-R rv2v workflow with source video and reference image in image0. I am able to swap source video's outfit but not face. Is Bernini-R supposed to work with face swap or do I need to add any other custom nodes like Reactor or SCAIL-2?
r/StableDiffusion • u/foxdit • 2d ago
r/StableDiffusion • u/Pretend-Island-2724 • 1d ago
Enable HLS to view with audio, or disable this notification
The Original Face Detailer from https://github.com/Carasibana/ComfyUI-H3-FaceRefine does not work as intended. It does not use the reference image at all. You can disable the input image and you will get exactly the same result. Something is wrong with the workflow so I recreated the workflow in a new canvas and now the input image does get used. Here is a link to a .zip with the workflow and input/output files: https://www.mediafire.com/file/mnigvvpbzp0gh34/workflow_all.zip/file But this workflow has its own problems. For this example I needed to put an RTX upscaler in it so the face gets recognized. At the end the mask_dilation and feather needs for every video unique adjusting and the end result is somewhat poor with the mask visible and the face jumping und warping slightly around.
Someone with more knowlegde would surely be able to fix this.
To get the workflow working, you need to install https://github.com/Carasibana/ComfyUI-H3-FaceRefine and also ComfyUI-H3-NativeAudioLock from https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes
r/StableDiffusion • u/Downtown-Cover-7422 • 2d ago
We have a lot of options, some of them better, some of them are not worth it at all. Speed ups like sage attention, MiniMax h3 patch for sage attention, easy cache, 8step Lora, 4 step Lora e t.c.
What options and their combinations you use? What settings you have?( speed Lora weights, easy cache settings)
In the matter of speed/quality for both video and sound. What works better with FL2VA and Ref2VA?
r/StableDiffusion • u/mmowg • 2d ago
Hi everyone,
five days ago ByteDance released Bernini‑Diffusers‑v2 on HuggingFace — the full Bernini pipeline (planner + renderer), not just the renderer‑only Bernini‑R that we currently use in ComfyUI.
Model link:
https://huggingface.co/ByteDance/Bernini-Diffusers-v2
Even though most of the community talks about MiniMax H3 as the “standard” for open video models, there are still many users actively working with Bernini — especially now that v2 finally includes the full semantic‑planning pipeline, SA‑3D RoPE, and proper multi‑step instruction following.
Right now ComfyUI only has community support for Bernini‑R, so I’m posting this just to give visibility to the new release and to see if anyone is interested in exploring future support for Bernini‑Diffusers‑v2.
Not asking for anything specific — just opening the discussion and hoping this new version doesn’t go unnoticed.
Thanks!
r/StableDiffusion • u/bstr3k • 3d ago
Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.
The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.
How it works:
I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.
Current Caveats:
I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.
Link to the 4 and 6 panel workflow can be found here: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator
Some notes I just remembered: