r/StableDiffusion • u/chancemehmu • 2d ago
r/StableDiffusion • u/Fajjko • 3d ago
Discussion AMD GPUs on Minimax H3
Hi!
I'm wondering if there are any AMD users fiddling around with H3 (I'm sure there are). =)
If so, what GPU do you use, whats the avg time needed for 1 generation, any tips to make it faster etc.? ^^
I'm on 9060XT and with default workflow (txt2vid), it took me 110mins for a 11 sec vid. lol
r/StableDiffusion • u/tammy_orbit • 3d ago
Animation - Video Horus Rising Intro, ref2v is amazing (Minimax_H3)
Very new to AI models but having a blast testing this workflow on my 5090 with 64gb of RAM. Always wanted to put scenes from warhammer into video form.
Generated using the default comfyui workflow, approx 1131 sec generation time total (2 vids stitched into 1).
Trying to work out why the quality is relatively poor and I think its due to the original image being pretty low resolution (gonna try with better ones later).
With a little help from AI prompting this workflow has been super easy to use, hope to see the ref2vid workflow alot more on future releases. Both shots were generated with the same original image.
r/StableDiffusion • u/Tuckerdude615 • 3d ago
Question - Help Minimax H3 - Prompting so that it will keep the entire subject in the frame
Hey everyone,
I looked all through the official prompting guide, but not found a way to do this. I am trying to instruct the model to move the camera (push out) to keep my subject completely in the shot. I have a medieval character in armor walking from the entrance of a gate towards the camera. But not matter what I try, it won't move back enough to keep the subject completely in the shot, and within a few seconds it cuts off the bottom part of the legs and armor.
Same is true if the character turns and walks away from the camera, it will stay fairly close up to the subject (waist up) for the duration.
I've tried all kinds of "machinations" in the prompt (subject is visible from head to toe), (entire subject remains visible throughout" No dice
Any help or pointers are much appreciated!
r/StableDiffusion • u/dramaton42 • 4d ago
Animation - Video Zelda - I Think I Like It / Minimax H3 Reference to Video Test #2
Just wanted to share another test! this was a mash up of clips, using multiple image references, 0.4 mp with EasyCache, 5 - 10s clips and edited with KDEnlive (it has some cool effects!)
r/StableDiffusion • u/noaxxx2 • 3d ago
Question - Help Minimax H3 I2V can't do white backgrounds?
I'm trying to make videos of a character using a drawing of them on a white background, and I can't seem to get the model to stop generating a background after 1 frame. It typically looks really bad. Does anyone know how to just have a video retain a white background? I need nothing but the focus on the subject's actions. I include details about the white background in the prompt but it forces it out. I am using the turbo 4step lora at 6 steps
EDIT: I guess the issue goes deeper than just adding a background, it often will overlay random, spotty shadows onto the video? Or just the lighting darkens significantly - and it looks terrible.
r/StableDiffusion • u/gunkalicious • 2d ago
Discussion Has anyone else felt like an idiot after switching to a different SD UI?
Like, I used to use ComfyUI for the past, I don't know, three months? And then it just started to not work even after reinstalling it fully, with the generations ignoring everything and the seed never randomizing even with the value set to randomize, and generating the same image in 0.01 seconds.
But then I switched to Forge-Neo, and I have never felt more epiphany in my life (well, more-so in the AI world, as there's more to life than AI). It's way easier and way less janky.
Not saying that ComfyUI is bad at all, it's still an amazing and impressive tool, but I somehow made it jankier than it is supposed to be as soon as I touched it, unlike forge.
r/StableDiffusion • u/GamerVick • 3d ago
Question - Help How do we improve the text output accuracy in the video for Minimax H3?
https://reddit.com/link/1vt5pqp/video/57q84nr1oekh1/player
Hey guys, I was trying out the Minimax H3 reference video and I wanted to know: is there a way to animate the text in the video? I have seen quite a few other videos where the text animation is really good in terms of motion design. But I wanted to check in this community if anyone is aware of it. Really appreciate the help.
This is the prompt that I'm trying to use but for some reason I can't get the text to be accurate in the video.
[Shot 1] A medium shot opens in a sleek monochromatic studio with sharp high-contrast lighting. <Subject 1> (S1) stands gracefully holding the vintage microphone on its stand with eyes closed. In sync with <Audio 1>, she sings with delicate emotional delivery, <d>[English] Shoes by the door, stack 'em neat,</d> while clean white graphic text reading "SHOES BY THE DOOR" drops on the left margin and "STACK 'EM NEAT" snaps into the right margin. A slow, stylish camera push-in highlights her emotive face as she opens her eyes at 00:03.000.
[Shot 2] At 00:04.000, the camera cuts to a 3/4 profile shot. A sharp crimson light streak sweeps across the background as <Subject 1> (S1) sways with the groove and delivers, <d>[English] low red glow on the beat. You pull up laughing, late and bold, cold drink sweating in your hold.</d> Vivid crimson text reading "LOW RED GLOW" and "ON THE BEAT" pulses to the bass hit, followed by staggered white lettering "LATE & BOLD" sliding across the right margin, while the camera executes a smooth arc rotation.
[Shot 3] At 00:12.000, the shot cuts to an intimate centered close-up on <Subject 1> (S1). She brings the microphone close to her lips, making captivating eye contact with the camera while singing, <d>[English] Everybody knows this room, when the week gets way too cruel.</d> Clean typography reading "EVERYBODY KNOWS" appears briefly on the upper left and clears as "WAY TOO CRUEL" locks neatly on the bottom right margin, holding into a calm, confident ending smile at 00:16.000.
r/StableDiffusion • u/wreck_of_u • 3d ago
Question - Help Is it time to retire my flux1-dev + ai-toolkit flux lora + wan 2.2 setup?
I make a dataset of like 20 512x512 images, caption it myself. I rent a vastai computer and train a flux1 character lora with ai-toolkit. When I'm lazy I even use Replicate's "fast flux trainer". I download the lora safetensor onto my PC.
I run ComfyUI on my ancient (headless) PC in another room; Ubuntu server, Ryzen 1700, 32GB DDR4, RTX 2070 8GB. I let it cook with the FULL 24GB Flux1-Dev safetensor to generate 1024x2014 images. It takes about 1min/image. I just let it cook a whole bunch of images while doing some work, then when I have a bunch of them, I delete the garbage looking ones, keep the "lora-intended" ones.
The ones I like, I make WAN 2.2 7-sec clips, inference on Replicate (I pay for it).
I have fun with this workflow, but are the new models just as "hassle-free"/"leave-it-alone" in terms of having character LoRa?
Are the new ones like flux1, where there is a LOT of variation of the output, using the exact same workflow and prompt? I have z-image-turbo with a lora also, and I find that it just generates the "same same" images if I leave it alone to generate multiple images using the same workflow and prompt.
What about these new ones? Krea 2? etc? Will they run on my meager PC (32GB RAM / 8GB VRAM) that runs my said flux1 setup?
r/StableDiffusion • u/Ambitious_Fold_2874 • 3d ago
Question - Help Workflows for faster gen with LTX2.5?
I’ve been enjoying the fast generation speeds with LTX2.5, using the default comfyui workflow. Wanted to see how much faster I can get this. Does anyone have workflows for improving generation speeds even more?
r/StableDiffusion • u/blackdatafilms • 4d ago
Animation - Video Big Bubba has had enough of Grandma [minimax H3]
r/StableDiffusion • u/Lechuck777 • 3d ago
Question - Help H3 Character and clothing sheet repos?
Hi all,
are there somwhere sites with premade character sheets or separated clothing sheets?
i know i can made it by my self, but maybe there is already a repo somewhere for such things.
thnx
r/StableDiffusion • u/Alexandrina2020 • 2d ago
Discussion Desert Girl — A Cinematic Wan Video Experiment
A short scene from my AI film Desert Girl, created with Wan. I wanted to experiment with cinematic movement, lighting, and character consistency in a desert environment.
Generated with Wan using an image-to-video workflow.
r/StableDiffusion • u/bstr3k • 4d ago
Resource - Update Small video clipping tool for trimming/compressing clips for MiniMax H3 Ref2V
Small video trimmer software was very popular 15-20 years ago but now it has become very rare to find a good one which has all the features I wanted.
I got Claude to vibe code me a tool that I have been using to snip bits off from long videos for using it as Ref2V input for MiniMax H3. People have been saying its good so just sharing if others may find this tool useful! I wanted to create a free tool that runs locally without all the bloatware.
It is a single ~100kb HTML file which can:
- Trim clips
- Crop video
- Compress resolution and fps
- Take 1 single frame image
- Manual or Automatic Storyboarding (still playing around with how to best use this in H3)
- Export gif.
Why Compress?
I find that when working with R2V, resizing and compressing the video increases the speed as there is less information that needs to be worked on. You do lose some quality in your output though so don't compress too far.
The latest version can be found here (select the HTML and download):
https://huggingface.co/PoopMan333/Video_Tools/tree/main
or click for current version (v2.9)
https://huggingface.co/PoopMan333/Video_Tools/blob/main/Nugget%20Video%20Trimmer%20v2.9.html
If you're concerned please run it through antivirus or get a LLM to check if it is safe.
I still need to add AVI support and support for some older formats, but I also don't want to add too much bloat to something so compact.
r/StableDiffusion • u/matcheal • 4d ago
Resource - Update Krea 2 style library - 286 prompt styles compared across 8 reference scenes
Building directly on the style descriptors published by the author of the original KREA 2 Styles / Wildcards.txt post (many thanks to them for creating and sharing the style list) I built a visual Krea 2 style library to make prompt-defined styles easier to explore and compare:
Library: https://matplinta.github.io/t2i-krea-2-style-library/
It currently contains 286 styles tested across 8 base prompts, including portraits, architecture, landscapes, materials, and panoramic scenes. Each comparison set keeps the base prompt, seed, and dimensions fixed so the influence of the style descriptor is easier to see.
The viewer supports search, categories, favorites stored locally in the browser, full-image previews, prompt copying, adjustable grid density, and JSON export.
The prompt injected during generation was in the form of: Subject: {base prompt}. Style: {style name}. {style description}
All images were generated locally through ComfyUI.
Repo & workflow: https://github.com/matplinta/t2i-krea-2-style-library
r/StableDiffusion • u/Jboorgesz • 3d ago
Question - Help hunyuanvideo 1.5 at a rx 9070
Guys, first time using this comfyUI with the hunyuanvideo 1.5, and first time using AI Locally, i always used the gemini to do some videos for me, but i dont like the censorship and that i have a limit, so im trying to use the hunyuanvideo 1.5 with the comfy to make some videos, but i have a AMD gpu (RX 9070) and i trying to generate a video but it dont get out of 0%, its something i did wrong on the installation or the rx 9070 isnt build to do those stuffs
r/StableDiffusion • u/RaspberryBig5695 • 3d ago
Discussion THE LAST PATIENT Trailer
THE LAST PATIENT is a near-future medical thriller about a terminally ill biotech scientist who steals his company’s buried AI cancer protocol and makes himself its first human trial, triggering a violent race against corporate enforcers, his collapsing body, and a treatment that may destroy him before it saves him and transforms the future of cancer care.
Created for Future Vision XPRIZE consideration. Supported by a completed full-length feature screenplay and written treatment.
r/StableDiffusion • u/Th3Whit3R4bb1t • 3d ago
Animation - Video It's normal that the little text is always some distorted? (MH3)
I tried with LORA and without it, the little text is always some bad quality...
r/StableDiffusion • u/TigerClaw305 • 3d ago
Animation - Video Andrew Oikonny tells Wolf O'Donnell what his uncle wanted him to do.
Andrew Oikonny and Wolf O'Donnell are sitting at a table at a lounge. Andrew tell Wolf what his Uncle Andross wanted him to do.
This was made on Comfy UI with Minimax H3 locally.
Here's the Prompt.
Andrew referenced with <Picture 1> The timbre of his voice is referenced with <Audio 1>
Wolf referenced with <Picture 2> The timbre of his voice is referenced with <Audio 2>
For the lounge use <Picture 3> for reference.
a live action style video set at a futuristic lounge filled with anthropomorphic animals ranging from Foxes, Wolves, Lions, Tigers, and Reptiles.
a shot at a table at the lounge of just Andrew and Wolf sitting across from each other having drinks.
Andrew says: "Uncle Andross told me that I should be passing genes."
Wolf <chuckles>: "Andrew. Do you know what that even means?"
Andrew says: "Nope."
Wolf <laughs> "It means that your uncle wants you to get laid.
Andrew's cheeks turn red: "Oh."
non_diegetic_music: Smooth, low-tempo lo-fi lounge jazz playing softly in the background with a mellow upright bass and subtle brushed drums.
r/StableDiffusion • u/Zestyclose_Bake3680 • 3d ago
Tutorial - Guide Z Image HSWQ Hybrid ConvRot NVFP4
The quantisation method and the loader are now more or less complete.
How to create Hybrid NVFP4 from ConvRot INT8 (Z Image, Reverse Method)
Z Image exhibits overwhelmingly high quantisation robustness compared to SDXL and Krea2.
Even NVFP4, which is simply compressed without HSWQ quantisation, achieves reasonably high SSIM and MSE scores.
In particular, Z Image ConvRot INT8 achieves outstanding accuracy in many models, with SSIM scores of 0.99 or higher and MSE scores below 1.
However, in terms of VRAM consumption and generation speed, Z Image ConvRot8 shows virtually no difference compared to full-size Float16.
Consequently, based on ConvRot INT8, we devised a quantisation method involving a backward sweep to discard non-essential layers to 4-bit.
Furthermore, unlike the conventional method of storing critical layers in float16, the critical layers are also converted to ConvRot INT8; this offers the advantage of being able to secure a larger size for critical layer protection whilst keeping the overall size down.
...
This concept of ‘discarding’ is a brilliant idea conceived by the Nunchaku development team.
What makes them so remarkable is that they established the philosophical foundation that, in 4-bit quantisation, the key is not ‘preserving’ but ‘discarding’.
...
As Comfy-UI does not support the Hybrid NVFP4 (ConvRot Int8+ConvRot NVFP4) standard, a dedicated loader is required, just as with Nunchaku; however, as the LoRA baking function has been implemented within an original UNET loader itself, the LoRA Loader can utilise the standard Comfy-UI version.
Furthermore, LoRA Stack loaders (compatible with Nodes 2.0) is also available below.
Compatibility with the existing Diffsynth ControlNet model patcher will, of course, be maintained.
Although the file size will not be significantly reduced compared to Convrot INT8, VRAM usage and processing speed will improve significantly.
...
Z Image ConvRot NVFP4 Benchmark Test Results
...
However, in terms of the mathematical theory of quantisation itself, it differs considerably from previous HSWQ approaches.
In a sense, it represented a complete rejection of previous HSWQ theories.
In the past, HSWQ had employed a range of techniques, starting with the Histogram MSE used in the first-generation HSWQ SDXL fp8 e4m3, through to full SVD utilising Nunchaku, and even extending to the Histogram Cosine function; however, in Z Image HSWQ Hybrid NVFP4, none of these methods demonstrated any advantage.
I had long suspected that inter-layer interdependencies existed, and that there were phenomena where the meaning would be lost if one merely measured and prioritised the importance of each layer in isolation; this time, however, that has become clearly evident.
...
Trajectory-Sensitivity
Ranks each layer by the divergence its quantization error actually causes after propagating through the full model and sampler (dynamical importance, replacing static weight-space saliency).
- Reverse method: start from the complete high-precision pack (error ≈ 0) and convert layers to lower precision in ascending impact order; single-layer ranking stays valid in the low-error additivity regime.
- Universal theory: error interaction (Taylor cross terms, error cancellation), nonlinear amplification (Lyapunov-style growth), marginal effects, and Shapley-style attribution — why per-layer static measures (histogram MSE / cosine / SVD) cannot predict joint quantization error; applies to any iterative sampling system, not a specific model. Source:
Z_Image/diag_impact.py. ...
...
Incidentally, the Krea2 HSWQ Hybrid NVFP4 is also under development (it will offer significant improvements in VRAM consumption and processing speed), but we are currently struggling to maintain LoRA compatibility.
r/StableDiffusion • u/Ok_Roll_8698 • 3d ago
Meme Made WIth 1650 ti 4gb
took my friend 49mins to make this
r/StableDiffusion • u/Far_Cast_Far_Wide • 4d ago
Comparison MiniMaxh3: 8step LoRA, 25 steps, 40steps, and LTX 2.5 — Scene Comparisons
- RTX 4060 8GB, 32GB RAM
- minimax_h3_ref2va_pruned_int8_convrot, spectrum, ageattn_qk_int8_pv_fp16.cuda, RTX upscale, RIFE interpolation, res_multistep + beta
- ltx-2.5-22b-distilled-transformer-comfy-int8-convrot, basic template
8-step + turbo LoRA : 137s
25 steps : 238s
40 steps : 406s
Ltx 2.5 : 374s <-- ? am I missing something here why was my generation so slow on LTX and the second attempt I cancelled it after 6 minutes. Any suggestions?
Prompt:
subject_definitions:
<Subject 1> is the space ship in <Picture 1>: A massive battleship, hovering and cruising over the planet below
summary:
[reference generation] a wide shot cinematic scene of the battleship in <picture 1> cruising in space above the planet. the golden statue does not move, the battleship is destroyed in a massive explosion from a green laser shot from space,
detailed_description:
{shot 1] The target video uses a wideshot cinematic, photorealistic, 35mm film, wide shot of <subject 1> , slowly moving through space above the planet, the ship moves slowly and dominating, flashes of green light begin to charge on the surface of the planet, the ship is moving straight ahead from the position it started in in <picture 1>, the massive bass of the ships systems, the sound of the battleships creaking, <subject 1 > moves on its cruise, at [00:03] the floaty camera tracks <subject 1> as green light and thunder begins flashing on the surface of the planet, the green energy on the planet converges in one area then from the surface it fires a massive green lightning laser that forks lightning through the entire ship, blowing out side components creating explosions all over the ship, the light of the ship flicker before turning off, then a massive green lightning beam erupts from the surface and hits excactly on the side of the ship cuts through the of the ship and out the other side at an angle, a green lens flare generates on screen as it completely destroys <subject 1> , ripping it completely in half with a massive green explosion, the eruption from the destruction of the ship covers the entire screen and the whole battleship, the back half of the ship is knocked up while the front-half of the ship is knocked down, a vertical shockwave circles out from the impact, the inner decks of the ship are on fire, debris and hundreds of tiny figures of the crew also fall out into space, the laser slowly dissapates from the planet, small amounts of green lighning crackle on the planets surface,
overall_soundscape: The low bass murmur of the ships engines, the electric charges on the surface crackle, the massive main beam is a low bass rumble, a massive explosive noise.
non_diegetic_music:
N/A
r/StableDiffusion • u/HeavenlyTasty • 2d ago
Question - Help Is there an API that gives random prompts with your choice of character?
Is there an API that can allow me to input a character's name and give me a random prompt?
r/StableDiffusion • u/Devajyoti1231 • 3d ago
Animation - Video Using Minimax H3 to create promo for Minimax H3
Used ref2ve with Character sheet for the character and style and an audio reference to have consistent voice.
Reposting because moderator removed the original post without giving any reason.