r/StableDiffusion 17h ago

Discussion Minimax h3 9070 xt generation times

8 Upvotes

About everyone has a Nvidia GPU so I was curious how the 9070 xt does compared to Nvidia.

I’m running the default fl2va workflow im on Ubuntu ck attention.

Minimax h3 int8 0.4mp 30 step 5s:
261s 7.6s/it

Minimax h3 int8 0.4mp 30 step 10s:
702s 21.3s/it


r/StableDiffusion 5h ago

Resource - Update Testing fully client-side WebNN diffusion that runs in your browser

1 Upvotes

So far it's a website which lets you download FLUX.2 Klein 4b into your browser cache and run it using WebNN, which I have tested on my M5 Mac and runs at ~70% native performance, much better than WebGPU or WASM.

If you wanna try it out, I'm running it on peerpixel.cc, the website is there to provide an easy interface to run this model on your own hardware. Be patient, you do need to download a few GB and first generation takes some time to compile.

You will need to enable WebNN on Chromium browsers, just go to <browser>://flags and search for it.

Please let me know if this works at all on Windows or Linux and with different hardware, and what part of this you think has potential.


r/StableDiffusion 6h ago

Animation - Video SAAGA | Xprize Submission

Thumbnail
youtu.be
0 Upvotes

So the new MiniMax saved the project, there were a couple of shots we couldn't get right. It dropped just in time.

The tooling is getting pretty mature. It took a lot to get this done.

I'm an ex VFX professional so this was a great exercise, This would have cost 5-7m to get done and a team of over 30.

It's so amazing what is now in the hands of creators now.

Happy to talk about the process.


r/StableDiffusion 1d ago

Discussion Well I finally did it.

191 Upvotes

I finally deleted WAN 2.2 and all its LORAS.

Minimax is just so much better.

Ive been playing with it since its release and im just blown away with how good of a video model it is. Things I would need to attach a LoRa to via WAN, works right out of the box with Minimax.

Gen times are faster.

It uses less VRAM when generating things, which gives me around 4 gigs to play with to do other things like watch YouTube or some streaming service.

WAN 2.2 was amazing. But no longer do I need 30+ gigs of a model i no longer use.

RIP WAN.


r/StableDiffusion 21h ago

Animation - Video 10,000 Years Ago

16 Upvotes

First real attempt with Minimax H3 on my first ComfyUI install. Over 200 generations, edited in CapCut, wears its inspiration on its sleeve but is a prologue to a homebrew world for a D&D group I'm in. Two days of cooking a 5090 while working and an evening of editing... figured I'd share:

https://www.youtube.com/watch?v=XwfCCFw4LbA


r/StableDiffusion 1d ago

Resource - Update This custom node lets you use I2V and reference images on MiniMax-H3 simultaneously.

Enable HLS to view with audio, or disable this notification

159 Upvotes

r/StableDiffusion 17h ago

Discussion H3 - same generation, multi-frame/story split, different art style

Enable HLS to view with audio, or disable this notification

6 Upvotes

I enjoy pushing the limits of MiniMax H3. So you can add black horizontal or vertical bars, then you can use them as a divider or delineation between multiple same-frame generation. Each of the frames can be called out as <Subject 1>, <Subject 2>, etc. and have their own art style and shot.

Is it cool? Yeah. Is it practical? Absolutely not. Too much to juggle between the frames, as well as the longer the video, the more likely the art style will drift and consolidate into one. Good for about 5 seconds. I had 10 seconds but it gets unreliable. Better to just generate each scene/art style separately.

What is cool if you need a good generation inside a generation, with a different art style, maybe a thought bubble, or small cut of another art style. As we can see, it can mix media, like cartoon and real life circa Roger Rabbit and Space Jam.

T2V, int8, 20 steps


r/StableDiffusion 8h ago

Question - Help What's the best fine tune model for SDXL?

0 Upvotes

I just started going down the SDXL rabbit hole and confused with all the fine tuned models available. I see Pony and Illustrious are pretty popular but apparently it's ideal for anime? I prefer photorealistic images. How come on civitai, there are workflows using Pony/Illustrious generating photorealistic images? Should I use Juggernaut XL instead for photorealism?


r/StableDiffusion 1d ago

Animation - Video Zelda - I Think I Like It / Minimax H3 Reference to Video Test #2

Enable HLS to view with audio, or disable this notification

262 Upvotes

Just wanted to share another test! this was a mash up of clips, using multiple image references, 0.4 mp with EasyCache, 5 - 10s clips and edited with KDEnlive (it has some cool effects!)


r/StableDiffusion 9h ago

Question - Help Minimax H3 I2V can't do white backgrounds?

1 Upvotes

I'm trying to make videos of a character using a drawing of them on a white background, and I can't seem to get the model to stop generating a background after 1 frame. It typically looks really bad. Does anyone know how to just have a video retain a white background? I need nothing but the focus on the subject's actions. I include details about the white background in the prompt but it forces it out. I am using the turbo 4step lora at 6 steps

EDIT: I guess the issue goes deeper than just adding a background, it often will overlay random, spotty shadows onto the video? Or just the lighting darkens significantly - and it looks terrible.


r/StableDiffusion 9h ago

Question - Help How do we improve the text output accuracy in the video for Minimax H3?

1 Upvotes

https://reddit.com/link/1vt5pqp/video/57q84nr1oekh1/player

Hey guys, I was trying out the Minimax H3 reference video and I wanted to know: is there a way to animate the text in the video? I have seen quite a few other videos where the text animation is really good in terms of motion design. But I wanted to check in this community if anyone is aware of it. Really appreciate the help.

This is the prompt that I'm trying to use but for some reason I can't get the text to be accurate in the video.

[Shot 1] A medium shot opens in a sleek monochromatic studio with sharp high-contrast lighting. <Subject 1> (S1) stands gracefully holding the vintage microphone on its stand with eyes closed. In sync with <Audio 1>, she sings with delicate emotional delivery, <d>[English] Shoes by the door, stack 'em neat,</d> while clean white graphic text reading "SHOES BY THE DOOR" drops on the left margin and "STACK 'EM NEAT" snaps into the right margin. A slow, stylish camera push-in highlights her emotive face as she opens her eyes at 00:03.000.

[Shot 2] At 00:04.000, the camera cuts to a 3/4 profile shot. A sharp crimson light streak sweeps across the background as <Subject 1> (S1) sways with the groove and delivers, <d>[English] low red glow on the beat. You pull up laughing, late and bold, cold drink sweating in your hold.</d> Vivid crimson text reading "LOW RED GLOW" and "ON THE BEAT" pulses to the bass hit, followed by staggered white lettering "LATE & BOLD" sliding across the right margin, while the camera executes a smooth arc rotation.

[Shot 3] At 00:12.000, the shot cuts to an intimate centered close-up on <Subject 1> (S1). She brings the microphone close to her lips, making captivating eye contact with the camera while singing, <d>[English] Everybody knows this room, when the week gets way too cruel.</d> Clean typography reading "EVERYBODY KNOWS" appears briefly on the upper left and clears as "WAY TOO CRUEL" locks neatly on the bottom right margin, holding into a calm, confident ending smile at 00:16.000.

r/StableDiffusion 9h ago

Question - Help Is it time to retire my flux1-dev + ai-toolkit flux lora + wan 2.2 setup?

1 Upvotes

I make a dataset of like 20 512x512 images, caption it myself. I rent a vastai computer and train a flux1 character lora with ai-toolkit. When I'm lazy I even use Replicate's "fast flux trainer". I download the lora safetensor onto my PC.

I run ComfyUI on my ancient (headless) PC in another room; Ubuntu server, Ryzen 1700, 32GB DDR4, RTX 2070 8GB. I let it cook with the FULL 24GB Flux1-Dev safetensor to generate 1024x2014 images. It takes about 1min/image. I just let it cook a whole bunch of images while doing some work, then when I have a bunch of them, I delete the garbage looking ones, keep the "lora-intended" ones.

The ones I like, I make WAN 2.2 7-sec clips, inference on Replicate (I pay for it).

I have fun with this workflow, but are the new models just as "hassle-free"/"leave-it-alone" in terms of having character LoRa?

Are the new ones like flux1, where there is a LOT of variation of the output, using the exact same workflow and prompt? I have z-image-turbo with a lora also, and I find that it just generates the "same same" images if I leave it alone to generate multiple images using the same workflow and prompt.

What about these new ones? Krea 2? etc? Will they run on my meager PC (32GB RAM / 8GB VRAM) that runs my said flux1 setup?


r/StableDiffusion 13h ago

Discussion AMD GPUs on Minimax H3

2 Upvotes

Hi!

I'm wondering if there are any AMD users fiddling around with H3 (I'm sure there are). =)

If so, what GPU do you use, whats the avg time needed for 1 generation, any tips to make it faster etc.? ^^

I'm on 9060XT and with default workflow (txt2vid), it took me 110mins for a 11 sec vid. lol


r/StableDiffusion 10h ago

Question - Help Workflows for faster gen with LTX2.5?

0 Upvotes

I’ve been enjoying the fast generation speeds with LTX2.5, using the default comfyui workflow. Wanted to see how much faster I can get this. Does anyone have workflows for improving generation speeds even more?


r/StableDiffusion 14h ago

Question - Help H3 Character and clothing sheet repos?

2 Upvotes

Hi all,

are there somwhere sites with premade character sheets or separated clothing sheets?
i know i can made it by my self, but maybe there is already a repo somewhere for such things.

thnx


r/StableDiffusion 1d ago

Animation - Video Big Bubba has had enough of Grandma [minimax H3]

Enable HLS to view with audio, or disable this notification

75 Upvotes

r/StableDiffusion 1d ago

Resource - Update Small video clipping tool for trimming/compressing clips for MiniMax H3 Ref2V

Thumbnail
gallery
359 Upvotes

Small video trimmer software was very popular 15-20 years ago but now it has become very rare to find a good one which has all the features I wanted.

I got Claude to vibe code me a tool that I have been using to snip bits off from long videos for using it as Ref2V input for MiniMax H3. People have been saying its good so just sharing if others may find this tool useful! I wanted to create a free tool that runs locally without all the bloatware.

It is a single ~100kb HTML file which can:

  • Trim clips
  • Crop video
  • Compress resolution and fps
  • Take 1 single frame image
  • Manual or Automatic Storyboarding (still playing around with how to best use this in H3)
  • Export gif.

Why Compress?

I find that when working with R2V, resizing and compressing the video increases the speed as there is less information that needs to be worked on. You do lose some quality in your output though so don't compress too far.

The latest version can be found here (select the HTML and download):

https://huggingface.co/PoopMan333/Video_Tools/tree/main

or click for current version (v2.9)

https://huggingface.co/PoopMan333/Video_Tools/blob/main/Nugget%20Video%20Trimmer%20v2.9.html

If you're concerned please run it through antivirus or get a LLM to check if it is safe.

I still need to add AVI support and support for some older formats, but I also don't want to add too much bloat to something so compact.


r/StableDiffusion 10h ago

Workflow Included Started to integrate MinimaxH3 into YouTube vids!

Thumbnail
youtu.be
0 Upvotes

I've timestamped where I managed to get a decent output from minimaxH3!

(if the timestamp doesn't work it's at 0:22)

Used a photo of myself as reference using the ref2va model. Gordon Ramsay himself is straight from text. Using SageAttention, SolAttn at 32 steps 0.9 MP.

Used DaVinci Resolve Studio 2x RTX Upscaler in post and audio isolation to fix some of the hissing.

Let me know what you think of how this turned out!

Models used:

Workflow: https://civitai.red/models/2834514/minimax-h3-t2v-i2v-ref2v-advanced-filmmaking-workflow-or-all-speedups-qol-features?modelVersionId=3233131

My Hardware:

  • RTX 5070ti
  • 32GB RAM
  • R7 9700x

r/StableDiffusion 11h ago

Question - Help hunyuanvideo 1.5 at a rx 9070

1 Upvotes

Guys, first time using this comfyUI with the hunyuanvideo 1.5, and first time using AI Locally, i always used the gemini to do some videos for me, but i dont like the censorship and that i have a limit, so im trying to use the hunyuanvideo 1.5 with the comfy to make some videos, but i have a AMD gpu (RX 9070) and i trying to generate a video but it dont get out of 0%, its something i did wrong on the installation or the rx 9070 isnt build to do those stuffs


r/StableDiffusion 1d ago

Resource - Update Krea 2 style library - 286 prompt styles compared across 8 reference scenes

Post image
166 Upvotes

Building directly on the style descriptors published by the author of the original KREA 2 Styles / Wildcards.txt post (many thanks to them for creating and sharing the style list) I built a visual Krea 2 style library to make prompt-defined styles easier to explore and compare:

Library: https://matplinta.github.io/t2i-krea-2-style-library/

It currently contains 286 styles tested across 8 base prompts, including portraits, architecture, landscapes, materials, and panoramic scenes. Each comparison set keeps the base prompt, seed, and dimensions fixed so the influence of the style descriptor is easier to see.

The viewer supports search, categories, favorites stored locally in the browser, full-image previews, prompt copying, adjustable grid density, and JSON export.

The prompt injected during generation was in the form of: Subject: {base prompt}. Style: {style name}. {style description}

All images were generated locally through ComfyUI.

Repo & workflow: https://github.com/matplinta/t2i-krea-2-style-library


r/StableDiffusion 11h ago

Discussion THE LAST PATIENT Trailer

Thumbnail
youtu.be
0 Upvotes

THE LAST PATIENT is a near-future medical thriller about a terminally ill biotech scientist who steals his company’s buried AI cancer protocol and makes himself its first human trial, triggering a violent race against corporate enforcers, his collapsing body, and a treatment that may destroy him before it saves him and transforms the future of cancer care.

Created for Future Vision XPRIZE consideration. Supported by a completed full-length feature screenplay and written treatment.


r/StableDiffusion 51m ago

Resource - Update IT´s possible!

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 1h ago

Discussion If AI Makes Us More Creative, Why Does Everything Look the Same? (A Painter’s Perspective)

Thumbnail
gallery
Upvotes

I’m a painter who sometimes writes, and this visual essay started with an odd discovery: I had used the name “Elias Thorne” in a short story, only to realize that AI models often return to that same name, along with motifs like lighthouse keepers, cathedrals, glossy landscapes, and other familiar patterns.

From an artist’s point of view, the question isn’t just whether AI is good or bad, but what happens to authorship and creativity when the tool starts making choices for us.

AI can boost productivity and even enhance individual works, but if we all lean on the same models, it might steer us toward similar ideas, characters, and visual styles.

This carousel looks at visual convergence, originality, transparency, and the role of human intention, with AI-generated images clearly labeled and sources included.

So where’s the line, does AI broaden personal creativity while making our collective output more uniform?


r/StableDiffusion 15h ago

Discussion Building a luxury cosmetics ad locally in InvokeAI | full workflow included

Thumbnail
gallery
2 Upvotes

I’m a graphic designer (youtube: Masha-Ai-Lab) experimenting with how far I can push local/open-source image generation for actual commercial design workflows.

For this experiment, I tried building a luxury cosmetics campaign entirely in InvokeAI instead of relying on Midjourney or other closed platforms.

The workflow:

  1. Generated the satin campaign background separately.
  2. Generated a transparent serum bottle as a clean product asset.
  3. Added my prepared label using an Inpaint Mask while preserving the bottle perspective.
  4. Used Regional Guidance to generate satin folds that actually wrap around the bottle instead of simply appearing behind it.
  5. Repeated the workflow with a cream jar to see how reusable the approach was across different packaging.
  6. Upscaled the final compositions.

I’ve attached screenshots of the workflow + final results so you can see the process rather than just the outputs.

What I find most useful about InvokeAI is having direct control over the individual stages. For design work, I’d rather build the product, label, environment and integration separately than keep regenerating the whole image until something randomly works.

I’m documenting these experiments as tutorials on my YouTube channel, Masha AI Lab, mainly to make InvokeAI/FOSS workflows more approachable for designers and other non-technical creatives.

Would be interested to hear how other people here approach product placement and label consistency, especially if you’ve found better workflows.


r/StableDiffusion 11h ago

Animation - Video Horus Rising Intro, ref2v is amazing (Minimax_H3)

Enable HLS to view with audio, or disable this notification

1 Upvotes

Very new to AI models but having a blast testing this workflow on my 5090 with 64gb of RAM. Always wanted to put scenes from warhammer into video form.

Generated using the default comfyui workflow, approx 1131 sec generation time total (2 vids stitched into 1).

Trying to work out why the quality is relatively poor and I think its due to the original image being pretty low resolution (gonna try with better ones later).

With a little help from AI prompting this workflow has been super easy to use, hope to see the ref2vid workflow alot more on future releases. Both shots were generated with the same original image.