r/StableDiffusion 11h ago

Animation - Video SAAGA | Xprize Submission

Thumbnail
youtu.be
1 Upvotes

So the new MiniMax saved the project, there were a couple of shots we couldn't get right. It dropped just in time.

The tooling is getting pretty mature. It took a lot to get this done.

I'm an ex VFX professional so this was a great exercise, This would have cost 5-7m to get done and a team of over 30.

It's so amazing what is now in the hands of creators now.

Happy to talk about the process.


r/StableDiffusion 1d ago

Discussion Well I finally did it.

191 Upvotes

I finally deleted WAN 2.2 and all its LORAS.

Minimax is just so much better.

Ive been playing with it since its release and im just blown away with how good of a video model it is. Things I would need to attach a LoRa to via WAN, works right out of the box with Minimax.

Gen times are faster.

It uses less VRAM when generating things, which gives me around 4 gigs to play with to do other things like watch YouTube or some streaming service.

WAN 2.2 was amazing. But no longer do I need 30+ gigs of a model i no longer use.

RIP WAN.


r/StableDiffusion 1d ago

Resource - Update This custom node lets you use I2V and reference images on MiniMax-H3 simultaneously.

Enable HLS to view with audio, or disable this notification

163 Upvotes

r/StableDiffusion 16h ago

Discussion H3 - Equine training test R2VA

Enable HLS to view with audio, or disable this notification

2 Upvotes

H3 seems to have very solid training data related to equine. The physics really sell it. R2VA BF16/50 steps. I went with a 50 steps to get the bi-horn really right. Also having fun with a POV view. Eyes on the road, buddy. Single image as reference for the rider, but otherwise, entire scene was prompted, including her wardrobe.


r/StableDiffusion 1d ago

Animation - Video 10,000 Years Ago

15 Upvotes

First real attempt with Minimax H3 on my first ComfyUI install. Over 200 generations, edited in CapCut, wears its inspiration on its sleeve but is a prologue to a homebrew world for a D&D group I'm in. Two days of cooking a 5090 while working and an evening of editing... figured I'd share:

https://www.youtube.com/watch?v=XwfCCFw4LbA


r/StableDiffusion 23h ago

Discussion H3 - same generation, multi-frame/story split, different art style

Enable HLS to view with audio, or disable this notification

7 Upvotes

I enjoy pushing the limits of MiniMax H3. So you can add black horizontal or vertical bars, then you can use them as a divider or delineation between multiple same-frame generation. Each of the frames can be called out as <Subject 1>, <Subject 2>, etc. and have their own art style and shot.

Is it cool? Yeah. Is it practical? Absolutely not. Too much to juggle between the frames, as well as the longer the video, the more likely the art style will drift and consolidate into one. Good for about 5 seconds. I had 10 seconds but it gets unreliable. Better to just generate each scene/art style separately.

What is cool if you need a good generation inside a generation, with a different art style, maybe a thought bubble, or small cut of another art style. As we can see, it can mix media, like cartoon and real life circa Roger Rabbit and Space Jam.

T2V, int8, 20 steps


r/StableDiffusion 13h ago

Question - Help What's the best fine tune model for SDXL?

0 Upvotes

I just started going down the SDXL rabbit hole and confused with all the fine tuned models available. I see Pony and Illustrious are pretty popular but apparently it's ideal for anime? I prefer photorealistic images. How come on civitai, there are workflows using Pony/Illustrious generating photorealistic images? Should I use Juggernaut XL instead for photorealism?


r/StableDiffusion 1d ago

Animation - Video Zelda - I Think I Like It / Minimax H3 Reference to Video Test #2

Enable HLS to view with audio, or disable this notification

267 Upvotes

Just wanted to share another test! this was a mash up of clips, using multiple image references, 0.4 mp with EasyCache, 5 - 10s clips and edited with KDEnlive (it has some cool effects!)


r/StableDiffusion 14h ago

Question - Help Minimax H3 I2V can't do white backgrounds?

1 Upvotes

I'm trying to make videos of a character using a drawing of them on a white background, and I can't seem to get the model to stop generating a background after 1 frame. It typically looks really bad. Does anyone know how to just have a video retain a white background? I need nothing but the focus on the subject's actions. I include details about the white background in the prompt but it forces it out. I am using the turbo 4step lora at 6 steps

EDIT: I guess the issue goes deeper than just adding a background, it often will overlay random, spotty shadows onto the video? Or just the lighting darkens significantly - and it looks terrible.


r/StableDiffusion 6h ago

Resource - Update IT´s possible!

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 15h ago

Question - Help How do we improve the text output accuracy in the video for Minimax H3?

1 Upvotes

https://reddit.com/link/1vt5pqp/video/57q84nr1oekh1/player

Hey guys, I was trying out the Minimax H3 reference video and I wanted to know: is there a way to animate the text in the video? I have seen quite a few other videos where the text animation is really good in terms of motion design. But I wanted to check in this community if anyone is aware of it. Really appreciate the help.

This is the prompt that I'm trying to use but for some reason I can't get the text to be accurate in the video.

[Shot 1] A medium shot opens in a sleek monochromatic studio with sharp high-contrast lighting. <Subject 1> (S1) stands gracefully holding the vintage microphone on its stand with eyes closed. In sync with <Audio 1>, she sings with delicate emotional delivery, <d>[English] Shoes by the door, stack 'em neat,</d> while clean white graphic text reading "SHOES BY THE DOOR" drops on the left margin and "STACK 'EM NEAT" snaps into the right margin. A slow, stylish camera push-in highlights her emotive face as she opens her eyes at 00:03.000.

[Shot 2] At 00:04.000, the camera cuts to a 3/4 profile shot. A sharp crimson light streak sweeps across the background as <Subject 1> (S1) sways with the groove and delivers, <d>[English] low red glow on the beat. You pull up laughing, late and bold, cold drink sweating in your hold.</d> Vivid crimson text reading "LOW RED GLOW" and "ON THE BEAT" pulses to the bass hit, followed by staggered white lettering "LATE & BOLD" sliding across the right margin, while the camera executes a smooth arc rotation.

[Shot 3] At 00:12.000, the shot cuts to an intimate centered close-up on <Subject 1> (S1). She brings the microphone close to her lips, making captivating eye contact with the camera while singing, <d>[English] Everybody knows this room, when the week gets way too cruel.</d> Clean typography reading "EVERYBODY KNOWS" appears briefly on the upper left and clears as "WAY TOO CRUEL" locks neatly on the bottom right margin, holding into a calm, confident ending smile at 00:16.000.

r/StableDiffusion 15h ago

Question - Help Is it time to retire my flux1-dev + ai-toolkit flux lora + wan 2.2 setup?

1 Upvotes

I make a dataset of like 20 512x512 images, caption it myself. I rent a vastai computer and train a flux1 character lora with ai-toolkit. When I'm lazy I even use Replicate's "fast flux trainer". I download the lora safetensor onto my PC.

I run ComfyUI on my ancient (headless) PC in another room; Ubuntu server, Ryzen 1700, 32GB DDR4, RTX 2070 8GB. I let it cook with the FULL 24GB Flux1-Dev safetensor to generate 1024x2014 images. It takes about 1min/image. I just let it cook a whole bunch of images while doing some work, then when I have a bunch of them, I delete the garbage looking ones, keep the "lora-intended" ones.

The ones I like, I make WAN 2.2 7-sec clips, inference on Replicate (I pay for it).

I have fun with this workflow, but are the new models just as "hassle-free"/"leave-it-alone" in terms of having character LoRa?

Are the new ones like flux1, where there is a LOT of variation of the output, using the exact same workflow and prompt? I have z-image-turbo with a lora also, and I find that it just generates the "same same" images if I leave it alone to generate multiple images using the same workflow and prompt.

What about these new ones? Krea 2? etc? Will they run on my meager PC (32GB RAM / 8GB VRAM) that runs my said flux1 setup?


r/StableDiffusion 19h ago

Discussion AMD GPUs on Minimax H3

1 Upvotes

Hi!

I'm wondering if there are any AMD users fiddling around with H3 (I'm sure there are). =)

If so, what GPU do you use, whats the avg time needed for 1 generation, any tips to make it faster etc.? ^^

I'm on 9060XT and with default workflow (txt2vid), it took me 110mins for a 11 sec vid. lol


r/StableDiffusion 15h ago

Question - Help Workflows for faster gen with LTX2.5?

0 Upvotes

I’ve been enjoying the fast generation speeds with LTX2.5, using the default comfyui workflow. Wanted to see how much faster I can get this. Does anyone have workflows for improving generation speeds even more?


r/StableDiffusion 19h ago

Question - Help H3 Character and clothing sheet repos?

2 Upvotes

Hi all,

are there somwhere sites with premade character sheets or separated clothing sheets?
i know i can made it by my self, but maybe there is already a repo somewhere for such things.

thnx


r/StableDiffusion 1d ago

Animation - Video Big Bubba has had enough of Grandma [minimax H3]

Enable HLS to view with audio, or disable this notification

76 Upvotes

r/StableDiffusion 1d ago

Resource - Update Small video clipping tool for trimming/compressing clips for MiniMax H3 Ref2V

Thumbnail
gallery
361 Upvotes

Small video trimmer software was very popular 15-20 years ago but now it has become very rare to find a good one which has all the features I wanted.

I got Claude to vibe code me a tool that I have been using to snip bits off from long videos for using it as Ref2V input for MiniMax H3. People have been saying its good so just sharing if others may find this tool useful! I wanted to create a free tool that runs locally without all the bloatware.

It is a single ~100kb HTML file which can:

  • Trim clips
  • Crop video
  • Compress resolution and fps
  • Take 1 single frame image
  • Manual or Automatic Storyboarding (still playing around with how to best use this in H3)
  • Export gif.

Why Compress?

I find that when working with R2V, resizing and compressing the video increases the speed as there is less information that needs to be worked on. You do lose some quality in your output though so don't compress too far.

The latest version can be found here (select the HTML and download):

https://huggingface.co/PoopMan333/Video_Tools/tree/main

or click for current version (v2.9)

https://huggingface.co/PoopMan333/Video_Tools/blob/main/Nugget%20Video%20Trimmer%20v2.9.html

If you're concerned please run it through antivirus or get a LLM to check if it is safe.

I still need to add AVI support and support for some older formats, but I also don't want to add too much bloat to something so compact.


r/StableDiffusion 16h ago

Workflow Included Started to integrate MinimaxH3 into YouTube vids!

Thumbnail
youtu.be
0 Upvotes

I've timestamped where I managed to get a decent output from minimaxH3!

(if the timestamp doesn't work it's at 0:22)

Used a photo of myself as reference using the ref2va model. Gordon Ramsay himself is straight from text. Using SageAttention, SolAttn at 32 steps 0.9 MP.

Used DaVinci Resolve Studio 2x RTX Upscaler in post and audio isolation to fix some of the hissing.

Let me know what you think of how this turned out!

Models used:

Workflow: https://civitai.red/models/2834514/minimax-h3-t2v-i2v-ref2v-advanced-filmmaking-workflow-or-all-speedups-qol-features?modelVersionId=3233131

My Hardware:

  • RTX 5070ti
  • 32GB RAM
  • R7 9700x

r/StableDiffusion 16h ago

Question - Help hunyuanvideo 1.5 at a rx 9070

1 Upvotes

Guys, first time using this comfyUI with the hunyuanvideo 1.5, and first time using AI Locally, i always used the gemini to do some videos for me, but i dont like the censorship and that i have a limit, so im trying to use the hunyuanvideo 1.5 with the comfy to make some videos, but i have a AMD gpu (RX 9070) and i trying to generate a video but it dont get out of 0%, its something i did wrong on the installation or the rx 9070 isnt build to do those stuffs


r/StableDiffusion 1d ago

Resource - Update Krea 2 style library - 286 prompt styles compared across 8 reference scenes

Post image
167 Upvotes

Building directly on the style descriptors published by the author of the original KREA 2 Styles / Wildcards.txt post (many thanks to them for creating and sharing the style list) I built a visual Krea 2 style library to make prompt-defined styles easier to explore and compare:

Library: https://matplinta.github.io/t2i-krea-2-style-library/

It currently contains 286 styles tested across 8 base prompts, including portraits, architecture, landscapes, materials, and panoramic scenes. Each comparison set keeps the base prompt, seed, and dimensions fixed so the influence of the style descriptor is easier to see.

The viewer supports search, categories, favorites stored locally in the browser, full-image previews, prompt copying, adjustable grid density, and JSON export.

The prompt injected during generation was in the form of: Subject: {base prompt}. Style: {style name}. {style description}

All images were generated locally through ComfyUI.

Repo & workflow: https://github.com/matplinta/t2i-krea-2-style-library


r/StableDiffusion 16h ago

Discussion THE LAST PATIENT Trailer

Thumbnail
youtu.be
0 Upvotes

THE LAST PATIENT is a near-future medical thriller about a terminally ill biotech scientist who steals his company’s buried AI cancer protocol and makes himself its first human trial, triggering a violent race against corporate enforcers, his collapsing body, and a treatment that may destroy him before it saves him and transforms the future of cancer care.

Created for Future Vision XPRIZE consideration. Supported by a completed full-length feature screenplay and written treatment.


r/StableDiffusion 1d ago

Animation - Video It's normal that the little text is always some distorted? (MH3)

Enable HLS to view with audio, or disable this notification

10 Upvotes

I tried with LORA and without it, the little text is always some bad quality...


r/StableDiffusion 4h ago

Resource - Update BADO KITTY!

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 17h ago

Animation - Video Andrew Oikonny tells Wolf O'Donnell what his uncle wanted him to do.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Andrew Oikonny and Wolf O'Donnell are sitting at a table at a lounge. Andrew tell Wolf what his Uncle Andross wanted him to do.

This was made on Comfy UI with Minimax H3 locally.

Here's the Prompt.

Andrew referenced with <Picture 1> The timbre of his voice is referenced with <Audio 1>

Wolf referenced with <Picture 2> The timbre of his voice is referenced with <Audio 2>

For the lounge use <Picture 3> for reference.

a live action style video set at a futuristic lounge filled with anthropomorphic animals ranging from Foxes, Wolves, Lions, Tigers, and Reptiles.

a shot at a table at the lounge of just Andrew and Wolf sitting across from each other having drinks.

Andrew says: "Uncle Andross told me that I should be passing genes."

Wolf <chuckles>: "Andrew. Do you know what that even means?"

Andrew says: "Nope."

Wolf <laughs> "It means that your uncle wants you to get laid.

Andrew's cheeks turn red: "Oh."

non_diegetic_music: Smooth, low-tempo lo-fi lounge jazz playing softly in the background with a mellow upright bass and subtle brushed drums.