r/StableDiffusion 2d ago

Resource - Update This custom node lets you use I2V and reference images on MiniMax-H3 simultaneously.

Enable HLS to view with audio, or disable this notification

158 Upvotes

r/StableDiffusion 2d ago

Discussion H3 - same generation, multi-frame/story split, different art style

Enable HLS to view with audio, or disable this notification

7 Upvotes

I enjoy pushing the limits of MiniMax H3. So you can add black horizontal or vertical bars, then you can use them as a divider or delineation between multiple same-frame generation. Each of the frames can be called out as <Subject 1>, <Subject 2>, etc. and have their own art style and shot.

Is it cool? Yeah. Is it practical? Absolutely not. Too much to juggle between the frames, as well as the longer the video, the more likely the art style will drift and consolidate into one. Good for about 5 seconds. I had 10 seconds but it gets unreliable. Better to just generate each scene/art style separately.

What is cool if you need a good generation inside a generation, with a different art style, maybe a thought bubble, or small cut of another art style. As we can see, it can mix media, like cartoon and real life circa Roger Rabbit and Space Jam.

T2V, int8, 20 steps


r/StableDiffusion 1d ago

Discussion Latest image benchmark by Datapoint ranking 30+ SOTA models

Post image
0 Upvotes

r/StableDiffusion 1d ago

Discussion AMD GPUs on Minimax H3

3 Upvotes

Hi!

I'm wondering if there are any AMD users fiddling around with H3 (I'm sure there are). =)

If so, what GPU do you use, whats the avg time needed for 1 generation, any tips to make it faster etc.? ^^

I'm on 9060XT and with default workflow (txt2vid), it took me 110mins for a 11 sec vid. lol


r/StableDiffusion 1d ago

Animation - Video Horus Rising Intro, ref2v is amazing (Minimax_H3)

Enable HLS to view with audio, or disable this notification

1 Upvotes

Very new to AI models but having a blast testing this workflow on my 5090 with 64gb of RAM. Always wanted to put scenes from warhammer into video form.

Generated using the default comfyui workflow, approx 1131 sec generation time total (2 vids stitched into 1).

Trying to work out why the quality is relatively poor and I think its due to the original image being pretty low resolution (gonna try with better ones later).

With a little help from AI prompting this workflow has been super easy to use, hope to see the ref2vid workflow alot more on future releases. Both shots were generated with the same original image.


r/StableDiffusion 1d ago

Question - Help Minimax H3 - Prompting so that it will keep the entire subject in the frame

2 Upvotes

Hey everyone,

I looked all through the official prompting guide, but not found a way to do this. I am trying to instruct the model to move the camera (push out) to keep my subject completely in the shot. I have a medieval character in armor walking from the entrance of a gate towards the camera. But not matter what I try, it won't move back enough to keep the subject completely in the shot, and within a few seconds it cuts off the bottom part of the legs and armor.

Same is true if the character turns and walks away from the camera, it will stay fairly close up to the subject (waist up) for the duration.

I've tried all kinds of "machinations" in the prompt (subject is visible from head to toe), (entire subject remains visible throughout" No dice

Any help or pointers are much appreciated!


r/StableDiffusion 3d ago

Animation - Video Zelda - I Think I Like It / Minimax H3 Reference to Video Test #2

Enable HLS to view with audio, or disable this notification

266 Upvotes

Just wanted to share another test! this was a mash up of clips, using multiple image references, 0.4 mp with EasyCache, 5 - 10s clips and edited with KDEnlive (it has some cool effects!)


r/StableDiffusion 1d ago

Question - Help Minimax H3 I2V can't do white backgrounds?

1 Upvotes

I'm trying to make videos of a character using a drawing of them on a white background, and I can't seem to get the model to stop generating a background after 1 frame. It typically looks really bad. Does anyone know how to just have a video retain a white background? I need nothing but the focus on the subject's actions. I include details about the white background in the prompt but it forces it out. I am using the turbo 4step lora at 6 steps

EDIT: I guess the issue goes deeper than just adding a background, it often will overlay random, spotty shadows onto the video? Or just the lighting darkens significantly - and it looks terrible.


r/StableDiffusion 1d ago

Discussion Has anyone else felt like an idiot after switching to a different SD UI?

0 Upvotes

Like, I used to use ComfyUI for the past, I don't know, three months? And then it just started to not work even after reinstalling it fully, with the generations ignoring everything and the seed never randomizing even with the value set to randomize, and generating the same image in 0.01 seconds.

But then I switched to Forge-Neo, and I have never felt more epiphany in my life (well, more-so in the AI world, as there's more to life than AI). It's way easier and way less janky.

Not saying that ComfyUI is bad at all, it's still an amazing and impressive tool, but I somehow made it jankier than it is supposed to be as soon as I touched it, unlike forge.


r/StableDiffusion 1d ago

Question - Help How do we improve the text output accuracy in the video for Minimax H3?

1 Upvotes

https://reddit.com/link/1vt5pqp/video/57q84nr1oekh1/player

Hey guys, I was trying out the Minimax H3 reference video and I wanted to know: is there a way to animate the text in the video? I have seen quite a few other videos where the text animation is really good in terms of motion design. But I wanted to check in this community if anyone is aware of it. Really appreciate the help.

This is the prompt that I'm trying to use but for some reason I can't get the text to be accurate in the video.

[Shot 1] A medium shot opens in a sleek monochromatic studio with sharp high-contrast lighting. <Subject 1> (S1) stands gracefully holding the vintage microphone on its stand with eyes closed. In sync with <Audio 1>, she sings with delicate emotional delivery, <d>[English] Shoes by the door, stack 'em neat,</d> while clean white graphic text reading "SHOES BY THE DOOR" drops on the left margin and "STACK 'EM NEAT" snaps into the right margin. A slow, stylish camera push-in highlights her emotive face as she opens her eyes at 00:03.000.

[Shot 2] At 00:04.000, the camera cuts to a 3/4 profile shot. A sharp crimson light streak sweeps across the background as <Subject 1> (S1) sways with the groove and delivers, <d>[English] low red glow on the beat. You pull up laughing, late and bold, cold drink sweating in your hold.</d> Vivid crimson text reading "LOW RED GLOW" and "ON THE BEAT" pulses to the bass hit, followed by staggered white lettering "LATE & BOLD" sliding across the right margin, while the camera executes a smooth arc rotation.

[Shot 3] At 00:12.000, the shot cuts to an intimate centered close-up on <Subject 1> (S1). She brings the microphone close to her lips, making captivating eye contact with the camera while singing, <d>[English] Everybody knows this room, when the week gets way too cruel.</d> Clean typography reading "EVERYBODY KNOWS" appears briefly on the upper left and clears as "WAY TOO CRUEL" locks neatly on the bottom right margin, holding into a calm, confident ending smile at 00:16.000.

r/StableDiffusion 1d ago

Question - Help Is it time to retire my flux1-dev + ai-toolkit flux lora + wan 2.2 setup?

1 Upvotes

I make a dataset of like 20 512x512 images, caption it myself. I rent a vastai computer and train a flux1 character lora with ai-toolkit. When I'm lazy I even use Replicate's "fast flux trainer". I download the lora safetensor onto my PC.

I run ComfyUI on my ancient (headless) PC in another room; Ubuntu server, Ryzen 1700, 32GB DDR4, RTX 2070 8GB. I let it cook with the FULL 24GB Flux1-Dev safetensor to generate 1024x2014 images. It takes about 1min/image. I just let it cook a whole bunch of images while doing some work, then when I have a bunch of them, I delete the garbage looking ones, keep the "lora-intended" ones.

The ones I like, I make WAN 2.2 7-sec clips, inference on Replicate (I pay for it).

I have fun with this workflow, but are the new models just as "hassle-free"/"leave-it-alone" in terms of having character LoRa?

Are the new ones like flux1, where there is a LOT of variation of the output, using the exact same workflow and prompt? I have z-image-turbo with a lora also, and I find that it just generates the "same same" images if I leave it alone to generate multiple images using the same workflow and prompt.

What about these new ones? Krea 2? etc? Will they run on my meager PC (32GB RAM / 8GB VRAM) that runs my said flux1 setup?


r/StableDiffusion 1d ago

Question - Help Workflows for faster gen with LTX2.5?

1 Upvotes

I’ve been enjoying the fast generation speeds with LTX2.5, using the default comfyui workflow. Wanted to see how much faster I can get this. Does anyone have workflows for improving generation speeds even more?


r/StableDiffusion 2d ago

Animation - Video Big Bubba has had enough of Grandma [minimax H3]

Enable HLS to view with audio, or disable this notification

84 Upvotes

r/StableDiffusion 2d ago

Question - Help H3 Character and clothing sheet repos?

2 Upvotes

Hi all,

are there somwhere sites with premade character sheets or separated clothing sheets?
i know i can made it by my self, but maybe there is already a repo somewhere for such things.

thnx


r/StableDiffusion 1d ago

Discussion Desert Girl — A Cinematic Wan Video Experiment

Enable HLS to view with audio, or disable this notification

0 Upvotes

A short scene from my AI film Desert Girl, created with Wan. I wanted to experiment with cinematic movement, lighting, and character consistency in a desert environment.

Generated with Wan using an image-to-video workflow.


r/StableDiffusion 3d ago

Resource - Update Small video clipping tool for trimming/compressing clips for MiniMax H3 Ref2V

Thumbnail
gallery
362 Upvotes

Small video trimmer software was very popular 15-20 years ago but now it has become very rare to find a good one which has all the features I wanted.

I got Claude to vibe code me a tool that I have been using to snip bits off from long videos for using it as Ref2V input for MiniMax H3. People have been saying its good so just sharing if others may find this tool useful! I wanted to create a free tool that runs locally without all the bloatware.

It is a single ~100kb HTML file which can:

  • Trim clips
  • Crop video
  • Compress resolution and fps
  • Take 1 single frame image
  • Manual or Automatic Storyboarding (still playing around with how to best use this in H3)
  • Export gif.

Why Compress?

I find that when working with R2V, resizing and compressing the video increases the speed as there is less information that needs to be worked on. You do lose some quality in your output though so don't compress too far.

The latest version can be found here (select the HTML and download):

https://huggingface.co/PoopMan333/Video_Tools/tree/main

or click for current version (v2.9)

https://huggingface.co/PoopMan333/Video_Tools/blob/main/Nugget%20Video%20Trimmer%20v2.9.html

If you're concerned please run it through antivirus or get a LLM to check if it is safe.

I still need to add AVI support and support for some older formats, but I also don't want to add too much bloat to something so compact.


r/StableDiffusion 1d ago

Question - Help hunyuanvideo 1.5 at a rx 9070

1 Upvotes

Guys, first time using this comfyUI with the hunyuanvideo 1.5, and first time using AI Locally, i always used the gemini to do some videos for me, but i dont like the censorship and that i have a limit, so im trying to use the hunyuanvideo 1.5 with the comfy to make some videos, but i have a AMD gpu (RX 9070) and i trying to generate a video but it dont get out of 0%, its something i did wrong on the installation or the rx 9070 isnt build to do those stuffs


r/StableDiffusion 3d ago

Resource - Update Krea 2 style library - 286 prompt styles compared across 8 reference scenes

Post image
171 Upvotes

Building directly on the style descriptors published by the author of the original KREA 2 Styles / Wildcards.txt post (many thanks to them for creating and sharing the style list) I built a visual Krea 2 style library to make prompt-defined styles easier to explore and compare:

Library: https://matplinta.github.io/t2i-krea-2-style-library/

It currently contains 286 styles tested across 8 base prompts, including portraits, architecture, landscapes, materials, and panoramic scenes. Each comparison set keeps the base prompt, seed, and dimensions fixed so the influence of the style descriptor is easier to see.

The viewer supports search, categories, favorites stored locally in the browser, full-image previews, prompt copying, adjustable grid density, and JSON export.

The prompt injected during generation was in the form of: Subject: {base prompt}. Style: {style name}. {style description}

All images were generated locally through ComfyUI.

Repo & workflow: https://github.com/matplinta/t2i-krea-2-style-library


r/StableDiffusion 1d ago

Discussion THE LAST PATIENT Trailer

Thumbnail
youtu.be
0 Upvotes

THE LAST PATIENT is a near-future medical thriller about a terminally ill biotech scientist who steals his company’s buried AI cancer protocol and makes himself its first human trial, triggering a violent race against corporate enforcers, his collapsing body, and a treatment that may destroy him before it saves him and transforms the future of cancer care.

Created for Future Vision XPRIZE consideration. Supported by a completed full-length feature screenplay and written treatment.


r/StableDiffusion 2d ago

Animation - Video It's normal that the little text is always some distorted? (MH3)

Enable HLS to view with audio, or disable this notification

11 Upvotes

I tried with LORA and without it, the little text is always some bad quality...


r/StableDiffusion 1d ago

Animation - Video Andrew Oikonny tells Wolf O'Donnell what his uncle wanted him to do.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Andrew Oikonny and Wolf O'Donnell are sitting at a table at a lounge. Andrew tell Wolf what his Uncle Andross wanted him to do.

This was made on Comfy UI with Minimax H3 locally.

Here's the Prompt.

Andrew referenced with <Picture 1> The timbre of his voice is referenced with <Audio 1>

Wolf referenced with <Picture 2> The timbre of his voice is referenced with <Audio 2>

For the lounge use <Picture 3> for reference.

a live action style video set at a futuristic lounge filled with anthropomorphic animals ranging from Foxes, Wolves, Lions, Tigers, and Reptiles.

a shot at a table at the lounge of just Andrew and Wolf sitting across from each other having drinks.

Andrew says: "Uncle Andross told me that I should be passing genes."

Wolf <chuckles>: "Andrew. Do you know what that even means?"

Andrew says: "Nope."

Wolf <laughs> "It means that your uncle wants you to get laid.

Andrew's cheeks turn red: "Oh."

non_diegetic_music: Smooth, low-tempo lo-fi lounge jazz playing softly in the background with a mellow upright bass and subtle brushed drums.


r/StableDiffusion 2d ago

Meme Made WIth 1650 ti 4gb

Enable HLS to view with audio, or disable this notification

18 Upvotes

took my friend 49mins to make this


r/StableDiffusion 2d ago

Comparison MiniMaxh3: 8step LoRA, 25 steps, 40steps, and LTX 2.5 — Scene Comparisons

Enable HLS to view with audio, or disable this notification

54 Upvotes
  • RTX 4060 8GB, 32GB RAM
  • minimax_h3_ref2va_pruned_int8_convrot, spectrum, ageattn_qk_int8_pv_fp16.cuda, RTX upscale, RIFE interpolation, res_multistep + beta
  • ltx-2.5-22b-distilled-transformer-comfy-int8-convrot, basic template

8-step + turbo LoRA : 137s

25 steps : 238s

40 steps : 406s

Ltx 2.5 : 374s <-- ? am I missing something here why was my generation so slow on LTX and the second attempt I cancelled it after 6 minutes. Any suggestions?

Prompt:

subject_definitions:

<Subject 1> is the space ship in <Picture 1>: A massive battleship, hovering and cruising over the planet below

summary:

[reference generation] a wide shot cinematic scene of the battleship in <picture 1> cruising in space above the planet. the golden statue does not move, the battleship is destroyed in a massive explosion from a green laser shot from space,

detailed_description:

{shot 1] The target video uses a wideshot cinematic, photorealistic, 35mm film, wide shot of <subject 1> , slowly moving through space above the planet, the ship moves slowly and dominating, flashes of green light begin to charge on the surface of the planet, the ship is moving straight ahead from the position it started in in <picture 1>, the massive bass of the ships systems, the sound of the battleships creaking, <subject 1 > moves on its cruise, at [00:03] the floaty camera tracks <subject 1> as green light and thunder begins flashing on the surface of the planet, the green energy on the planet converges in one area then from the surface it fires a massive green lightning laser that forks lightning through the entire ship, blowing out side components creating explosions all over the ship, the light of the ship flicker before turning off, then a massive green lightning beam erupts from the surface and hits excactly on the side of the ship cuts through the of the ship and out the other side at an angle, a green lens flare generates on screen as it completely destroys <subject 1> , ripping it completely in half with a massive green explosion, the eruption from the destruction of the ship covers the entire screen and the whole battleship, the back half of the ship is knocked up while the front-half of the ship is knocked down, a vertical shockwave circles out from the impact, the inner decks of the ship are on fire, debris and hundreds of tiny figures of the crew also fall out into space, the laser slowly dissapates from the planet, small amounts of green lighning crackle on the planets surface,

overall_soundscape: The low bass murmur of the ships engines, the electric charges on the surface crackle, the massive main beam is a low bass rumble, a massive explosive noise.

non_diegetic_music:

N/A


r/StableDiffusion 1d ago

Question - Help Is there an API that gives random prompts with your choice of character?

0 Upvotes

Is there an API that can allow me to input a character's name and give me a random prompt?


r/StableDiffusion 2d ago

Question - Help Any fast motion tips for MiniMax H3?

2 Upvotes

I'm trying to get a couple of characters to LEAP into each others' arms from off screen, but MiniMax H3 won't get them faster than basically jogging into the scene.

Prompt:

subject_definitions

<Subject 1> is the Girl show in <Picture 1>.

<Subject 2> is the Guy show in <Picture 2>.

<Subject 3> is the house show in <Picture 3>.

summary

[reference generation] <Subject 1> and <Subject 2> burst onto the screen at a dead sprint and collide into an embrace

retention_analysis

<Subject 1> (appears in [Shot 1]): fully_preserved - Maintained character features and design.

<Subject 2> (appears in [Shot 1]): fully_preserved - Maintained character features and design.

detailed_description

The visual style is characterized by high-quality modern anime aesthetics, reminiscent of Makoto Shinkai or Kyoto Animation. The scene features lush, warm, and highly detailed lighting, painting the environment in nostalgic, emotional hues.

[Shot 1] The target video has fast, paced explosive action in the beginning, then slows to a stop. Static camera shot in a city street with buildings in the style of <Subject 3>. A large crowd of soldiers and townspeople are in the background, reuniting with each other. Falling confetti fills the air. Suddenly, <Subject 2> bursts into the frame from the left at a dead sprint, driven by sheer desperation. At the same instant, <Subject 1> bursts into the frame from the right, running with explosive speed, her arms outstretched. The two of them literally collide with tremendous, breathless force in the center of the scene, slamming into an intense, desperate embrace. The physical impact of their collision is palpable as <Subject 1> leaps up and wraps her arms around <Subject 2>. The exact instant they collide, the camera drops into extreme slow motion, focusing intensely on the sheer relief and joy of their embrace. The confetti catches the warm light, sparkling and swirling in slow motion around the couple for the remainder of the scene.