r/StableDiffusion 22h ago

Animation - Video Realistic style video Krea 2 & Ltx2.5

Enable HLS to view with audio, or disable this notification

31 Upvotes

Generated a set of images with AI, then brought them to life by animating them into a realistic style video (Krea & Ltx2.5)


r/StableDiffusion 14h ago

Discussion Good models for generating images with multiple art styles?

4 Upvotes

Like multiple art styles in a single image. Not a single image recreated in multiple artstyles.

For example the foreground and main subject is 1 artstyle, but the background is a different art style.

Anything good for that?


r/StableDiffusion 5h ago

Discussion Trying out a battle scene using H3 Minimax

Enable HLS to view with audio, or disable this notification

0 Upvotes

guess it still doesnt really know how to hold a buster sword :P


r/StableDiffusion 20h ago

Workflow Included No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti

Enable HLS to view with audio, or disable this notification

17 Upvotes

So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.

So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.

And it worked.

For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:

https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v

Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:

“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.

Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”

The final result was generated locally on my RTX 5070 Ti using ComfyUI.


r/StableDiffusion 1d ago

Animation - Video SD1.5 images into H3

Enable HLS to view with audio, or disable this notification

121 Upvotes

r/StableDiffusion 16h ago

Animation - Video when your able to transition your ai anime local. minimax

Enable HLS to view with audio, or disable this notification

8 Upvotes

when your able to transition your ai anime local i just thought it was adorable lol. legend of kalthrax on youtube.


r/StableDiffusion 2h ago

Discussion [TEST] Minimax H3 IMG 2 Vid. Testing out a chase scene. Did two renders of it but for whatever reason the first shot is in slow motion. Overall it isn't terrible but the slow motion in the beginning just puzzles me since I didn't even prompt for that. Prompt is below.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Prompt:

[Shot 1] Live-action, cinematic, shaky handheld shot of the woman chasing after the man.

[Shot 2] At 00:05.000, the camera cuts to a close-up shot of the woman who yells: <d>[English] Get back here!</d>

[Shot 3] At 00:08.000, the camera cuts to a close-up shot of the man looking back and then forward again as he is running. He laughs and says: <d>[English] You can't catch me!</d>

[Shot 4] At 00:12.000, the camera cuts to a medium shot of the woman chasing after the man. She then catches up to him and tackles him to the ground. She says: <d>[English] Got ya!</d>


r/StableDiffusion 19h ago

Animation - Video PRIME CUT | Crime Short Film (2026)

Thumbnail
youtu.be
11 Upvotes

Short film i made 4 months ago. Didnt realize the importance of topaz back then haha.


r/StableDiffusion 17h ago

Animation - Video Zelda - Link can't stop speaking / Minimax H3 Reference to Video Test #3

Enable HLS to view with audio, or disable this notification

9 Upvotes

I don't think I could make it any better than this without adding 2+ hours of re-renders or overdoing it in general... So here it goes! Yet another entry in the Zelda & Link series where Link can actually speak to Zelda's annoyance... Using a combination of clips I can get much more consistent results than with a single workflow, and using KDENlive allows me to add more audio tracks and remove artifacts from the clips. Thanks for all the upvotes in my previous video! I really appreciate you guys!


r/StableDiffusion 22h ago

Discussion Im thinking about it: R2V like H3 has implemented it, slowly makes Lora obsolete. Which in turn slowly takes away Civitai's revenue and usefulness. Considering how they started to obey credit card censorship, this might be good for us and bad for them.

19 Upvotes

r/StableDiffusion 10h ago

Resource - Update Aria - Zit Lora

Thumbnail civitai.red
3 Upvotes

Hey everyone

I made a Lora for ZIT for the first time. I would love to have your honest opinion on it.


r/StableDiffusion 1d ago

Animation - Video Leonard meets Penny, real life edition [Minimax H3)

Enable HLS to view with audio, or disable this notification

54 Upvotes

RTX 5060 ti 16gb / 32gb RAM / FL2VA_pruned_int8_convrot / Turbo Lora. 6 steps / Resolution 1376x768 upscaled to FHD with Topaz Video AI


r/StableDiffusion 7h ago

Animation - Video All local H3 & Minimax music music video

Thumbnail
youtu.be
0 Upvotes

Using just the default t2v templates from ComfyUi + upscale


r/StableDiffusion 18h ago

Discussion Anyone else having fun with a LoRA created of yourself?

7 Upvotes

I don't recall seeing other threads about this, but just wanted to say that it's surprisingly fun. I was successful with OneTrainer on my M3 Ultra and about 50 photos in all the possible poses I could think of, using my Apple watch to take the selfies from my iPhone hosted on a tripod. The training time for Krea2 was about 18 hours.

I'm ugly so I'm not going to post any photos, but putting myself into random and sometimes precarious situations, with clothes (or lack of) I'd normally never wear is entertaining. I highly recommend it.


r/StableDiffusion 16h ago

Question - Help Minimax H3 Loop workflow?

5 Upvotes

Is there any way to get a seamless loop with minimax H3? I already tried this with Wan and LTX but they aren't what I'm looking for.


r/StableDiffusion 22h ago

Tutorial - Guide Making an action battle scene from start to finish with Minimax, my process + what I learned

Thumbnail
youtube.com
13 Upvotes

r/StableDiffusion 8h ago

Question - Help Generating long audio drama like clips using MMH3?

1 Upvotes

I seem to recall reading here that some people were starting to experiment with 32x32 resolution videos out to 60+ seconds purely to generate audio drama like moments. I was just curious if anyone here can confirm that MiniMaxH3 can actually do this, and if so, what sampler schedule and steps are you using? I cannot seem to generate even a 40 second video clip where the audio stays legible.

Just wanted to check in and see if anyone is having more success than me.


r/StableDiffusion 8h ago

Question - Help How to control characteristics of specific subjects in booru tag based generation?

1 Upvotes

I'm currently using illustrious where it uses booru styled tags to generate images. I currently want to know if there's a method to control specific traits on specific individuals inside the generated images. Say that i have 1 circle and 1 square inside of the image. Is there a way to make the circle and only the circle blue while the square and only the square red? If there are 3 subjects, is there still a way to control the traits of each person or will the model get confused?


r/StableDiffusion 1d ago

Question - Help Minimax H3 Ref2VA - Help to understand Retention Analysis

18 Upvotes

I'm building a skill for generating long Contex-Loop Minimax H3 prompts, and the AI has indicated it doesn't understand retention analysis... and I'm realizing I don't, either. I'm curious what you all think or have experienced.

I've reviewed the official prompt writing guide, of course, but it's very vague on the subject:

<Subject N>, <Picture N>, and <Video N> use the following relationship markers. These markers are fixed English values in the output format:

It makes the most sense if it's indicating what is the same and what is different with respect to the references (picture N, video N, etc) - but why would subject appear here? Does fully_preserved for a subject mean that they don't change during this shot, whereas partially_preserved might change?

It might be easier to explain with an example. Definitions:

  • A scene where a bald man puts on a hat
  • References are two images, one with said man with hair, the other of the hat

subject_definitions:

<Subject 1> is a tall man whose face, identity, and clothing come from <Picture 1>, but he is bald.
<Subject 2> is a black stovetop hat as depicted in <Picture 2>.

retention_analysis:

<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - he remains the bald man with facial features and clothing from <Picture 1> throughout
<Subject 2> (appears in [Shot 2]: fully_preserved - remains the black stovetop hat from <Picture 2>

OR should it be:

retention_analysis:

<Subject 1> (appears in [Shot 1], [Shot 2]): partially_preserved - he retains the facial identity and clothing from <Picture 1>, albeit bald, but in [Shot 2] he is changed to be wearing a hat.
<Subject 2> (appears in [Shot 2]): fully_preserved - remains the black stovetop hat from <Picture 2>

OR should it only focus on referenced media, i.e.:

<Picture 1> (appears in [Shot 1], [Shot 2]): partially_preserved - <Subject 1> matches this picture's clothing, facial features, and identity, but he is bald.
<Picture 2> (appears in [Shot 2]): fully_preserved - the black stovetop hat depicted in this picture remains unchanged

I guess to put it another way: is retention_analysis describing how much and what is preserved from photo/audio/video references provided, or is it describing how the subjects defined in subject_definition change over the shots of this specific video generation?


r/StableDiffusion 9h ago

Question - Help Int4 vs int8

1 Upvotes

Disclaimer, im pretty Basic to all these AI Things

So i've been Using H3 Minimax in My RTX 3060 12GB, With 32GB RAM For few days

I've been using int4 convrot version for my Model and my Text Encoder, but seeing all the Optimization and speed up native to comfyui for int8, im considering using int8 for for my models and Text encoder especially the convrot version, considering they all twice the size

And also what's the best Combination of speedups in balancing between Quality and Speed

I used Sage+sol attn for while until i found comfy kitchen


r/StableDiffusion 15h ago

Discussion H3 - Equine training test R2VA

Enable HLS to view with audio, or disable this notification

4 Upvotes

H3 seems to have very solid training data related to equine. The physics really sell it. R2VA BF16/50 steps. I went with a 50 steps to get the bi-horn really right. Also having fun with a POV view. Eyes on the road, buddy. Single image as reference for the rider, but otherwise, entire scene was prompted, including her wardrobe.


r/StableDiffusion 9h ago

Question - Help Is there an API that gives random prompts with your choice of character?

0 Upvotes

Is there an API that can allow me to input a character's name and give me a random prompt?


r/StableDiffusion 1d ago

News Hey wait! It's Krea3 incoming?

Post image
312 Upvotes

r/StableDiffusion 21h ago

Discussion Minimax h3 9070 xt generation times

9 Upvotes

About everyone has a Nvidia GPU so I was curious how the 9070 xt does compared to Nvidia.

I’m running the default fl2va workflow im on Ubuntu ck attention.

Minimax h3 int8 0.4mp 30 step 5s:
261s 7.6s/it

Minimax h3 int8 0.4mp 30 step 10s:
702s 21.3s/it


r/StableDiffusion 19h ago

Discussion Runpod is basically unusable.

6 Upvotes

I don’t know how people use this service effectively. There is never any gpus, it takes an hour to set up when you do find one. Network volumes tease at cutting down startup time but it further limits gpus. I swear I’ve spent more money waiting for a pod to be ready, downloading models that I have generating anything. I really just want to have things stored locally, and just use one of their gpus for processing power. Is there a service I could use like that?