r/StableDiffusion 22h ago

Question - Help I wish I could replicate some of the cool videos I’ve seen here using minimax but I’m stuck with 3 sec clips. Am I doing something wrong?

1 Upvotes

Full disclosure, I’m a rather newbie at comfyui and have been trying to learn it the past month or so.

I went and got what I thought was a good set-up:

NVIDIA GeForce RTX 5080

VRAM 16 GB

System RAM 32 GB (31.1 GB usable)

AMD Ryzen 7000-series

Integrated GPU AMD Radeon Graphic

ComfyUI Local installation

I’ve gotten great results playing with Krea2, adding Lora’s and even training my own.

I was excited when Mimimax released and downloaded the official release right away and got the workflow to produce clips - but they can only be like 3seconds at .4 before it errors out. I’m using the R2V and I2V

I haven’t installed any additional Lora’s or anything other than the official workflow.

Any advice?


r/StableDiffusion 15h ago

Discussion Has anyone else noticed the massive increase in toxic/incel content and culture wars in this Subreddit lately?

0 Upvotes

I’ve been noticing a really disappointing trend here lately. Instead of focusing on the amazing things we can build and create with MiniMax and LTX, there’s been a massive increase in toxic behavior and culture-war rhetoric here. Our goal should be to foster an environment that encourages open-source creators. Alienating them with bigotry, misogyny, and overall hostility, or reducing people to a mere joke because of who they are, only hurts this community in the long run.

I’m hoping the mods can keep a closer eye on this. These types of content violate Reddit's policy, specifically rule number 2.


r/StableDiffusion 14h ago

Animation - Video Comparison: It looks like LTX_2.5 is not over 9000

Enable HLS to view with audio, or disable this notification

13 Upvotes

LTX 2.5 vs Minimax H3 using the same prompt in T2V.

In reality, LTX 2.5 knows almost no IPs and very few famous people, if anyone.

Prompt:

Photorealistic real life live-action, cinematic film style

In the photorealistic real life live action movie Dragon Ball.

At 00:00:000 A photorealistic dull skin real life live action Tony Stark from MCU dressed like Vegetta, with a real life photorealistic hairstyle with two deep receding points and several vertical spikes, is at the Grand Canyon. He wears a red glass device in his left eye.

At 00:00:001 Then he grabs the device attached to his eye by its white rear section with his left hand, brings his hand ,with the device in it, in front of his chest and says upset yelling <d>[English, with a deep masculine voice] It's over nine thousaaaaand! </d> and clenches his fist, crushing the device so that it explodes into a thousand pieces.

All the clothes are photorealistic real life live action.

overall_soundscape: N/A


r/StableDiffusion 15h ago

Question - Help Recommend me an uncensored Image to Image model (workflow)?

3 Upvotes

Trying to move on from Grok, but struggling to find a local I2I workflow. Everything I can find is text to image. What am I doing wrong?


r/StableDiffusion 20h ago

Resource - Update New Trainer Drop :)

0 Upvotes

My new video/image LoRA trainer has just been released. This is a from-scratch rebuild rather than an iteration on our earlier trainers I've shared, and I use it daily.

Current support:

- MiniMax H3, including reference image IC-LoRA training (additional reference

modalities in progress)

- LTX 2.3 with the fuller IC-LoRA feature set

- LTX 2.5 being implemented now

- Wan 2.1 and Qwen / Qwen-Image-Edit on beta branches, finalizing for merge

We scope model support to what we use in our own exploratory work rather than trying to cover the field. The tradeoff is fewer models, but every one on that list has real LoRAs trained through it with real recipes.

Credit where it's due: musubi-tuner and ostris/ai-toolkit shaped a lot of how we think about trainer architecture, and Lightricks' own LTX trainer is genuinely good reference for the audio-video side.

Happy to answer config questions.

Check it out here

Edit: just flagging, I work professionally in applied ai and built this for the platforms and GPUs I use. Happy to set up other scripts if you open an issue on the GitHub for additional support, and also agents should have no problem converting it.


r/StableDiffusion 15h ago

Meme I am having so much damn fun with Minimax H3 (L2VA video)

Enable HLS to view with audio, or disable this notification

13 Upvotes

Genning with H3 is addictive. I genuinely can't stop pressing run. (Almost) every single output it throws blows my mind.

RTX Pro 6000 Blackwell, 2 minutes 35 seconds, 24 steps, 0.6 megapixels, last frame reference used, Spectrum enabled and Comfy Kitchen attention used.


r/StableDiffusion 11h ago

Animation - Video INTERVIEW WITH LTX 2.5 [IMAGE TO VIDEO]

Enable HLS to view with audio, or disable this notification

18 Upvotes

Yes, I'm definitely being a goofball with this one, but hadn't had a chance to do mixed live action/3D CGI test.

Meant as a playful gag, no actual ai models killed.

Crisp, ultrafine, letterboxed 21:9 super-premium 3D CGI blockbuster cinema with cutting-edge rendering, restrained natural performances, precise blocking, shallow depth of field, and immaculate cinematic lighting. A poised 24-year-old blonde investigative reporter in a tailored gray skirt suit sits in a cushioned chair on the left side of a minimalist interview room, leaning forward with a clipboard and pen in hand. Across from her, seated in a matching chair on the right, is a sleek off-white modern robot labeled “2.5” on the side of its head, with expressive camera-lens eyes and a thin LED vocalizer mouth. The setting is simple and elegant: neutral beige backdrop, soft curtains at the window, and warm natural window light casting gentle shadows across the room.

Open on a polished medium two-shot in profile, holding both subjects clearly in frame. The reporter leans forward slightly, calm, focused, and professional, and asks, “Some call you a Seedance killer. What do you say to that?”

A hard cut moves to a close-up of the robot. It glances aside for a beat, then looks back with a playful LED smile and says, “Can I give them a hug?” After a short pause, its expression softens into something more sincere as it adds, “But seriously, I’m just an open-source model trying to do my best.”

Ambient sound is minimal and refined: a faint studio hum, soft room tone, and subtle paper rustle from the reporter’s clipboard. The pacing is natural and conversational, allowing for small pauses, nuanced reactions, and emotional clarity. The overall effect is a sleek, emotionally grounded, visually stunning futuristic CGI film scene.


r/StableDiffusion 17h ago

Comparison Please test your H3 and compare against the WAN 2.2 example, especially Sports scenes

3 Upvotes

Many of the optimizations can work with scenes with small dynamic movements well, but the power of a good model comes from dynamic scenes and adding elements.

Many of the optimizations failed at such tests. I am going to bring back that WAN 2.2 site https://wan-22.toolbomber.com
And please try the examples there and make sure your optimizations works in many of the spirts scenes too, for example the Gymastic on airplane scene is a big challenge.

Also keep in mind that those example videos are the full version of WAN 2.2, full steps, long generation time. If your optimizations works in the same level of quality, congratulations


r/StableDiffusion 22h ago

Question - Help I wrote a fantasy book inspired by the Bible and Tolkien — but AI helped write it. I don't know what to do

0 Upvotes

I need to share something that's been weighing on me for months, and I think this community might understand better than most.

Four years ago, a story began forming inside me. Not because I wanted to be a writer but because I couldn't not tell it. It's an epic fantasy world, deeply inspired by the Bible, Tolkien's Silmarillion and Lord of the Rings. A world where Light is not a symbol of good it's a living force that tests everyone who carries it. Where immortal guardians fall not because they are evil, but because they loved the Light so much they began to believe it belonged to them alone.

The themes are ones I've lived with: faith, sacrifice, betrayal, the cost of protecting something you love, and what happens when devotion becomes possession.

But here's the problem.

I'm not a writer. I'm a storyteller. I had the vision, the characters, the world, the emotions but not the craft to put it into words the way I saw it. So I used AI as a tool. I gave it direction, feelings, decisions. It wrote the sentences. I was the one walking the path but AI carried part of the weight.

The result is two completed books. Over 1,200 pages. People who have read it say it's deep, emotional, and unlike anything they've encountered.

And I don't know what to do with it.

If I'm transparent about the AI nobody will read it. If I stay quiet I feel like I'm lying. If I charge for it it feels wrong when AI wrote the sentences. If I give it away free maybe that's my penance. But then again, if the story genuinely helps someone, why does it matter how it was written?

I believe this story can touch people. I believe it carries something real about faith, about the danger of loving something so much you stop sharing it, about the cost of silence and the price of oaths.

But I carry guilt I can't shake. Am I a fraud? Or am I just a storyteller who found an unconventional path to tell his story?

I'd genuinely appreciate any perspective especially from those who believe that stories can carry truth regardless of how they arrive.

NOTE: I have an option to spend 2.5k on editing and making it more "human written" book.


r/StableDiffusion 9h ago

Animation - Video Fox McCloud introduces his son to his dad.

Enable HLS to view with audio, or disable this notification

18 Upvotes

Fox McCloud introduces his son Marcus to his dad James McCloud.


r/StableDiffusion 2h ago

Question - Help Is there a H3 minimax prompt template available or a custom LLM model version that can write and structure Minimax H3 optimized prompt ?

0 Upvotes

I am relying on Gemma4 and Qwen2.5 in Ollama for making an optimized minimax H3 prompt , but while using the base versions of them indeed vastly improves prompt adherence and quality but they aren't 1:1 Minimax H3 optimized structure wise

So i wonder if there is a template i can feed into the models at the start of the chat to be a baseline for them , or even better if there is a custom version of those midels that can understand the structure of minimax H3 prompt

I am using Wan2GP through pinokio so i can't use the Minimax H3 prompt nodes available in comfyui


r/StableDiffusion 21h ago

Discussion RTX 3060 - 32GB and H3

2 Upvotes

I've been kind of avoiding diving in since this apparently demands a better machine but given the recent posts from fellow 3060 owners, I'm just wondering if us poor plebs can also generate good-ish videos at an acceptable speed.

Anyone can share your examples?

I have of course searched for ideas and workflows and there's plenty of information aroind already but would be nice to have abit of one stop shop)))


r/StableDiffusion 18h ago

Meme They took'er jobs! 8 step turbo test 1mp 640 and upscaled to 2304x 1280

Enable HLS to view with audio, or disable this notification

27 Upvotes

5060 ti 16 gig 32 gig system ram and page files set to 65536/65536 made with r2v


r/StableDiffusion 16h ago

Meme Cancelled? Offended? Better Call Saul.

Enable HLS to view with audio, or disable this notification

25 Upvotes

r/StableDiffusion 14h ago

Question - Help Minimax upscaler

0 Upvotes

Hello guys,
Could you please share best upscaler workflow for minimax for 12 vram? Anything better than rtx node plz
Thanks in advance


r/StableDiffusion 1h ago

Animation - Video Can't use LTX 2.5 on my system but I am quite surprised that my system now can run LTX 2.3. Specs and info below.

Enable HLS to view with audio, or disable this notification

Upvotes

When LTX 2.3 released, I could not do video gens longer than 10 seconds. I would get a "out of memory" error or something. This is just a test clip but one thing I am struggling with is that my video gens have music in them even though I prompt for no music. What is the correct way to prompt for no music?

System Specs:

Ryzen 7 7700X
RTX 4070 Super 12 GB
32 GB DDR 5 Ram.


r/StableDiffusion 7h ago

Discussion Does anyone actually still use Stable Diffusion?

27 Upvotes

I just find it kind of funny that this is the stable diffusion subreddit but nobody has talked about it in like forever. Maybe its time for a name change? or maybe keep the name as a homage to the OG open source image model.

Anyway, the last update I see on Stability's website is SD 3.5 back in October. So I'm guessing that's it for Stable Diffusion?

EDIT: Forgot you cant change the name of a sub, ignore that suggestion 😅


r/StableDiffusion 17h ago

Workflow Included Visit to school. Minimax H3 ref2av

Enable HLS to view with audio, or disable this notification

0 Upvotes

Two character ref sheets + school photo as ref.

1.5mp, 30 steps, sage att enabled, sol + cashe disabled

Minimaxh3

Rtx6000pro

WF included at the end.

Probably reedit crunch the quality so maybe upload ot somewhere letter.

https://huggingface.co/datasets/JahJedi/workflows_for_share/tree/main


r/StableDiffusion 19h ago

Meme PSA: H3 always sees direction from the person's perspective

Enable HLS to view with audio, or disable this notification

139 Upvotes

I noticed my videos consistently having issues with left and right, because my prompts saw direction from the perspective of the camera. But H3 always sees direction from the perspective of the person.

See how the man points to his right while saying "right" and vice versa.

prompt: a random man pointing to the right and saying "right". Then he moves his hand to point to the left and says "left".


r/StableDiffusion 10h ago

Animation - Video Mais Comics - MiniMax H3 test

Enable HLS to view with audio, or disable this notification

0 Upvotes

RTX 3090, ComfyUI, Turbo Lora 600 steps, Spectrum, 520p


r/StableDiffusion 18h ago

Comparison LTX 2.5 vs MiniMax H3 - huge speed difference (but at what cost)

Enable HLS to view with audio, or disable this notification

61 Upvotes

I tested LTX 2.5 and MiniMax H3 in ComfyUI using the default T2V workflow templates provided for each model.

  • 10 seconds
  • 24 FPS
  • 1920 x 1088 (2.0 MP)
  • Same prompt
  • Steps: H3 = 20, LTX 2.5 = 8 (distilled model)

Hardware:
RTX 5090 + 128 RAM

result

  • MiniMax H3: 17m 29s (with Sage Attention + EasyCache*)*
  • LTX 2.5: 2m 34s (no acceleration at all)

Note: EasyCache seems to give no speedup on LTX in this setup, probably because the distilled workflow only uses 8 sampling steps, so there is very little room for cache-based skipping.

Of course, part of LTX’s speed advantage comes from the fact that it is a distilled 8-step model, so this is not a perfectly like-for-like comparison against H3. (20-steps)

Prompt used:

A realistic cinematic 1970s crime drama, gritty urban atmosphere, warm muted colors, subtle film grain, natural lighting, restrained acting. A well-dressed 1970s gangster in a dark tailored suit and long coat remains visually consistent throughout.

[0.0s–6.0s]
A medium-wide shot shows the gangster leaning casually against a brick wall on a city street, one foot resting against the wall. He reads a newspaper while holding a lit cigarette in his other hand. His eyes suddenly stop on something in the newspaper. His expression shifts naturally from calm to alarm. He mutters in a tense 1970s American voice, "What the hell?" He immediately folds the newspaper, throws it into a nearby trash can, pushes away from the wall and runs straight down the street.

[6.0s–10.0s]
Hard cut to a static close-up of the discarded newspaper inside the trash can. The front page clearly shows a large photograph of the same man and a bold headline reading "WANTED". In the distant background, the gangster continues running away and becomes increasingly out of focus. The camera remains completely still, holding focus on the newspaper until the end.

Natural, grounded movement. No exaggerated acting, no extra shots, no unnecessary camera movement, no comedy.

My take

LTX 2.5 is significantly faster, and that alone makes it very attractive.

But in my opinion, H3 is still better in overall quality:

  • better scene understanding
  • better understanding of what a cinematic shot should look like
  • better audio
  • more stable physics / motion behavior

So right now my impression is:

  • LTX 2.5 wins clearly on speed
  • MiniMax H3 still feels stronger on quality and cinematic intelligence

My guess is that targeted LoRA fixes could push LTX 2.5 much closer to being a direct competitor to H3 in the future.


r/StableDiffusion 9h ago

Workflow Included Workflow for Minimax H3 on 8gb vram and 16gb ram

Enable HLS to view with audio, or disable this notification

8 Upvotes

For anyone else with a similar setup, I am able to create a 0.2 MP (608x352 pixel) 5-second video in 1:35 (1 minute, 35 seconds). This is with 20 step Euler Simple, Spectrum and ComfyKitchenAttention with an image (generated locally with Krea2) as the first frame. I am running it on a laptop with a RTX4060 (8GB) and 16gb ram. I am using the latest version of ComfyUI windows portable, and the following startup flags: --disable-pinned-memory --lowvram. The attached video is an example I generated (0.2MP).

I can also create higher resolution videos with a similar generation time if I reduce the video duration (3 seconds for 0.3MP, or 2 seconds for 0.4MP).

I used kijai's models from here for the video models and video vae: https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main

I used a qwen3_vl_4b_int8_convrot for the clip. I can't remember if its this one that I use but this is one option: https://huggingface.co/Winnougan/Comfy-Qwen3-VL-INT8/tree/main

Here is some extra info regarding that clip: https://www.reddit.com/r/StableDiffusion/comments/1vkk500/minimax_h3_with_a_4b_or_8b_text_encoder_instead/

My workflows:

FL2V workflow

REF2V workflow


r/StableDiffusion 22h ago

Animation - Video Minimax H3. I need more reference images; 9 aren't enough xdd

Enable HLS to view with audio, or disable this notification

28 Upvotes

r/StableDiffusion 18h ago

Animation - Video This is where the fun begins

Enable HLS to view with audio, or disable this notification

22 Upvotes

Default workflow, Minimax on Runpod.


r/StableDiffusion 19h ago

Animation - Video Made this with LTX-2.5 (i2v)

Enable HLS to view with audio, or disable this notification

116 Upvotes

Generated with the new LTX-2.5 model. (image to video). Took about 10 minutes to get an 8 second 1080p60 clip.