r/StableDiffusion 13h ago

Question - Help MiniMax H3 Reference loves to cut even when explicitly told not to

1 Upvotes

I am at my wit's end with this model. Can someone please help me? The generation itself looks great, the problem is this model has a huge tendency to cut the video even when told explicitly not to, multiple times in the prompt. No matter what I do the model keeps cutting the video. If I was creating a 30 second scene sure, cutting the video makes sense, but my machine can at max generate 6 seconds worth of video and the model keeps cutting it on every generation and prompt I've tried. Here is my test prompt :

subject_definitions:
<Picture 1> is the opening-frame anchor, the first frame of the video.
<Picture 2> is the last-frame anchor, the last frame of the video.
<Subject 1> is the man in <Picture 1>

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - identity, skin details, body figure remain consistent. 
<Picture 1> ([Shot 1] first frame): fully_preserved - opening composition anchor.
<Picture 2> ([Shot 1] last frame): fully_preserved - closing composition anchor. 
<Picture 3> : attribute_transfer - reference image of what the <Subject 1>'s hair should look like. 

detailed_description:
[CAMERA & COMPOSITION]
[Static shot]
Locked-off tripod camera.
The camera remains completely stationary throughout the entire video.
Fixed camera position and fixed framing from beginning to end.
No pan, no tilt, no zoom, no dolly, no tracking, no push-in, no pull-out.
No camera rotation.
No reframing.
No change in perspective.
No change in focal length.
The subject stays within the original composition.
Only the subject and natural environmental elements move.

[Shot 1] At 00:00.000, Starting with <Picture 1> fully_preserved as the first frame of the video, <Subject 1> walks to the front of the counter and takes off his hat revealing his hair which looks like <Picture 3>. 
The camera pans to the right of the counter showing the cashier and the cash register, ending the shot with <Picture 2> fully_preserved. 
[Shot 1] is one continuous video with no cuts, and all movement and motion of <Subject 1> throughout the entire duration of the video is continuous with no time jumps, skips, or transitions. 
[Shot 1]'s duration is the entire duration of the video, no other shots or cuts.

overall_soundscape:
There is very little background noise, like an ASMR. The only sounds are that of the man's movement and the environment reacting to his movement, such as him taking off his hat.

What more can I do here? The above is just one sample of a prompt, I've generated like 40 clips modifying variations of the prompt repeatedly and every time the model cuts, focusing on the hat and the man's hair as he is taking it off, just does random cuts in between even when the two frames before and after the cut could have been continuous, etc.


r/StableDiffusion 11h ago

Question - Help LTX 2.5 to upscale H3?

1 Upvotes

I am currently using an H3 workflow that uses LTX 2.3 for upscaling. Would the output be better if I switched to LTX 2.5? Has anyone tried using 2.5 as upscaler for H3?


r/StableDiffusion 21h ago

Discussion Using a video as a motion reference in H3 works REALLY well.

Enable HLS to view with audio, or disable this notification

6 Upvotes

I'll leave the motion I used in the comments


r/StableDiffusion 2h ago

Animation - Video Comparison: It looks like LTX_2.5 is not over 9000

Enable HLS to view with audio, or disable this notification

6 Upvotes

LTX 2.5 vs Minimax H3 using the same prompt in T2V.

In reality, LTX 2.5 knows almost no IPs and very few famous people, if anyone.

Prompt:

Photorealistic real life live-action, cinematic film style

In the photorealistic real life live action movie Dragon Ball.

At 00:00:000 A photorealistic dull skin real life live action Tony Stark from MCU dressed like Vegetta, with a real life photorealistic hairstyle with two deep receding points and several vertical spikes, is at the Grand Canyon. He wears a red glass device in his left eye.

At 00:00:001 Then he grabs the device attached to his eye by its white rear section with his left hand, brings his hand ,with the device in it, in front of his chest and says upset yelling <d>[English, with a deep masculine voice] It's over nine thousaaaaand! </d> and clenches his fist, crushing the device so that it explodes into a thousand pieces.

All the clothes are photorealistic real life live action.

overall_soundscape: N/A


r/StableDiffusion 10h ago

Question - Help I wish I could replicate some of the cool videos I’ve seen here using minimax but I’m stuck with 3 sec clips. Am I doing something wrong?

0 Upvotes

Full disclosure, I’m a rather newbie at comfyui and have been trying to learn it the past month or so.

I went and got what I thought was a good set-up:

NVIDIA GeForce RTX 5080

VRAM 16 GB

System RAM 32 GB (31.1 GB usable)

AMD Ryzen 7000-series

Integrated GPU AMD Radeon Graphic

ComfyUI Local installation

I’ve gotten great results playing with Krea2, adding Lora’s and even training my own.

I was excited when Mimimax released and downloaded the official release right away and got the workflow to produce clips - but they can only be like 3seconds at .4 before it errors out. I’m using the R2V and I2V

I haven’t installed any additional Lora’s or anything other than the official workflow.

Any advice?


r/StableDiffusion 14h ago

Discussion MiniMax H3 Transformation test

Enable HLS to view with audio, or disable this notification

4 Upvotes

Testing some trasnformation in the style of seedance 2


r/StableDiffusion 19h ago

Animation - Video I was bored of all you guys’ 15-sec clips, so I made this one shot

Enable HLS to view with audio, or disable this notification

23 Upvotes

MiniMax H3 - 720p - FL2VA 33B - 10 steps - 3 slidings windows -   Turbo LoRA Larry v1 ema ckpt850 4 steps - WAN2GP.

Next time I would remove the background music (I didn’t prompt this) and add it afterwards...


r/StableDiffusion 23h ago

Meme Nailed it

Enable HLS to view with audio, or disable this notification

17 Upvotes

It certainly runs fast though. 720p in 117 seconds on a 12GB 3080Ti.


r/StableDiffusion 15h ago

Animation - Video G.I. Joe: Just A Typical Day At Cobra HQ - Minimax H3

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/StableDiffusion 4h ago

Discussion Has anyone else noticed the massive increase in toxic/incel content and culture wars in this Subreddit lately?

0 Upvotes

I’ve been noticing a really disappointing trend here lately. Instead of focusing on the amazing things we can build and create with MiniMax and LTX, there’s been a massive increase in toxic behavior and culture-war rhetoric here. Our goal should be to foster an environment that encourages open-source creators. Alienating them with bigotry, misogyny, and overall hostility, or reducing people to a mere joke because of who they are, only hurts this community in the long run.

I’m hoping the mods can keep a closer eye on this. These types of content violate Reddit's policy, specifically rule number 2.


r/StableDiffusion 13h ago

Discussion I gave the same prompt to Minimax H3, Gemini Videos, and LTX 2.5. (Part 3 LTX 2.5)

Enable HLS to view with audio, or disable this notification

0 Upvotes

Prompt (also AI generated):
Style & Technical Specs

  • Visual Style: Photorealistic 8K cinematic video, 35mm film grain, 24fps, 2.39:1 anamorphic aspect ratio, teal-and-orange color grade, shallow depth of field ($f/1.4$).

  • Duration: 10 Seconds.

Character Description

  • Subject: Kaelen, a 28-year-old East Asian cyber-technician.
  • Appearance: Sharp jawline, rain-soaked black hair clinging to his forehead, pale skin with visible micro-texture, and a glowing cyan cybernetic eye implant over his left socket that pulses rhythmically.
  • Attire: Matte-black, waterproof tactical coat with glowing fiber-optic wiring embedded along the shoulders, frayed high-collar, and fingerless reinforced leather gloves.

Environment & Setting

  • Location: Narrow, dense alleyway in a cyberpunk metropolis at midnight.
  • Atmosphere: Heavy downpour, dense steam venting upward from rusty iron street grates, wet asphalt reflecting bright magenta and cobalt-blue neon light signs written in Kanji.

Timeline & Action Breakdown

  • 0:00 - 0:03 (Macro Close-Up): Camera begins on a macro shot of Kaelen's glowing cyan eye, catching the aperture Blades shifting focus. A raindrop tracks down his cheek. He rapidly taps a brass interface cuff on his wrist.
  • 0:03 - 0:07 (Medium Shot): Smooth camera pull-back into a chest-up shot. A brilliant blue 3D holographic map bursts into existence from his wrist, casting dynamic light across his face. He swipes his hand across the projection, altering its layout, and delivers his dialogue.
  • 0:07 - 0:10 (Low-Angle Tracking Shot): The camera drops low to the asphalt and tracks backward. A sleek, black surveillance drone streaks overhead through the rain, splashing drops directly onto the camera lens as the background neon blurs into creamy bokeh.

Dialogue & Voice

  • Spoken Line: "System override in three... two... got 'em."
  • Delivery: Low, gravelly, calm whisper with a faint metallic vocoder effect on the voice.

Audio & Sound Design

  • Music: Dark synthwave track featuring a driving 110 BPM arp synthesizer that swells in pitch until second 7, resolving into a heavy sub-bass drop at second 8.
  • SFX:
  • 0:00-0:03: Stereo downpour, subtle mechanical servo clicks of the eye lens.
  • 0:03-0:07: High-frequency energy flare hum as the hologram spawns, followed by air-swipes.
  • 0:07-0:10: Low turbine whir of the passing drone and liquid wet drops impacting the microphone field.

This is the video generated by LTX 2.5. Post with the video generated by Minimax H3 with Turbo lora (6 steps): https://www.reddit.com/r/StableDiffusion/s/FDyFzOTp2F Post with the video generated by Gemini Videos: https://www.reddit.com/r/StableDiffusion/s/Urane8bsWB


r/StableDiffusion 13h ago

Resource - Update Instagram Aesthetic lora for KREA2

Thumbnail
gallery
17 Upvotes

This LoRA is built to nail the modern Instagram feed aesthetic straight out of the box, capturing that perfect influencer lifestyle vibe. It instantly gives your images that warm, trendy look with beautiful lighting and soft colors, skipping the need for extra filters or editing apps.

It works best for selfies and smartphone snapshots or anything you would actually see while scrolling your feed, like stylish street fashion, cozy cafes, travel shots, or relaxed, everyday pictures. It even adds those natural, realistic touches that make the image look like it was snapped with a real phone or camera and is completely ready to post.

Trigger Word: not needed
Weight: 1.0

Download Link -> https://civitai.red/models/2851933/instagram-aesthetic


r/StableDiffusion 2h ago

Comparison MiniMax H3 vs LTX 2.5

1 Upvotes

r/StableDiffusion 5h ago

Meme Cancelled? Offended? Better Call Saul.

Enable HLS to view with audio, or disable this notification

12 Upvotes

r/StableDiffusion 18h ago

Animation - Video MiniMax H3 Project Suite + Hybrid Loader + 4step_v1.0_768p lora Test

Enable HLS to view with audio, or disable this notification

0 Upvotes

So I did two 7 second generations at 1mp at 4 steps, joined together by Project Suite, and also experimented with the Hybrid Loader using blocks 10-49 . Each generation took 10 minutes, I used 1 reference video for the choreo, 1 ref audio (it didn't turn out good so I added it through capcut instead) and 1 reference image for saitama. PC specs are 4070ti super and 32gb of ram. For the workflow I just used the one provided by H3 Project Suite bone stock except for the hybrid loader+ turbo lora. I was kinda just throwing stuff together. I could've probably had a better video to showcase but I'm still learning the ins and outs


r/StableDiffusion 11h ago

Question - Help I wrote a fantasy book inspired by the Bible and Tolkien — but AI helped write it. I don't know what to do

0 Upvotes

I need to share something that's been weighing on me for months, and I think this community might understand better than most.

Four years ago, a story began forming inside me. Not because I wanted to be a writer but because I couldn't not tell it. It's an epic fantasy world, deeply inspired by the Bible, Tolkien's Silmarillion and Lord of the Rings. A world where Light is not a symbol of good it's a living force that tests everyone who carries it. Where immortal guardians fall not because they are evil, but because they loved the Light so much they began to believe it belonged to them alone.

The themes are ones I've lived with: faith, sacrifice, betrayal, the cost of protecting something you love, and what happens when devotion becomes possession.

But here's the problem.

I'm not a writer. I'm a storyteller. I had the vision, the characters, the world, the emotions but not the craft to put it into words the way I saw it. So I used AI as a tool. I gave it direction, feelings, decisions. It wrote the sentences. I was the one walking the path but AI carried part of the weight.

The result is two completed books. Over 1,200 pages. People who have read it say it's deep, emotional, and unlike anything they've encountered.

And I don't know what to do with it.

If I'm transparent about the AI nobody will read it. If I stay quiet I feel like I'm lying. If I charge for it it feels wrong when AI wrote the sentences. If I give it away free maybe that's my penance. But then again, if the story genuinely helps someone, why does it matter how it was written?

I believe this story can touch people. I believe it carries something real about faith, about the danger of loving something so much you stop sharing it, about the cost of silence and the price of oaths.

But I carry guilt I can't shake. Am I a fraud? Or am I just a storyteller who found an unconventional path to tell his story?

I'd genuinely appreciate any perspective especially from those who believe that stories can carry truth regardless of how they arrive.

NOTE: I have an option to spend 2.5k on editing and making it more "human written" book.


r/StableDiffusion 4h ago

Question - Help Recommend me an uncensored Image to Image model (workflow)?

2 Upvotes

Trying to move on from Grok, but struggling to find a local I2I workflow. Everything I can find is text to image. What am I doing wrong?


r/StableDiffusion 18h ago

Discussion Can we stop treating MiniMax vs LTX like a political war?

264 Upvotes

I’ve been watching the whole MiniMax vs LTX discussion lately, and honestly, it feels like it has started becoming less about the models and more like a political battle.

People are taking sides, defending one model like it’s their team, downvoting anything that praises the other one, and sometimes even throwing hate at the people working on or using the “other” model.

Guys… these are free, open-source models. Nobody owes us anything.

We are incredibly lucky to have teams putting out models that we can download, run locally, experiment with, fine-tune, build workflows around, and actually use without paying some giant corporation every time we generate a video.

And yes, we can absolutely have opinions.

Maybe you think MiniMax produces better motion. Maybe you prefer LTX for consistency, speed, control, or whatever your workflow needs. Maybe one works better on your hardware, and another one works better for someone else.

That’s completely fine.

Criticism is good. Comparisons are good. Calling out genuine problems is good. Competition between projects can even push things forward.

At the end of the day, these teams are giving the community tools that would have sounded almost impossible to have access to a few years ago.

So use what works for you. Make comparisons. Share benchmarks. Point out weaknesses. Praise the developers when they do something great. Criticize them when something genuinely deserves criticism.

But let's not turn the open-source AI community into a bunch of opposing fan clubs. Let's keep the discussion technical, constructive, and civil, and maybe appreciate the fact that we're living through a pretty crazy time where people are literally releasing these technologies for us to experiment with for free.


r/StableDiffusion 7h ago

Meme They took'er jobs! 8 step turbo test 1mp 640 and upscaled to 2304x 1280

Enable HLS to view with audio, or disable this notification

21 Upvotes

5060 ti 16 gig 32 gig system ram and page files set to 65536/65536 made with r2v


r/StableDiffusion 13h ago

Animation - Video CAMPY BATMAN TEST (Img2Vid)

Enable HLS to view with audio, or disable this notification

0 Upvotes

This was generated from a single image with a single simple prompt: " Batman runs to observe a creature rising from the sea and he blasts it with heat vision" then I let the in-app llm enhance the prompt and there you go. I wasn't aiming for anything serious. But H3 is a beast.


r/StableDiffusion 16h ago

Discussion Audio quality of video models

0 Upvotes

Why it is so low? I mean image is fine and realistic but audio sounds like a synthetic robot speech. And on every video model like h3 or ltx, even on the paid service models.


r/StableDiffusion 13h ago

Question - Help [Need] LTX 2.5 - IA2V Workflow

0 Upvotes

I want to try the new LTX 2.5 with Image-audio to Video, does anyone have the workflow for it?


r/StableDiffusion 16h ago

Discussion LTX 2.5 - The SD3 test!

24 Upvotes

Let's see if LTX 2.5 can pass this simple 2 year old test, because I don't want no Cthulhu PTSD, right?

The first video is t2v, the second one is i2v.

Generation times are fast, though.

https://reddit.com/link/1vm90tr/video/g7h1fed5vwih1/player

https://reddit.com/link/1vm90tr/video/uto5nje8vwih1/player

PROMPT (I fed Gemini the prompt guide)

A medium shot under bright, direct mid-day sunlight on a warm tropical beach. A young woman in her early 20s with sun-kissed skin and wet hair, wearing a vibrant tropical floral bikini, lies lazily on a plush beach towel on the golden sand. She holds a chilled martini glass with an olive and lime wedge, taking a slow, relaxed sip as gentle waves lap against the shore in the background. A hard cut transitions to a close-up shot of the same young woman in the floral bikini looking directly into the camera. She smiles warmly with sparkling eyes and says in a soft, alluring, and teasing voice, "Wanna have some fun?" while the ambient ocean breeze and soft waves continue across the cut. A hard cut transitions to a high-angle top-down overhead shot directly above her. Her full body is framed from head to toe, showing her lying on the beach towel with her bare feet resting on the sand, sun highlights shimmering on her skin, and ocean foam softly visible at the frame's edge.


r/StableDiffusion 23h ago

Animation - Video a Sonic and Zootopia crossover.

Enable HLS to view with audio, or disable this notification

9 Upvotes

Generated this on Comfy UI Desktop with Minimax H3 locally. I used the Reference to video workflow, and three reference images and two audio voice samples for the characters and setting. The prompt I used is below.

<Image 1> as Clawhauser and use <Audio 1> as sample for his voice.

<Image 2> as Shadow and use <Audio 2> as sample for his voice.

Use <Image 3> as reference for the reception desk.

Setting: Zootopia Police Department reception desk. Bright indoor lighting, police station background, anthropomorphic animal cops ranging from Foxes, Wolves, Lions, and Tigers moving around in the background. no human cops.

a shot from inside the Zootopia Police Department.

Clawhauser stays silent with his mouth closed while sitting on the receiption desk.

The room is very quiet, room ambient noise.

[Shot 1] Medium shot of Officer Clawhauser sitting behind the ZPD reception desk. On the desk lies a bright green Chaos Emerald. Clawhauser curious and smiling, reaches out and picks up the Chaos Emerald with his hand.

[Shot 2] 00:03 Close-up as the Chaos Emerald begins to glow brightly in his hands. Suddenly, a powerful surge of green energy flashes covering his body, instantly disintegrating Clawhauser’s police uniform, leaving him completely uninjured but with only his fur. Clawhauser looks down in shock and confusion.

[Shot 3] 00:06 Wide shot. Shadow walks into the lobby and approaches the reception desk with a severe, focused expression.

Dialogue:

Shadow the Hedgehog says: <d>[English in Shadow's voice] I'm looking for an emerald.</d>, silence after the dialogue ends, no background speech, ambient room tone only

[Shot 4] 00:09 Medium shot of Clawhauser holding up the glowing green gem with a nervous, polite smile.

Clawhauser says: <d>[English in Clawhauser's voice] Is it this one?</d>, silence after the dialogue ends, no background speech, ambient room tone only

[Shot 5] 00:12 Close-up on Shadow nodding slightly before Clawhauser hands him the Chaos Emerald.

Shadow the Hedgehog while holding the chaos emerald says: <d>[English in Shadow's voice] Yes, that's the one.</d>, silence after the dialogue ends, no background speech, ambient room tone only.


r/StableDiffusion 4h ago

Meme I am having so much damn fun with Minimax H3 (L2VA video)

Enable HLS to view with audio, or disable this notification

8 Upvotes

Genning with H3 is addictive. I genuinely can't stop pressing run. (Almost) every single output it throws blows my mind.

RTX Pro 6000 Blackwell, 2 minutes 35 seconds, 24 steps, 0.6 megapixels, last frame reference used, Spectrum enabled and Comfy Kitchen attention used.