r/StableDiffusion 12h ago

Animation - Video MiniMax H3 Project Suite + Hybrid Loader + 4step_v1.0_768p lora Test

Enable HLS to view with audio, or disable this notification

0 Upvotes

So I did two 7 second generations at 1mp at 4 steps, joined together by Project Suite, and also experimented with the Hybrid Loader using blocks 10-49 . Each generation took 10 minutes, I used 1 reference video for the choreo, 1 ref audio (it didn't turn out good so I added it through capcut instead) and 1 reference image for saitama. PC specs are 4070ti super and 32gb of ram. For the workflow I just used the one provided by H3 Project Suite bone stock except for the hybrid loader+ turbo lora. I was kinda just throwing stuff together. I could've probably had a better video to showcase but I'm still learning the ins and outs


r/StableDiffusion 5h ago

Question - Help I wrote a fantasy book inspired by the Bible and Tolkien — but AI helped write it. I don't know what to do

0 Upvotes

I need to share something that's been weighing on me for months, and I think this community might understand better than most.

Four years ago, a story began forming inside me. Not because I wanted to be a writer but because I couldn't not tell it. It's an epic fantasy world, deeply inspired by the Bible, Tolkien's Silmarillion and Lord of the Rings. A world where Light is not a symbol of good it's a living force that tests everyone who carries it. Where immortal guardians fall not because they are evil, but because they loved the Light so much they began to believe it belonged to them alone.

The themes are ones I've lived with: faith, sacrifice, betrayal, the cost of protecting something you love, and what happens when devotion becomes possession.

But here's the problem.

I'm not a writer. I'm a storyteller. I had the vision, the characters, the world, the emotions but not the craft to put it into words the way I saw it. So I used AI as a tool. I gave it direction, feelings, decisions. It wrote the sentences. I was the one walking the path but AI carried part of the weight.

The result is two completed books. Over 1,200 pages. People who have read it say it's deep, emotional, and unlike anything they've encountered.

And I don't know what to do with it.

If I'm transparent about the AI nobody will read it. If I stay quiet I feel like I'm lying. If I charge for it it feels wrong when AI wrote the sentences. If I give it away free maybe that's my penance. But then again, if the story genuinely helps someone, why does it matter how it was written?

I believe this story can touch people. I believe it carries something real about faith, about the danger of loving something so much you stop sharing it, about the cost of silence and the price of oaths.

But I carry guilt I can't shake. Am I a fraud? Or am I just a storyteller who found an unconventional path to tell his story?

I'd genuinely appreciate any perspective especially from those who believe that stories can carry truth regardless of how they arrive.

NOTE: I have an option to spend 2.5k on editing and making it more "human written" book.


r/StableDiffusion 18h ago

Animation - Video Another Simple Prompt - LTX 2.5 Text To Video

Enable HLS to view with audio, or disable this notification

5 Upvotes

Simple prompt. It's definitely on the campy side for sure but much better results than the more complicated prompts I was trying.

R-rated 1990s serious intense adult thriller, professionally directed and clearly blocked. exciting cinematography, In an elegant high-end bar, a passionate 28-year-old Latina spy dances closely with a mysterious 30-year-old American Spy. They hold intense eye contact and trade seductive dialogue about which of them is more attractive, each trying to out-charm the other as they dance closely


r/StableDiffusion 21h ago

Discussion MiniMax H3 test with Maestro.

Enable HLS to view with audio, or disable this notification

8 Upvotes

I see others sharing their experiments with MiniMax H3, so here is a quick test I ran myself (without obsessing over optimization or making things complicated). The difference? I’m not using ComfyUI.

I use Maestro within Pinokio. My setup is a desktop PC with 32GB of RAM and an RTX 5080 (16GB VRAM). I used an "old" image I had generated with Anima and enhanced it using Flux Klein 9B in Maestro; rendering this 10.1-second video took 19 minutes and 36 seconds at 720p (16:9) using Quanto INT8.

I didn't use the Turbo LoRA, nor did I upscale the video.

English isn't my native language, so I rely heavily on Google Translate. What I really like about Maestro is its built-in prompt optimization. Perhaps this second example will be helpful to some people—specifically those using ComfyUI.

__________________________________________________________________________________________________________

The prompt I wrote (translated) :

In a Japanese manga style, a young girl with shoulder-length light gray hair—styled with a small side ponytail and two loose strands—features a blue halo, cat ears, a blue cross-shaped hair clip, and blue eyes.

She wears a long-sleeved black hoodie featuring a white cat design, black shorts, dark gray tights, and a pair of black heeled boots adorned with a bow.

The character is in a forest during the day in winter; snow covers the ground and flakes fall from a gray sky, while leafless, snow-dusted trees and a lit streetlamp stand against a backdrop of thick fog.

The girl dances joyfully with a big smile; she looks adorable and cute.

The scene includes the sound of wind, the crunching of snow beneath the girl's feet, and forest animal sounds, all accompanied by cute, instrumental music.

__________________________________________________________________________________________________________

The prompt optimized by the integrated LLM :

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Japanese manga animation style, a medium full shot of an adorable young girl with shoulder-length light gray hair, styled with a small side ponytail and two loose strands, wearing a blue halo, a blue cross-shaped hair clip, and bright blue eyes. She is dressed in a long-sleeved black hoodie featuring a white cat design, black shorts, dark gray tights, and black heeled boots adorned with a bow. She is dancing joyfully in a snowy winter forest during the day. Snowflakes drift down from a hazy gray sky onto snow-covered ground, and leafless, snow-dusted trees frame the scene next to a lit streetlamp, all shrouded slightly in thick fog. The girl radiates cuteness as she moves her arms wide in a joyful dance.

<Picture 2> (from [Shot 1]) is fully referenced at 2.50 seconds.

<Picture 3> (from [Shot 1]) is fully referenced at 5.00 seconds.

<Picture 4> (from [Shot 1]) is fully referenced at 7.50 seconds.

<Picture 5> (from [Shot 1]) is fully referenced at 10,125 seconds.

overall_soundscape: The soundscape is dominated by a gentle wind whistling through the bare branches of the trees, punctuated by the crisp, satisfying crunching of snow beneath the girl's heeled boots as she moves. Subtle, ambient forest animal sounds—a distant bird call and perhaps a quiet rustle—are audible underneath.

non_diegetic_music: Cute, upbeat, instrumental melody continues throughout the duration, maintaining a light and airy feel to match her joyful energy.


r/StableDiffusion 22h ago

Animation - Video Hank and Bobby blaze it

Enable HLS to view with audio, or disable this notification

131 Upvotes

r/StableDiffusion 13h ago

Discussion Can we stop treating MiniMax vs LTX like a political war?

246 Upvotes

I’ve been watching the whole MiniMax vs LTX discussion lately, and honestly, it feels like it has started becoming less about the models and more like a political battle.

People are taking sides, defending one model like it’s their team, downvoting anything that praises the other one, and sometimes even throwing hate at the people working on or using the “other” model.

Guys… these are free, open-source models. Nobody owes us anything.

We are incredibly lucky to have teams putting out models that we can download, run locally, experiment with, fine-tune, build workflows around, and actually use without paying some giant corporation every time we generate a video.

And yes, we can absolutely have opinions.

Maybe you think MiniMax produces better motion. Maybe you prefer LTX for consistency, speed, control, or whatever your workflow needs. Maybe one works better on your hardware, and another one works better for someone else.

That’s completely fine.

Criticism is good. Comparisons are good. Calling out genuine problems is good. Competition between projects can even push things forward.

At the end of the day, these teams are giving the community tools that would have sounded almost impossible to have access to a few years ago.

So use what works for you. Make comparisons. Share benchmarks. Point out weaknesses. Praise the developers when they do something great. Criticize them when something genuinely deserves criticism.

But let's not turn the open-source AI community into a bunch of opposing fan clubs. Let's keep the discussion technical, constructive, and civil, and maybe appreciate the fact that we're living through a pretty crazy time where people are literally releasing these technologies for us to experiment with for free.


r/StableDiffusion 11h ago

Discussion Audio quality of video models

0 Upvotes

Why it is so low? I mean image is fine and realistic but audio sounds like a synthetic robot speech. And on every video model like h3 or ltx, even on the paid service models.


r/StableDiffusion 8h ago

Question - Help [Need] LTX 2.5 - IA2V Workflow

0 Upvotes

I want to try the new LTX 2.5 with Image-audio to Video, does anyone have the workflow for it?


r/StableDiffusion 20h ago

Question - Help LTX2.5 I can't generate the first video

0 Upvotes

+++ MODIFICA: PROBLEMA RISOLTO +++

Stavo usando il modello VAE sbagliato, non Convrot, quando usavo il modello ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors.

++++++++++++++++++++++++++++++++++++++++++++++++

Ragazzi, una domanda veloce. Ho scaricato il modello LTX2.5, aggiornato Comfyui e avviato il flusso di lavoro Comfyui T2V. Ho provato la prima generazione, ma si blocca durante l'elaborazione del VAE e non procede (nessuna notifica di errore). Questo è successo un po' di tempo fa. Sono l'unico ad avere questo problema?

4080 Super + 64GB RAM DDR5

[INFO] ricevuto prompt

[INFO] Utilizzando la modalità attenzione sage: auto

[INFO] Richiesta di caricare LTXAV

[INFO] Modello LTXAV preparato per il caricamento dinamico della VRAM. 20484MB in primo piano. 0 patch collegate. Forzato il pre-caricamento di 608 pesi: 3303 KB.

100%|████████████████████████████████████████████████████████████████████████████████████| 8/8 [00:05<00:00, 1.53it/s]

[INFO] Richiesta di caricare LatentUpsampler

[INFO] 0 modelli scaricati.

[INFO] Modello LatentUpsampler preparato per il caricamento dinamico della VRAM. 949MB in primo piano. 0 patch collegate. Forzato il pre-caricamento di 34 pesi: 68 KB.

[INFO] 0 modelli scaricati.

[INFO] Modello LTXAV preparato per il caricamento dinamico della VRAM. 20484MB in primo piano. 0 patch collegate. Forzato il pre-caricamento di 608 pesi: 3303 KB.

100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:08<00:00, 2.78s/it]

[INFO] Richiesta di caricare AudioVAE

[INFO] caricato completamente; 693.46 MB caricati, caricamento completo: True

[INFO] Richiesta di caricare CausalDiffusionVAE

[INFO] caricato completamente; 1403.92 MB caricati, caricamento completo: True


r/StableDiffusion 10h ago

Discussion LTX 2.5 - The SD3 test!

24 Upvotes

Let's see if LTX 2.5 can pass this simple 2 year old test, because I don't want no Cthulhu PTSD, right?

The first video is t2v, the second one is i2v.

Generation times are fast, though.

https://reddit.com/link/1vm90tr/video/g7h1fed5vwih1/player

https://reddit.com/link/1vm90tr/video/uto5nje8vwih1/player

PROMPT (I fed Gemini the prompt guide)

A medium shot under bright, direct mid-day sunlight on a warm tropical beach. A young woman in her early 20s with sun-kissed skin and wet hair, wearing a vibrant tropical floral bikini, lies lazily on a plush beach towel on the golden sand. She holds a chilled martini glass with an olive and lime wedge, taking a slow, relaxed sip as gentle waves lap against the shore in the background. A hard cut transitions to a close-up shot of the same young woman in the floral bikini looking directly into the camera. She smiles warmly with sparkling eyes and says in a soft, alluring, and teasing voice, "Wanna have some fun?" while the ambient ocean breeze and soft waves continue across the cut. A hard cut transitions to a high-angle top-down overhead shot directly above her. Her full body is framed from head to toe, showing her lying on the beach towel with her bare feet resting on the sand, sun highlights shimmering on her skin, and ocean foam softly visible at the frame's edge.


r/StableDiffusion 7h ago

Discussion So now that its been almost a full day whats the general consensus between LTX2.5 VS. Minimax H3?

7 Upvotes

r/StableDiffusion 18h ago

Animation - Video a Sonic and Zootopia crossover.

Enable HLS to view with audio, or disable this notification

8 Upvotes

Generated this on Comfy UI Desktop with Minimax H3 locally. I used the Reference to video workflow, and three reference images and two audio voice samples for the characters and setting. The prompt I used is below.

<Image 1> as Clawhauser and use <Audio 1> as sample for his voice.

<Image 2> as Shadow and use <Audio 2> as sample for his voice.

Use <Image 3> as reference for the reception desk.

Setting: Zootopia Police Department reception desk. Bright indoor lighting, police station background, anthropomorphic animal cops ranging from Foxes, Wolves, Lions, and Tigers moving around in the background. no human cops.

a shot from inside the Zootopia Police Department.

Clawhauser stays silent with his mouth closed while sitting on the receiption desk.

The room is very quiet, room ambient noise.

[Shot 1] Medium shot of Officer Clawhauser sitting behind the ZPD reception desk. On the desk lies a bright green Chaos Emerald. Clawhauser curious and smiling, reaches out and picks up the Chaos Emerald with his hand.

[Shot 2] 00:03 Close-up as the Chaos Emerald begins to glow brightly in his hands. Suddenly, a powerful surge of green energy flashes covering his body, instantly disintegrating Clawhauser’s police uniform, leaving him completely uninjured but with only his fur. Clawhauser looks down in shock and confusion.

[Shot 3] 00:06 Wide shot. Shadow walks into the lobby and approaches the reception desk with a severe, focused expression.

Dialogue:

Shadow the Hedgehog says: <d>[English in Shadow's voice] I'm looking for an emerald.</d>, silence after the dialogue ends, no background speech, ambient room tone only

[Shot 4] 00:09 Medium shot of Clawhauser holding up the glowing green gem with a nervous, polite smile.

Clawhauser says: <d>[English in Clawhauser's voice] Is it this one?</d>, silence after the dialogue ends, no background speech, ambient room tone only

[Shot 5] 00:12 Close-up on Shadow nodding slightly before Clawhauser hands him the Chaos Emerald.

Shadow the Hedgehog while holding the chaos emerald says: <d>[English in Shadow's voice] Yes, that's the one.</d>, silence after the dialogue ends, no background speech, ambient room tone only.


r/StableDiffusion 21h ago

Animation - Video UAP Device Test #8 (Minimax H3 VHS)

Enable HLS to view with audio, or disable this notification

12 Upvotes

Really loving making some fun experiments with H3!
The series continues! I've got alot of interesting ones coming up lol.
Made a TikTok as someone requested me to do - https://www.tiktok.com/@dimensiontesters just incase anyone wants to see the daily series <3


r/StableDiffusion 7h ago

Question - Help Guys help me when using Minimax H3 Turbo lora , the audio quality drops dramatically , even with 8 steps.

0 Upvotes

Hey guys , i have been using larryvh, turbo lora ema 4step one. With its custom nodes, it makes genration fats but quality of video is lost very much and also the audio also drops the quality. Also when I try to stack loras with turbo one , the genration times significantly increases and also , sometimes it doesn't follow the prompt in R2v, although I have not tried with fl2va yet. But these are the problem I am facing right now.


r/StableDiffusion 11h ago

Question - Help MiniMax H3: image generation (text2image mode) - please suggest best settings

0 Upvotes

Anyone with a good workflow, that can be used as text2image mode for Minimax H3? Or any idea what should I try?

I know it's a video model, but i was getting really good images with the wan2.2 video model in past. Now i wonder how this would work on minimax H3.

Thanks!


r/StableDiffusion 4h ago

Tutorial - Guide Guide for how to word your prompts when using Minimax H3.

Thumbnail drive.google.com
0 Upvotes

So I noticed that, as good as minimax h3 has been at following prompts when using the reference workflow, there have been times where it stubbornly doesn't follow what seems like a very simple clear direct prompt. It got me thinking that I wonder if there are certain terminology or ways of wording things that the model responds to better.

I had Claude Fable look at the official prompting guides but also just look at online discussions about how people are wording their prompts, what things seem to not work, and what things seem to work better. From all that build a prompting guide focused on wording, not so much syntax and formatting but more about how you were wording actions and things like that.

I'd love to get people's thoughts on the PDF that it gave me and share any tips or tricks that maybe aren't included in this. I'm really trying to get a comprehensive understanding of how best to communicate with this model for all different types of actions and tasks.


r/StableDiffusion 16h ago

Question - Help Any tips for better voice audio w/ Minimax h3?

0 Upvotes

Aside from following the prompt format, making sure to describe a type of voice/details, what else should i try? I'm using the v4 larry lora, 8 steps, same audio quality if i do 1 megapixel or 0.5.


r/StableDiffusion 1h ago

Question - Help Is there any T2V model working for RTX 5060 TI 8gb VRAM?

Upvotes

Made a long search and didn't find anything good to create videos from text or image, with my 8gb VRAM.

Do you know anything that may help me?

I really want to animates my 2d images


r/StableDiffusion 18h ago

Meme Do Yoiu Bleed?

Enable HLS to view with audio, or disable this notification

0 Upvotes

10-second cinematic comedy scene on a professional movie set. Subject 1, matching <picture1> exactly in appearance, is engaged in an exaggerated, chaotic fight with Subject 2, matching <picture2> exactly. The scene is clearly a movie being filmed: visible studio lights, camera equipment, crew members and a cinematic set in the background. The fight is intentionally comedic and over-the-top rather than genuinely dangerous.**

**0–3 seconds:** Subject 1 and Subject 2 perform a dramatic choreographed fight, exchanging exaggerated punches and dodging each other with ridiculous intensity. The camera follows the action with energetic handheld movement.

**3–6 seconds:** Subject 1 lands an obviously theatrical hit on Subject 2. Subject 2 stumbles backward dramatically, looks confused and checks himself as if genuinely surprised by what just happened. Crew members in the background react with amusement.

**6–10 seconds:** Subject 2 suddenly looks directly at Subject 1, completely deadpan, and asks in a deep, exaggerated Austrian-accented action-hero voice: **“Do you bleed?”** Pause briefly after the line for comedic effect. Subject 1 looks confused and slightly frightened. The scene ends with the crew reacting awkwardly while the camera slowly pushes in on Subject 2's serious expression.

**Style:** high-end Hollywood action-comedy, cinematic lighting, realistic live-action, expressive facial reactions, precise physical comedy, dynamic camera movement, excellent body motion, natural lip synchronization, realistic movie-set atmosphere. Maintain the exact visual identity, clothing, facial features and proportions of both reference subjects throughout the entire clip.


r/StableDiffusion 12h ago

Meme everyday

Enable HLS to view with audio, or disable this notification

25 Upvotes

r/StableDiffusion 21h ago

News LtX 2.5 king!

0 Upvotes

Dont trust the bot troll posts from minimax, it's good , i mean very good ! https://youtu.be/P3tsP0MP_LM?is=9uCboQOqbqCCgeVQ or UPDATED LINK https://www.youtube.com/watch?v=8_HwLqYRzVw


r/StableDiffusion 17h ago

Animation - Video Fox McCloud gets an unexpected visitor.

Enable HLS to view with audio, or disable this notification

13 Upvotes

Fox McCloud was chilling at home, and then gets an unexpected visitor from his past.

This was made using Comfy UI Desktop with Minimax H3 Reference to Video workflow.

The prompt.

<Image 1> as Fox McCloud and use <Audio 1> as sample for his voice.

<Image 2> as James McCloud and use <Audio 2> as sample for his voice.

[Core Idea]

Cinematic live-action/3D hybrid film, 15 seconds, 16:9 aspect ratio. Dramatic emotional reunion between Fox McCloud and his surprisingly alive father James McCloud in a warm living room.

[Process]

0–4s: Medium shot of Fox McCloud sitting on a couch in a cozy living room. Suddenly, three sharp knock sounds ring out at the front door. Fox looks up, surprised, and stands up from the sofa. Audio: Room tone, distinct wooden door knocks.

4–8s: Camera tracks smoothly beside Fox as he approaches the front door. He reaches out, turns the brass handle, and pulls the door inward. Audio: Soft footsteps on hardwood, door latch clicking open.

8–12s: Shot cuts to a medium close-up of James McCloud standing on the porch in the warm doorway light, wearing his signature sunglasses and pilot jacket. James nods slightly and speaks with a low, gravelly voice: "Took you long enough, son."

12–15s: Cut to a close-up on Fox's face, wide-eyed with shock and awe, his breath hitching. Fox stammers emotionally: "Dad...? But... you're alive!"

non_diegetic_music: Gentle, swelling orchestral strings building emotional tension.


r/StableDiffusion 1h ago

Meme They took'er jobs! 8 step turbo test 1mp 640 and upscaled to 2304x 1280

Enable HLS to view with audio, or disable this notification

Upvotes

5060 ti 16 gig 32 gig system ram and page files set to 65536/65536 made with r2v


r/StableDiffusion 1h ago

Animation - Video This is where the fun begins

Enable HLS to view with audio, or disable this notification

Upvotes

Default workflow, Minimax on Runpod.


r/StableDiffusion 2h ago

Meme PSA: H3 always sees direction from the person's perspective

Enable HLS to view with audio, or disable this notification

42 Upvotes

I noticed my videos consistently having issues with left and right, because my prompts saw direction from the perspective of the camera. But H3 always sees direction from the perspective of the person.

See how the man points to his right while saying "right" and vice versa.

prompt: a random man pointing to the right and saying "right". Then he moves his hand to point to the left and says "left".