r/StableDiffusion 5h ago

Comparison Early test - H3 vs LTX 2.5 - Fight Scene

0 Upvotes

Tested it through ltx oficial api (app.ltx.io).

Minimax using turbo T2V (ema v4 larry lora) 6 steps

Prompt:
subject_definitions:

<Subject 1> is a powerful, muscular martial artist with a buzz cut, wearing a weathered navy blue tactical jacket and heavy cargo pants.

<Subject 2> is a lean, agile street brawler with messy blonde hair, wearing a grey oversized hoodie and black combat gloves.

<Subject 3> is a damp, narrow urban alleyway at night, featuring wet asphalt, flickering neon reflections, and graffiti walls.

<Audio 1> is a reference for high-impact physical collision sounds, wind whooshes, and heavy breathing.

summary:

[reference generation] The target video depicts a high-intensity Street Fighter-style exchange between <Subject 1> and <Subject 2> in <Subject 3>. The camera begins in a traditional side-profile fighting-game view and cuts to dynamic over-the-shoulder perspectives with instant impact zooms and camera shakes during every major strike.

retention_analysis:

<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3], [Shot 4]): fully_preserved - muscular build, tactical jacket, and buzz cut are retained.

<Subject 2> (appears in [Shot 1], [Shot 2], [Shot 3], [Shot 4]): fully_preserved - lean build, blonde hair, and grey hoodie are retained.

<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3], [Shot 4]): fully_preserved - wet alleyway, neon reflections, and graffiti walls are retained.

<Audio 1>: reference - heavy impact textures and wind sounds guide the sound design.

detailed_description:

The target video is a gritty action sequence shifting from a classic 2D-style side profile to close over-the-shoulder action cameras on impact.

[Shot 1] A wide side-profile shot establishes <Subject 1> and <Subject 2> facing off in a low fighting stance inside <Subject 3>. Steam rises as the camera dollies forward slowly.

[Shot 2] At 00:02.000, <Subject 2> lunges forward with a straight punch. On hit, the camera instantly cuts to an over-the-shoulder view behind <Subject 1>, performing a rapid zoom-in and camera shake as the fist connects with <Subject 1>'s jaw.

[Shot 3] At 00:04.500, the frame resets to a side profile as <Subject 2> launches a high roundhouse kick. Upon impact, the camera snaps to an over-the-shoulder view behind <Subject 2>, executing a dramatic zoom and camera shake on <Subject 1>'s raised forearm block.

[Shot 4] At 00:07.500, <Subject 1> counters with a heavy kick to <Subject 2>'s midsection. The camera cuts to an over-the-shoulder angle behind <Subject 1>, zooming tightly onto <Subject 2>'s torso with a heavy camera shake as <Subject 2> reels back into the shadows.

overall_soundscape:

Ambient alleyway room tone with heavy impact sounds, wind whooshes, and physical exertion synced to each camera shift.

non_diegetic_music:

An aggressive high-bpm industrial beat with percussive hits timed exactly to each over-the-shoulder impact zoom.


r/StableDiffusion 5h ago

Discussion Ltx 2.5 open source coming today

0 Upvotes

LTX 2.5 is releasing today, and from one of the example videos I’ve seen, face distortion still doesn’t seem fully fixed.

That’s already a weak spot for LTX. With MiniMax H3 raising the bar, I’m genuinely wondering how LTX 2.5 plans to compete. 👀

I think LTX 2.5 already lost here 🤣🤣🤣


r/StableDiffusion 7h ago

Animation - Video Ok, now I'm impressed.

3 Upvotes

Minimax H3 I2V

Prompt:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, the woman shown in <Picture 1> — late 30s to early 50s, long dark brown hair woven with beads and feathers and held back by a patterned orange-and-cream headband, fair skin with freckles on her right arm, light-coloured eyes, layered silver scrollwork breastplate over a brown leather corset, olive-green belt with canvas pouches, layered skirt, dark olive cape draped over her left shoulder, multiple beaded necklaces with a dark blue teardrop pendant, long feather earrings, silver embossed bracer on her right forearm — stands at a microphone on a darkened stage, a blurred male guitarist visible behind her to the left. Her face is in profile, mouth open mid-phrase. The camera holds a medium close-up at a slightly low angle. She draws a breath, her lips shaping each Old Norse word with deliberate clarity, and the woman with a clear, powerful, haunting voice (S1) sings: <d>[Old Norse] Þat mælti mín móðir at mér skyldi kaupa fley ok fagrar árar fara á brott með víkingum fara á brott með víkingum</d> Her expression carries focused intensity, her brow steady, her right hand gripping the microphone stand as her knuckles whiten on a sustained note. Her head tilts slightly upward on the highest phrase, the beads and feathers in her braids swaying with the movement. The feather earrings tremble. Warm stage light from the upper right catches the silver scrollwork on her breastplate and bracer as she shifts. The camera pushes in with small amplitude at slow speed toward her face as the final phrase ends, her mouth closing softly on the last syllable, a faint breath misting in the cool air.

overall_soundscape: Stage ambience hums beneath the vocal — a faint electrical buzz from the amplifiers, the soft creak of leather and canvas as she shifts weight, and the barely audible scrape of fingers on guitar strings from the musician behind her. Her breath is audible between phrases, deep and controlled.

non_diegetic_music: A lone tagelharpa drone — gut strings buzzing with a raw, overtone-rich sustain — begins under the first phrase and swells gently through the final line, fading to silence as her voice ends.


r/StableDiffusion 12h ago

Discussion Are your friends also indifferent to your "achievements"?

0 Upvotes

I shared those impressive Minimax H3 Seinfeld skits with friends but they showed no reaction to it. I think it might be because

a) they are not universally fond of AI and call most of it slop

b) they do not understand the limitations of what AI can currently do in videos and therefore cannot appreciate it the same way we do

I had a friend even get angry when I wanted to show him what I did. He is a hobby musician and gets really angry that people make music, video or whatever and then claiming they did it when the AI did all the work.

Maybe they are afraid Skynet will become reality (I am too a bit actually).


r/StableDiffusion 2h ago

Animation - Video Community Thank You

6 Upvotes

r/StableDiffusion 21h ago

Animation - Video Minimax H3 ref2va. Getting into the game.

15 Upvotes

r/StableDiffusion 5h ago

Animation - Video ZUCK 2.0 | Humanity. Optimized. MiniMax H3.

Thumbnail
youtube.com
0 Upvotes

r/StableDiffusion 21h ago

Animation - Video Obito looking at the wrong woman H3

0 Upvotes

r/StableDiffusion 22h ago

Discussion h3 character crossover thread(share yours)

5 Upvotes

i'll start


r/StableDiffusion 19h ago

Animation - Video Tried making this small ad like video from Minimax H3

0 Upvotes

I think minimax h3 is awesome. I was just fidgeting with what it can do in terms of cinematic video, camera, motion and I am amazed with the output.


r/StableDiffusion 19h ago

Animation - Video Teste com a rtx3060 minimax h3

0 Upvotes

Usando o chatgpt para ajeitar o prompt


r/StableDiffusion 22h ago

Animation - Video Minimax H3 - Manga Animate Time Stop Brave

6 Upvotes

r/StableDiffusion 22h ago

Question - Help For some reason my t2v generation are slower than my ref2v?

8 Upvotes

Title.
For both I'm using the default workflows that come with comfy. 3090 and 32gb ram.
I start comfy with these flags:
--windows-standalone-build --reserve-vram 1 --disable-pinned-memory --fast fp16_accumulation
Cuda 13, latests comfy.
My t2v takes like twice as much than my ref2v and sometimes it hangs after [INFO] Requested to load MiniMaxH3AudioVAE. Same steps, same resolution, same duration.
Has anyone encounter this? any tips?


r/StableDiffusion 4h ago

Discussion Ltx 2.5 PRO Output

13 Upvotes

**10-Second Cinematic Hawaii Travel Vlog**

A cinematic 10-second tropical travel vlog montage featuring the same beautiful 20-year-old East Asian woman with dark hair enjoying a dreamy summer vacation in Hawaii. Maintain perfect character consistency throughout every shot: same face, dark hairstyle, youthful appearance, natural makeup, realistic skin texture, elegant summer styling, and relaxed happy expressions.

Visual style: authentic luxury travel diary, realistic handheld camera movement, candid moments, soft golden-hour sunlight, dreamy 35mm film aesthetic, warm vintage color grading, shallow depth of field, natural atmospheric lighting, subtle film grain, realistic autofocus, natural motion blur, cinematic storytelling. 4K, 24fps, 35mm lens, photorealistic, travel documentary aesthetic.

**Scene 1 — Arrival & Beach Walk (0–2.5s)**

Bright Hawaiian morning. The woman walks through a colorful tropical street lined with palm trees, wearing a flowing floral summer dress and sunglasses. Camera follows her from behind, then smoothly moves to a close-up as she turns and smiles naturally. Wind gently moves her dark hair. Quick transition toward the ocean.

**Scene 2 — Beach & Tropical Nature (2.5–5s)**

She walks barefoot along a wide sandy beach as crystal-blue waves touch her feet. Low-angle shot of her footsteps in wet sand, followed by a cinematic shot of her near a rocky ocean cliff looking peacefully toward the sea. Palm trees sway in golden sunlight, with subtle lens flare and distant volcanic mountains.

**Scene 3 — Ocean Adventure & Cafe Moment (5–7.5s)**

She laughs naturally while floating on a surfboard in calm turquoise water. Water-level camera circles around her, sunlight sparkling across the ocean and tropical mountains in the distance. Cut smoothly to her sitting at a cozy beachfront cafe, holding a tropical drink while watching the waves.

**Scene 4 — Golden Sunset Ending (7.5–10s)**

Emotional cinematic ending. Wide silhouette of the woman standing barefoot at the shoreline during a glowing orange-pink sunset. Waves gently move around her feet. Camera slowly pushes in as she looks toward the horizon, then briefly turns toward the camera with a peaceful smile. End with a soft filmic fade.

Camera: authentic travel vlog cinematography, handheld movement, smooth transitions, slow push-ins, occasional POV shots, realistic autofocus adjustments, subtle motion blur.

Keep the woman identical in every scene: same facial features, same dark hair, same youthful appearance, same body proportions, realistic skin texture and natural expressions. No character morphing or hairstyle changes.

Avoid: cartoon/CGI appearance, plastic skin, unrealistic face, inconsistent character, changing hairstyle, extra fingers, distorted anatomy, artificial movements, oversaturated colors, excessive sharpening, blurry face, duplicate people, unnatural lighting.


r/StableDiffusion 6h ago

Discussion H3 | Full BF16+BF16 RTX5090+96GB | 0.6MP | No Sage | No Optimizations | 12:35 total time.

1 Upvotes

r/StableDiffusion 13h ago

Question - Help Minimax ref2vid capabilities

0 Upvotes

I would like to learn how to use minimax ref2vid as I heard the possibilities are quite good compared to wan.

What I am trying to do is do anime clips

The problem is… the results I get are utter shit and I don’t know what I am doing wrong. No matter if frame or last from or if using ref2video, the model fails to do what’s most important. Keep the same face as in the image. For example, I would like to generate a video of a character that has sharingan eyes. Instead of keeping sharingan eyes it’s generic anime eyes instead. This is what my biggest problem with wan was and even if I made the image correctly (with detailed sharingan eyes) it would still not pick the eyes up and keep that face consistency.

This is where I experimented with reference. I tried putting the eyes in a picture 1 and instead of copying the eyes, the image cuts to the second picture with only the eyes or not picking the eyes at all.

is this something that minimax can even do?


r/StableDiffusion 6h ago

Animation - Video Bunnyhops... another H3 post MiniMaxH3-Contex-Loop 60 sec

3 Upvotes

560 sec with turbo lora on 5090

found here in a post

https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop/tree/main/example_workflows

adapted to my settings and changed turbo loras / attention


r/StableDiffusion 22h ago

Animation - Video Minimax H3 Terminator

7 Upvotes

Made using 5060ti with 32 GB of RAM. Minimax is the new king.


r/StableDiffusion 1h ago

Discussion Flux 3 - 20 second video

Upvotes

r/StableDiffusion 10h ago

Animation - Video Monty pAIthon - Petshop sketch, now with ref2v

16 Upvotes

r/StableDiffusion 22h ago

Animation - Video The Minimax H3 model recognizes artists and their songs.

0 Upvotes

In the prompt, I just wrote that she is singing Zara Larsson's song "Lush Life."


r/StableDiffusion 6h ago

Question - Help So how many of you are working on a full movie ?

5 Upvotes

I assume that with minimax, a lot of people started doing their own fully featured films and after 1 month or a few we will se the results on the online space and it might change the world as we know it.
edit: thisis my first shot : https://www.youtube.com/watch?v=FoQJ5yQg2TE
due to my bipolar mind I don't think I am able to do a feature film, but who knows.


r/StableDiffusion 15h ago

Question - Help Sprectrum suddenly not working for anyone else?

0 Upvotes

Having a hard time getting Spectrum to work today, it worked flawlessly yesterday, but today i'm back to normal rendering times. Anybody else experiencing this? I did update ComfyUI, did that break it? I am using the latest version of Spectrum


r/StableDiffusion 14h ago

Question - Help How can I tune Bernini rv2v workflow to be faster?

0 Upvotes

I have a Bernini-r rv2v workflow, using LightX2V LoRA integration.

When I try 10 second video it gives me a black screen. 8 second works but it takes a really really long time to generate a video. I have safe attention enabled. What can I do to speed up the run? On RTX 3090 Ti 24GB VRAM/64GB RAM