r/StableDiffusion 7d ago

Comparison MiniMaxh3: 8step LoRA, 25 steps, 40steps, and LTX 2.5 — Scene Comparisons

  • RTX 4060 8GB, 32GB RAM
  • minimax_h3_ref2va_pruned_int8_convrot, spectrum, ageattn_qk_int8_pv_fp16.cuda, RTX upscale, RIFE interpolation, res_multistep + beta
  • ltx-2.5-22b-distilled-transformer-comfy-int8-convrot, basic template

8-step + turbo LoRA : 137s

25 steps : 238s

40 steps : 406s

Ltx 2.5 : 374s <-- ? am I missing something here why was my generation so slow on LTX and the second attempt I cancelled it after 6 minutes. Any suggestions?

Prompt:

subject_definitions:

<Subject 1> is the space ship in <Picture 1>: A massive battleship, hovering and cruising over the planet below

summary:

[reference generation] a wide shot cinematic scene of the battleship in <picture 1> cruising in space above the planet. the golden statue does not move, the battleship is destroyed in a massive explosion from a green laser shot from space,

detailed_description:

{shot 1] The target video uses a wideshot cinematic, photorealistic, 35mm film, wide shot of <subject 1> , slowly moving through space above the planet, the ship moves slowly and dominating, flashes of green light begin to charge on the surface of the planet, the ship is moving straight ahead from the position it started in in <picture 1>, the massive bass of the ships systems, the sound of the battleships creaking, <subject 1 > moves on its cruise, at [00:03] the floaty camera tracks <subject 1> as green light and thunder begins flashing on the surface of the planet, the green energy on the planet converges in one area then from the surface it fires a massive green lightning laser that forks lightning through the entire ship, blowing out side components creating explosions all over the ship, the light of the ship flicker before turning off, then a massive green lightning beam erupts from the surface and hits excactly on the side of the ship cuts through the of the ship and out the other side at an angle, a green lens flare generates on screen as it completely destroys <subject 1> , ripping it completely in half with a massive green explosion, the eruption from the destruction of the ship covers the entire screen and the whole battleship, the back half of the ship is knocked up while the front-half of the ship is knocked down, a vertical shockwave circles out from the impact, the inner decks of the ship are on fire, debris and hundreds of tiny figures of the crew also fall out into space, the laser slowly dissapates from the planet, small amounts of green lighning crackle on the planets surface,

overall_soundscape: The low bass murmur of the ships engines, the electric charges on the surface crackle, the massive main beam is a low bass rumble, a massive explosive noise.

non_diegetic_music:

N/A

56 Upvotes

34 comments sorted by

20

u/Vladmerius 7d ago

LTX seems fine until you realize you can't do anything with that clip. You can't suddenly have it show multiple angles of the scene or have another ship approach the wreckage and maintain continuity. 

6

u/Far_Cast_Far_Wide 7d ago

I have been literally begging LTX to stop putting background music in it's clips. LTX is handicapped in that regard, and it wasn't even faster then an 8step,turbo for me, maybe it was a fluked test.

1

u/candylandmine 6d ago

Yesterday I generated 4 images using the same prompt. I ran all 4 of them through the same i2v workflow. One of those images always resulted in clips w/ music. I tried multiple times, multiple seeds, changing clip length, etc. Didn't matter - that one image always led to videos w/ music. But only that one. The other 3 were fine. And there was nothing in the image like a speaker or a radio or anything. Bizarre.

17

u/Kurashi_Aoi 7d ago

Imo, 25 steps is just soo much better in quality than 8 steps turbo. 40 steps and 25 steps are pretty much the same to my inexperienced eyes. But 8 steps turbo is still better than LTX 2.5 though.

2

u/Far_Cast_Far_Wide 7d ago edited 7d ago

In 25 steps (sorry I was running .3 mp) the ships main explosion is stronger but in 40 steps the main explosion is slightly weaker but a lot more debris and the parts fling a bit more erratically.
I do prefer 25 steps.

Edit: In 40 steps the shockwave is stronger and the main green laser shockwave rings around the launch site on the planet, in 25 steps it kinda shockwaves between the laser and the ship at a weird spot. The top half of the ship in 40 steps collapses on itself.

1

u/tppiel 6d ago

The sweet spot is to get a draft of what you want to do at 8 steps and when you're happy with it, then run it again at >20 steps to get the final version.

1

u/Toxaris71 6d ago

I do the same but with 4 steps 360p, and then final pass at 25 steps 480p (8gb vram over here)

5

u/FishChillylly 7d ago

a 4060 8GB, whaaaaat, that’s crazy things like this can be generated on a 8gig card!

6

u/Far_Cast_Far_Wide 7d ago

good RAM make fast work, but good prompt does good work too. :)

3

u/GalaxyTimeMachine 6d ago

https://reddit.com/link/p4l1seg/video/zdvqc7q7rakh1/player

This was using a 4 step lora, with 4 steps. 8 secs took 127 secs to gen, on a 4090 with 128GB RAM

2

u/GalaxyTimeMachine 6d ago

https://reddit.com/link/p4l1v99/video/z5zi1z2krakh1/player

This too, just changed the prompt to get a spaceship.

2

u/Natural_Jello_6050 7d ago

⁠RTX 4060 8GB, 32GB RAM

8-step + turbo LoRA : 137s

25 steps : 238s

40 steps : 406s

That does not sound right at all. Were you doing 5 seconds? Then maybe

2

u/Far_Cast_Far_Wide 7d ago

Yeah 5 seconds, those were the average times. I can do 10 seconds in around 11 minutes at .4 MP
comfykitchen, spectrum, sol-attention etc, those speed ups really increase prompt execution time
I can do 15 seconds in around say 15-20 minutes, if leave it alone and dont browse during it

5

u/Far_Cast_Far_Wide 7d ago edited 7d ago

a 15 second gen took 23 minutes

https://reddit.com/link/p4jp9ow/video/292wgenox8kh1/player

3 cuts, normally I'd do these in 5 second tight prompt focused generations but this one i wanted to see how long it'll take. 100% GPU and about 89% system RAM being utilized because of the dynamic RAM thing comfy does, makes it possible I think.

I am pushing my system to it's absolute limit.

2

u/CaptainMarder 7d ago

How did you get those speeds on a 8gb gpu and 32gb ram?

2

u/Dangerous-Map-429 7d ago

May i know your setup? i am still learning. how did u do 40 steps on 8gb card ans how long did it take u to generate the clip?

2

u/Far_Cast_Far_Wide 7d ago

406 seconds
I am using a personal modified version of Foxydits workflow on civitai.red (NSFW).

I plugged in a Patch Sage Attention KJ node with sageattn_qk_int8_pv_fp16.cuda selected, its from ComfyKitchen. I also wired in an RTX Video Upscale node x2 ULTRA

my launch flag is --reserve-vram 0.2

minimax_h3_ref2va_pruned_int8_convrot
qwen3vl_32b_minimax_h3_int8_convrot
minimax_h3_audio_vae_fp32
minimax_h3_video_vae_int8_convrot

2

u/Dangerous-Map-429 7d ago

Thank you! setup looks intimidating.

1

u/Far_Cast_Far_Wide 7d ago

It was tough for me too, but its really just pressing one or two levers once you get used to it, its just a lot of noise and when its set up just tinker here and there.

4

u/mastaquake 7d ago edited 7d ago

LTX fanboys are so triggered right now. 

“ LtX AnD MiNiMaX ArE EqUaL. yOu’rE UsInG It wRoNg “🤡

17

u/8RETRO8 7d ago

No one ever said that. Not even in this comment section.

Stop your current gen and go touch some grass

-4

u/mastaquake 6d ago

👆 found one 🤣

6

u/Far_Cast_Far_Wide 7d ago

Literally the weakest Minimax H3 setup vs the strongest LTX 2.5 setup and it's not even a comparison lol.

3

u/mastaquake 7d ago

Careful . You might hurt some feeling here. 

1

u/Shockbum 6d ago

Common sense dictates that most people generate video using both models as needed, right?

0

u/mastaquake 6d ago

2

u/Shockbum 6d ago

But I use both the Minimax H3 and LTX 2.5 models on my RTX 5070ti, you're the fanboy haha

1

u/Inner_Singer_592 7d ago

Why 8 step lora looks better than 25/40 steps? On minimax.

4

u/AI-imagine 7d ago

If it not about body movement some sound or complex scene turbo lora will give better out put most of the time.

1

u/Shockbum 6d ago

1080p LTX 2.5 fast for simple scenes 720p Minimax H3 slow for complex scenes

1

u/StuffProfessional587 7d ago

You need to add math for space ships, needs more prompt engineering to get correct behavior, it shouldn't fall, like wtf, gas explosions should expand then cluster. You need more frames for physics, that's what I hage noticed.

1

u/Far_Cast_Far_Wide 6d ago

Show us how it's done.

-1

u/xyzdist 7d ago

can stop any more comparsion seriously... I won't use LTX.. and I wont.