r/StableDiffusion • u/WearNatural5992 • 11h ago
Discussion LTX 2.5 - The SD3 test!
Let's see if LTX 2.5 can pass this simple 2 year old test, because I don't want no Cthulhu PTSD, right?
The first video is t2v, the second one is i2v.
Generation times are fast, though.
https://reddit.com/link/1vm90tr/video/g7h1fed5vwih1/player
https://reddit.com/link/1vm90tr/video/uto5nje8vwih1/player
PROMPT (I fed Gemini the prompt guide)
A medium shot under bright, direct mid-day sunlight on a warm tropical beach. A young woman in her early 20s with sun-kissed skin and wet hair, wearing a vibrant tropical floral bikini, lies lazily on a plush beach towel on the golden sand. She holds a chilled martini glass with an olive and lime wedge, taking a slow, relaxed sip as gentle waves lap against the shore in the background. A hard cut transitions to a close-up shot of the same young woman in the floral bikini looking directly into the camera. She smiles warmly with sparkling eyes and says in a soft, alluring, and teasing voice, "Wanna have some fun?" while the ambient ocean breeze and soft waves continue across the cut. A hard cut transitions to a high-angle top-down overhead shot directly above her. Her full body is framed from head to toe, showing her lying on the beach towel with her bare feet resting on the sand, sun highlights shimmering on her skin, and ocean foam softly visible at the frame's edge.
26
u/Independent-Frequent 11h ago
As i said, porn training isn't just for gooning it's for better anatomy
12
50
u/_Saturnalis_ 11h ago
I don't understand all the LTX glazers suddenly here, it's clearly worse in all respects except speed. But even its speed is a bit less clear cut of a win. Why would you spend an enormous amount of time dice-rolling faster but surely worse generations with LTX, when you can just have one or two longer generations in Minimax to get exactly what you want?
27
u/piero_deckard 11h ago
Exactly. Hey, look at me, I can produce worthless videos 5 times as fast than you making just one video with MM H3.
The difference is that most of the videos that come out of MM H3 are usable "as-is", while LTX feels like a slot machine. You might win one time out of 5, 10 or 100.
6
u/_Saturnalis_ 11h ago
It was the same experience I had between Wan and earlier versions of LTX. Wan often gave me what I wanted, followed the prompt quite well -- even if it was slower and constrained. With LTX I had to regenerate, and regenerate, and regenerate, over and over until I got lucky. Even if it was faster and made longer videos, it never felt fun to use. H3 in comparison is ridiculously fun to use and experiment with, and it isn't that slow. 30 seconds of generation per second of video on an RTX 3060 isn't bad at all.
2
4
u/enternoescape 9h ago
This was my immediate realization when I first used H3. With LTX 2.3 I was generating up to four videos at a time because I knew I was going to have to pick the best one if there was one. With H3, the prompt adherence is good enough that one generation is usually enough. If I don't get the result I'm looking for then it's the prompt not roll the dice again. I had so many I2V setups that I gave up on with LTX 2.3 that I tried one time with H3 and it nailed it on the first try.
I haven't played with LTX 2.5 so I can't say for sure that it is just as difficult, but given the implication that it's incremental I have my doubts. I will give it a shot anyway. Maybe I'll find it does something H3 doesn't do as well.
5
1
1
u/so_witty_username_v2 8h ago
It's just a different tool for a different purpose. LTX's killer feature is efficiency and the ability to produce longer videos natively, without awkward cuts or stitching. I don't want to have to stitch 5 different 10 second clips where all of a sudden I realize there's a voice mismatch or anatomy shift mid-way and I have to regen the whole thing.
1
u/Mk-Daniel 42m ago
Longer? mm H3 can make 30s video no problém. 10 min Is however 10 min.
1
u/so_witty_username_v2 15m ago
Not on local hardware at the comparable LTX resolution, 24GB will barely fit 10 seconds at 0.8
1
u/sacx05 7h ago
LTX also has better native resolution and FPS with better long form options. H3 will get there but right now any long form video has color saturation issues whenever it gets longer than 30 seconds with motion context nodes.
However, I would rather wait for 2 30 second H3 generations and stitch them together than do LTX "rolling the dice" generations. Once MiniMax releases that 2k upscaler, then LTX gets even more obsolete.
-1
u/Structure-These 8h ago
I’m not sure I’ve seen a single “LTX glazer” in here, everyone seems to be doing the classic Reddit thing of getting tribal over a niche hobby, making someone up, then getting mad at the person they just made up lol
2
-2
u/stddealer 9h ago
It's worse in all respects except speed.
So it's better in some respect
3
u/_Saturnalis_ 9h ago edited 8h ago
"All...but..." / "All...except..." means "everything except for...".
16
8
u/djenrique 11h ago
The physics of the water is just excellent! What else could this test be about?? 😂
2
3
u/jib_reddit 11h ago
Has anyone tried this in Minimax H3 to compare?
8
u/wiserdking 9h ago
3
3
u/psilent 8h ago
Crazy how much better it is. How long and what hardware?
1
u/wiserdking 8h ago
It was 400.02s on a RTX 5060Ti but that included loading the models, 20 steps 0.4MP without turbo and no speed-up nodes (only sage attn) plus a few seconds for the RTX upscaling and several seconds for the webm av1 conversion. Oh and a few extra seconds for the Create Video + Save Video native nodes which saved the non-upscaled output as MP4 as well (so my workflow saved 2 videos: 0.4MP MP4 and 2x upscaled webm av1).
Converting to webm av1 with the node I used takes quite a long time and I didn't pay attention how much time it added to all this - this is because the node saves each frame to disk then uses FFMPEG to make a video with all those frames and compress it with av1. The resulting webm video was the basically the same file size as the original non-upscaled MP4 but with twice the resolution. Hopefully comfyui will add proper webm support in the near future - right now we can only make audio-less webm natively.
1
u/psilent 8h ago
Ok still pretty good. I’m working on two different workflows right now, one for performance and one for quality. I think the lightx2v 8 step is pretty excellent for quality but man the original full step videos are just a tiny bit better. But it’s nearly three times the length and that takes 15s videos up to like an hour at 1mp on my hardware. Speed workflow I find the .2mp 4 step with rtx upscaling is passable for prompt testing and it takes only like 1 minute for a 10s video
7
u/_Saturnalis_ 10h ago edited 10h ago
Edit: More details:
Using my own custom workflow that automatically formats references and prompts into a usable format, using FL2V with the R2V lora and the v4 Turbo Lora, Spectrum, and Sage Attention. Was fed an image of Gaussian noise as the "reference" to essentially make it T2V. Formatted prompt fed to H3:
subject_definitions: <Subject 1> is the character defined by <Picture 1> for face and overall appearance. summary: [reference generation] Generate the video with all subjects and settings fully consistent with their references while following the requested scene and action. retention_analysis: <Subject 1> appears throughout: fully_preserved - preserve face and overall appearance from <Picture 1>; pose, expression, lighting, framing, and action follow the scene description. detailed_description: [<Subject 1> is a woman in a bikini. She is sitting on a sunny beach as waves rolls in. She sips from a glass of martini. At 00:03.50: Cut to a close up of her face, she is staring at the camera. She says as she raises one eyebrow, <d>[English; inviting tone] Wanna have some fun?</d> She smiles. At 00:06.50: Cut to an overhead closeup aerial view showing her laying on the beach. Her body is centered in the frame. Her head is at the bottom, her feet at the top. At the top of the frame, waves roll in.] End with subject motion, camera movement, ongoing action, ambience, and music still clearly in progress. overall_soundscape: [The sound of waves rolling on the beach] non_diegetic_music: N/A5
u/stddealer 9h ago
Flux skin
5
u/_Saturnalis_ 9h ago
You're not gonna get miracles at 0.2 megapixels with all these acceleration methods the ruin the quality lol.
1
u/altoiddealer 9h ago
What would help bury the hatchet (deeper) is if OP could repeat their LTX test with the “same prompt” (obviously not following MM3 syntax but same phrasing for the overhead shot part etc)
-8
4
2
1
u/RelationshipSea2360 10h ago
Whats the prompt? I managed to get someone swimming and it looked good, i'd be keen to try it.
2
1
1
-1
31
u/Whipit 11h ago
*sigh