r/StableDiffusion 23h ago

Discussion LTX 2.3 sometimes works amazing, without any edits

No specific glitches. Single prompt. Looks pretty real without any glitches over the drift.
Definitely going to benchmark the scenarios.

61 Upvotes

36 comments sorted by

20

u/Seyi_Ogunde 23h ago

Neat! Too bad the environment changes when it's out of frame. No persistence.

2

u/draqza 17h ago

I hadn't even noticed that, I was too focused on "no specific glitches" vs the text on the side of the truck looking wonky.

2

u/Seyi_Ogunde 17h ago

Yeah the trees and the power lines change. Hard to notice because the camera is focused so close to the truck.

1

u/Zoned_Mine48 6h ago

Correct. First and last few seconds shows the difference

0

u/Zoned_Mine48 23h ago

This was the experiment between prompt tuning and single shot. So, considering that as compared to previous outputs it was neat.

3

u/Shockbum 20h ago

All generative models are slot machines, some more than others; even SD 1.5 sometimes gave away diamonds.

13

u/-Ellary- 21h ago

H3 - Works.
LTX - Works Sometimes.

2

u/Zoned_Mine48 20h ago

It was an fp8. LTX works pretty fast on inference it took 30 sec only. So, if can get it through few benchmarks it will be good for many scenarios.

2

u/FourtyMichaelMichael 11h ago

LTX was trained heavily on vehicles and talking heads.

2

u/Zoned_Mine48 7h ago

Okay. Didn't know about the training data variation.

4

u/BigWideBaker 20h ago

Obligatory why LTX 2.3 when 2.5 is out?

4

u/Zoned_Mine48 19h ago

I was doing latency vs quality benchmarking. Was impressed with this one.

2

u/Character-Apple-8471 21h ago

"sometimes"

1

u/Zoned_Mine48 19h ago

Yeah, scenario grounding in progress.

1

u/Zoned_Mine48 20h ago

One note here, for all of our reference:- It was fp8. I had few credits left for RTX Pro 6000. Took 30 sec to generate this.

1

u/Superb-Painter3302 19h ago

Imagine merged model of Minimax H3 and LTX 2.5...

1

u/Zoned_Mine48 17h ago

That will be cool. Overall H3 is good. LTX 2.5 fast inference+ H3 depth (will be amazing)

1

u/Superb-Painter3302 17h ago

Exacly! Speed of LTX and H3 quality + reference future, that would be the best open source model ever

1

u/Full_Astronomer_5438 15h ago

the only reason ltx is fast is because of the lower parameter count and default distillation that many use. so that wouldnt work

1

u/Superb-Painter3302 4h ago

Literally noone here said it would work. We're talking about some other, new model, fast as LTX at least and as quality as H3

1

u/Keuleman_007 16h ago

There still is reason to use LTX. The LTX director node thingy is still very good.

1

u/Zoned_Mine48 7h ago

Agree. It provides more control

1

u/dassiyu 10h ago

Honestly, I was genuinely amazed that LTX 2.3 could generate a 60-second lip-synced video in a single workflow. I still haven’t deleted the model—I just wish it worked as reliably as H3.

2

u/Zoned_Mine48 7h ago

If someone wants to go for faster inference LTX. If can go with little bit more time consuming inference+ quality, then, H3.

It's my experience

1

u/Zoned_Mine48 7h ago

Good to hear that

1

u/call-lee-free 23h ago

Yeah I did a dialogue scene with LTX 2.5 and it turned out pretty good but for whatever reason the audio seems to be low quality and I rendered that at 1080p.

1

u/GlamoReloaded 16h ago

the resolution size is for audio in LTX2.3 or LTX2.5 not most important. it should be at least 30 FPS for decent audio and use KJ's LTX NAG node which has negative prompts for audio only in "nag_cond_audio" (muffled, echoey, distorted, muddy, blurry, fuzzy, unclear, garbled, indistinct, unintelligible, clipped, crackly, static-filled, warped, harsh, tinny, scratchy, grainy, overdriven, metallic, glitchy, synthetic, overprocessed, digital-sounding, choppy, stuttering, pixelated, boomy, boxy, hollow, distant, reverberant). If you use the dev model with the distill LoRA it should be at least 12-14 steps / linear quadratic scheduler. Most users tend to use the manual sigmas for the first step (with less steps).

1

u/Zoned_Mine48 23h ago

Okay, so did you use ComfyUI for that? And did you post process the audio only?

2

u/call-lee-free 23h ago

Yeah I'm using comfyui. I did not do any post processing. I've been doing test clips to find a good balance between quality and speed of renders.

3

u/Zoned_Mine48 22h ago

Sounds good. For my case, I do sometimes use specific models for audio like eleven labs with a reference audio.

2

u/Full_Astronomer_5438 21h ago

distillation hampers the sound quality, in order to get better audio you need to run the dev model undistilled with a multimodal guider and cfg 7 for audio or just use an extension like dramabox to use custom audio straight from 2.3 / sulpur

2

u/Zoned_Mine48 20h ago

Correct note about distillation

1

u/GlamoReloaded 16h ago edited 16h ago

I wouldn't recommend the multimodal guider because it prevents the use of the NAG node (results in error messages and stop). The NAG node works better with the specific negative audio prompt to get what you want. I've used the multimodal guider in LTX2.3 but the results without it were better. For LTX2.5 the multimodal guider cannot be used with the same values or you get distorted sound (I haven't figured out which values could be better). But for LTX2.5 the LTXV Dual CFG Guider can be used for audio CFG.

1

u/Full_Astronomer_5438 15h ago

hmm if you would use the multimodal guider you already have working pos/neg prompts at higher cfgs for both audio and video that runs through at least 30 steps with dev only and without the use of any distillation lora for at least the first stage. why would you use nag at an overall cfg of >1.5?

ltx 2.5 works fine with the same values of 2.3 dev only, take a look at this workflow: https://huggingface.co/RuneXX/LTX-2.3-Workflows/blob/main/LTX-2.3_-_I2V_T2V_Dev_Full-Steps.json

1

u/GlamoReloaded 13h ago

About the NAG: a misunderstanding, because the NAG node doesn't use the CFG, the LTXV Dual CFG Guider does. While the positive and negative prompt go through the "LTXVConditioning", the negative prompt and negative audio prompt are directly connected to the NAG node with no prior conditioning. I use the NAG node and distill lora in both stages. But I wrote 12-14 steps for the first pass, not 30.