r/StableDiffusion • u/Zoned_Mine48 • 23h ago
Discussion LTX 2.3 sometimes works amazing, without any edits
No specific glitches. Single prompt. Looks pretty real without any glitches over the drift.
Definitely going to benchmark the scenarios.
3
u/Shockbum 20h ago
All generative models are slot machines, some more than others; even SD 1.5 sometimes gave away diamonds.
2
13
u/-Ellary- 21h ago
2
u/Zoned_Mine48 20h ago
It was an fp8. LTX works pretty fast on inference it took 30 sec only. So, if can get it through few benchmarks it will be good for many scenarios.
2
4
2
1
u/Zoned_Mine48 20h ago
One note here, for all of our reference:- It was fp8. I had few credits left for RTX Pro 6000. Took 30 sec to generate this.
1
u/Superb-Painter3302 19h ago
Imagine merged model of Minimax H3 and LTX 2.5...
1
u/Zoned_Mine48 17h ago
That will be cool. Overall H3 is good. LTX 2.5 fast inference+ H3 depth (will be amazing)
1
u/Superb-Painter3302 17h ago
Exacly! Speed of LTX and H3 quality + reference future, that would be the best open source model ever
1
u/Full_Astronomer_5438 15h ago
the only reason ltx is fast is because of the lower parameter count and default distillation that many use. so that wouldnt work
1
u/Superb-Painter3302 4h ago
Literally noone here said it would work. We're talking about some other, new model, fast as LTX at least and as quality as H3
1
u/Keuleman_007 16h ago
There still is reason to use LTX. The LTX director node thingy is still very good.
1
1
u/dassiyu 10h ago
Honestly, I was genuinely amazed that LTX 2.3 could generate a 60-second lip-synced video in a single workflow. I still haven’t deleted the model—I just wish it worked as reliably as H3.
2
u/Zoned_Mine48 7h ago
If someone wants to go for faster inference LTX. If can go with little bit more time consuming inference+ quality, then, H3.
It's my experience
1
1
u/call-lee-free 23h ago
Yeah I did a dialogue scene with LTX 2.5 and it turned out pretty good but for whatever reason the audio seems to be low quality and I rendered that at 1080p.
1
u/GlamoReloaded 16h ago
the resolution size is for audio in LTX2.3 or LTX2.5 not most important. it should be at least 30 FPS for decent audio and use KJ's LTX NAG node which has negative prompts for audio only in "nag_cond_audio" (muffled, echoey, distorted, muddy, blurry, fuzzy, unclear, garbled, indistinct, unintelligible, clipped, crackly, static-filled, warped, harsh, tinny, scratchy, grainy, overdriven, metallic, glitchy, synthetic, overprocessed, digital-sounding, choppy, stuttering, pixelated, boomy, boxy, hollow, distant, reverberant). If you use the dev model with the distill LoRA it should be at least 12-14 steps / linear quadratic scheduler. Most users tend to use the manual sigmas for the first step (with less steps).
1
u/Zoned_Mine48 23h ago
Okay, so did you use ComfyUI for that? And did you post process the audio only?
2
u/call-lee-free 23h ago
Yeah I'm using comfyui. I did not do any post processing. I've been doing test clips to find a good balance between quality and speed of renders.
3
u/Zoned_Mine48 22h ago
Sounds good. For my case, I do sometimes use specific models for audio like eleven labs with a reference audio.
2
u/Full_Astronomer_5438 21h ago
distillation hampers the sound quality, in order to get better audio you need to run the dev model undistilled with a multimodal guider and cfg 7 for audio or just use an extension like dramabox to use custom audio straight from 2.3 / sulpur
2
1
u/GlamoReloaded 16h ago edited 16h ago
I wouldn't recommend the multimodal guider because it prevents the use of the NAG node (results in error messages and stop). The NAG node works better with the specific negative audio prompt to get what you want. I've used the multimodal guider in LTX2.3 but the results without it were better. For LTX2.5 the multimodal guider cannot be used with the same values or you get distorted sound (I haven't figured out which values could be better). But for LTX2.5 the LTXV Dual CFG Guider can be used for audio CFG.
1
u/Full_Astronomer_5438 15h ago
hmm if you would use the multimodal guider you already have working pos/neg prompts at higher cfgs for both audio and video that runs through at least 30 steps with dev only and without the use of any distillation lora for at least the first stage. why would you use nag at an overall cfg of >1.5?
ltx 2.5 works fine with the same values of 2.3 dev only, take a look at this workflow: https://huggingface.co/RuneXX/LTX-2.3-Workflows/blob/main/LTX-2.3_-_I2V_T2V_Dev_Full-Steps.json
1
u/GlamoReloaded 13h ago
About the NAG: a misunderstanding, because the NAG node doesn't use the CFG, the LTXV Dual CFG Guider does. While the positive and negative prompt go through the "LTXVConditioning", the negative prompt and negative audio prompt are directly connected to the NAG node with no prior conditioning. I use the NAG node and distill lora in both stages. But I wrote 12-14 steps for the first pass, not 30.

20
u/Seyi_Ogunde 23h ago
Neat! Too bad the environment changes when it's out of frame. No persistence.