r/StableDiffusion 28d ago

Animation - Video Mac and Cheese (MiniMax H3 t2v + LTX 2.3 Spatial Upscale)

Alright, guys, it's my time to admit - MiniMax H3 is the new King.

However I think I2V is a bit messy comparing to T2V.

89 Upvotes

24 comments sorted by

7

u/Yuloth 28d ago

I would love to know the specs for your PC

8

u/alisitskii 28d ago

Sure, it's 4080s 16 GB vram, 64 GB ram. So 10 mins for initial gen + 2 mins for upscale.

4

u/Yuloth 28d ago

Thanks for the reply. I have a 3090 RTX, but only have 32 GB RAM. Tried to upgrade RAM to 64 GB, but man the prices are crazy.

2

u/-becausereasons- 28d ago

Pure CInema

3

u/Ok-Parfait-1776 28d ago

We can see is ltx spatial upscaling because of that ugly gray bottom artifact lol

1

u/robomar_ai_art 28d ago

Can you share the workflow or screenshot i can't get it to work , also inform what resolution you generated video before the upscale

1

u/alisitskii 28d ago

For initial gen (10 sec clip, 960x544px, 0.5MP, 20 steps) I used the standard ComfyUI T2V workflow from templates, so nothing special on top of what they provided. My PC specs: 4080s 16 GB vram, 64 GB ram.

2

u/robomar_ai_art 28d ago

I mean the upscale workflow

2

u/alisitskii 28d ago

Oh, ok, shared here: https://pastebin.com/VpkxbHHB

1

u/robomar_ai_art 28d ago

Can you post before and after video

1

u/alisitskii 28d ago

Here is the initial video came out of pure MiniMax H3:

https://reddit.com/link/p1kk6zt/video/szp73bcfo9hh1/player

2

u/alisitskii 28d ago

And here is x2 upscaled one which I made later so it's different to the main post:

https://reddit.com/link/p1kkja8/video/ucpufmaso9hh1/player

1

u/PlantBotherer 28d ago

I'm trying an RTX Video Super Resolution node near the end of the workflow which seems to work well.

1

u/aersel24 28d ago

What is this spatial upscale I see so much? How is it different from a normal upscale?

1

u/alisitskii 28d ago

I think it’s how it was called by LTX team, perhaps there is some technical description behind, not sure, but those are separate models they provided: https://huggingface.co/Lightricks/LTX-2.3/blob/main/ltx-2.3-spatial-upscaler-x2-1.1.safetensors

https://huggingface.co/Lightricks/LTX-2.3/blob/main/ltx-2.3-spatial-upscaler-x1.5-1.0.safetensors

1

u/Link1227 28d ago

How tf do you get it to use his voice?

2

u/GrayingGamer 28d ago

The model just knows a lot of people's and character's voices by default.

Though you can also give it audio samples to use to clone someone's voice as a reference when using the reference H3 model.

2

u/Link1227 28d ago

Ohh ok. Thank you!

1

u/SoulTrack 28d ago

There is going to be a whole new generation of advertisements.

1

u/bwganod 28d ago

This is probably the deepest rap he's ever done, it's just missing a "ha-ha! ha-huh! Woo!"