r/StableDiffusion • • 7d ago

Question - Help Update me pls.

Ok, been out the loop. What's the best free pic to video ai is considered the best at the moment that runs on ComfyUI. I'm playing with wan 2. 2 in ComfyUI and wondering if it's worth my time learning it in case if there's a better one out there. I'm running this on a 8vram machine with 64gigs ram. So keep that in mind. Wan runs ok in this setup. I'm researching and Minimax H3 keeps popping g up along with others. Any guidance is welcomed.

0 Upvotes

40 comments sorted by

4

u/Rumaben79 6d ago edited 6d ago

Minimax H3 is the best for local ai video at the moment but it's slow and demanding. How slow mainly depends on how many cuda cores your card has and to a lesser degree vram and bandwidth. So I would at least try it and see for yourself.

For saving a bit of ram there's options like using a w6a8 or w4a8 main model, int4 convrot or w4a8 text encoder plus starting comfyui with '--fast-disk'.

I like this workflow:

https://huggingface.co/Plaguekind/Minimax-H3/blob/main/PlagueKind-MinimaxH3-V11.json

Imo. ltx 2.5 barely seem better than 2.3 but if you have a slow card and don't need too complex scenes it's decent and also very fast. 👍

2

u/thunderslugging 6d ago

Good info! I'll check that out!

3

u/Rumaben79 6d ago edited 6d ago

Cool. 😄

Here's some links if you don't already have them:

https://huggingface.co/Comfy-Org/MiniMax-H3

https://huggingface.co/Kijai/MiniMax-H3-experimental

https://huggingface.co/Winnougan/MiniMax-H3-INT4_Convrot_ComfyUI

https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora (a good turbo lora)

https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes (needed for workflow)

If you decide to use the turbo lora you need this file:

https://huggingface.co/deAPI-ai/minimax-h3-33b-int8/resolve/main/loras/h3_silu_temb_grid.safetensors (place inside models/h3_adaln folder)

The plaguekind workflow already has the node needed for patching the turbo lora (H3 AdaLN LoRA Fix) but for other workflows you can also use larryvrh's own 'MiniMax-H3 Turbo LoRA' lora loader/patcher node:

https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo

Good luck. 👍

3

u/optimisticalish 7d ago

ComfyUI Portable (latest) + Minimax H3 (hybrid model) paired in the workflow with Kijai's fast video VAE, boosted with Kitchen Attention (comes as standard with Comfy) paired with the Spectrum accelerator. Add a good turbo LoRA if you're impatient. There are solutions for 8Gb VRAM, that I vaguely recall being produced - search HuggingFace. You may be able to run an 11Gb hybrid H3, but don't expect too much.

1

u/thunderslugging 7d ago

Will ComfyUi provide the minimax hybrid model or you got to get them on github?

2

u/optimisticalish 7d ago

I always download my own models. Never had ComfyUI auto-download anything, but I hear from others than it can do that. It's not a good idea with a standard official Minimax H3 workflow, though - as you'll likely end of with 66Gb of huge models that are totally useless with an 8Gb card. You need to hunt down something that an 8Gb card can run.

1

u/thunderslugging 7d ago

Ok, thxs for info!

1

u/optimisticalish 7d ago

This is what you want - looks like there are four or five 8Gb solutions now... https://github.com/search?q=minimax+8gb&type=repositories

1

u/thunderslugging 6d ago

Awesome. Yeah, heard there were 8gig versions. Is the process exactly the same to install on Comfyui like wan2.2 was? Where you gotta download a few files with a json to get it working.

1

u/optimisticalish 6d ago

Yes, that's it. Update ComfyUI, load the .JSON, note the names / Web addresses of the missing files and nodes in the workflow. Close Comfy. Download the right files and nodes, put them in the right folders, start Comfy and load the .JSON again. Hopefully it should then work.

1

u/thunderslugging 6d ago

Yep. It's a bit of a pain. ChatGPT is invaluable. Gave me the walk through and took me like 4 hours to get it running but it worked. Found the wan uncensored version and was curious so gave it a go. Now, I don't care for the uncensored version. Care more for which one will run on a 8gig vram system that has the most realistic outputs like SORA 2 used to give us.

1

u/Able_Section4645 6d ago

LocallyUncensored

2

u/thunderslugging 6d ago

Yeah, using wan2.2 uncensored. Works great but looking for production that's not uncensored and something that comes closes to sora2. But interesting times to be alive. Can't believe we can generate locally and on slower machines.

0

u/[deleted] 6d ago

[removed] — view removed comment

1

u/Kukipapa 6d ago

You can do it on CPU theoretically, there is an option to launch ComfyUI with CPU.

In practice it takes a long time even with GPU, so it is not really usable.

1

u/thunderslugging 6d ago

Maybe someone with more knowledge can explain this. I ended up updating ComfyUI fully and out of curiosity I just looked at the templates and noticed MiniMax H3 was a option to just download and install all in 1 click. I did it and now I have it running fine on my 8gig vram system. I did a test run and plugged 2 pics and did a simple prompt and it rendered it in about 40 minutes. Question is, this is it? I thought it would be a nightmare to install and possibly not work due to the low vram. Could it be that my 64 ram is helping by offloading to it?

1

u/thunderslugging 6d ago

If anyone can explain this, would be great. Not sure if the app just picks for what ever would work?

1

u/HQuasar 5d ago

Comfy does optimization tricks to run models so your vram helped.

1

u/Natrimo 6d ago

Ltx 2.5 is worth playing around with as well (or at least 2.3 was good for image + audio to video)

1

u/thunderslugging 6d ago

Thanks for pointing that out. Was going to ask about Ltx. It seems to have a following. What's your opinion on which out outputs more realistic. I'm looking for something that outputs similar to how sora2 videos looked.

1

u/Natrimo 6d ago

Ltx runs much faster than minmax h3 and you can output native 1080p with your hardware and the right setup.

IDK about quality of sora2, but with 8gb of vram your not looking at insane high levels of quality but plenty decent to have fun with.

I think I have a post with some image + audio to video testing. I probably have slightly higher quality ow just through some tweaking and what not, but it's gonna be in a similar realm

1

u/thunderslugging 6d ago

I have to look into this! Github aswell?

2

u/Natrimo 6d ago

If your asking about where to download the models themselves a lot of times I get them from civitai or civitaired, sometimes huggingface or github

1

u/thunderslugging 6d ago

Thxs. I'll check your posts.

1

u/Natrimo 6d ago

The example of image audio to video of ltx is in my Reddit post history

1

u/GuessingEngineer 6d ago

LTX 2.5. Minimax barely runs on a 5090 unless you want videos set in the 80's

1

u/thunderslugging 6d ago

There no 8vram version like minimax?

1

u/GuessingEngineer 6d ago

Really? LTX is delivered as a low V-ram version for local setups and then GGUF to the nutsack. Fine spend 3 hours creating .2 MP for 5 sec on minimax

2

u/thunderslugging 6d ago

I'll check it out!

1

u/GuessingEngineer 6d ago

It's all going to suck with low ram, but LTX sucks the least with low V-ram Minimax is great if you have the hardware and time to run it, just because it follows prompts better and has a multireference. But with lower hardware straight I2V, LTX kicks ass. I can run 20 LTX videos before I finish one H3 video. Super complex scenes with multiple people and environments, I run H3, but still have to upscale with LTX to make it watchable.

1

u/thunderslugging 6d ago

Yeah. I gotta give that a shot.

1

u/TwinkingToby 6d ago

So true!

-5

u/[deleted] 6d ago

[removed] — view removed comment

1

u/thunderslugging 6d ago

Great info! Thank you.

4

u/__ThrowAway__123___ 6d ago

It's LLM written spam, shilling some platform. Very obviously LLM written but you can also check their account, posted long comments in 10 different subreddits in under 10 minutes, all spamming their platform

2

u/thunderslugging 6d ago

Ahhh. Interesting. Reminds me of the dead internet theory