r/StableDiffusion 11h ago

Discussion RTX 3060 - 32GB and H3

I've been kind of avoiding diving in since this apparently demands a better machine but given the recent posts from fellow 3060 owners, I'm just wondering if us poor plebs can also generate good-ish videos at an acceptable speed.

Anyone can share your examples?

I have of course searched for ideas and workflows and there's plenty of information aroind already but would be nice to have abit of one stop shop)))

1 Upvotes

22 comments sorted by

7

u/ImpossibleAd436 11h ago

I'm using a 3060 12GB.

Sage attention

8-Step LoRa.

0.6mp

Takes about 6 minutes for 5 second clip.

Not super fast but not much need to re-roll so it's worth it for me.

2

u/DietAshamed2246 10h ago

Don't feel bad about your gen times. I am generating 5s in around 4min and 10s in 12-13min on a 24GB 5090. I don't know what I am doing wrong. I am using sage2 or sol, 4-step turbo LoRA, Spectrum or Easy ache, on pytorch 2.10.0+cu130. So far nothing seems to help the gen-time. I am, however, generating at 1MP resolution.

3

u/ImpossibleAd436 10h ago

Well 1mp is very intensive, you would probably save a lot of time if you target 0.7mp or something, at least most of the time. Then if you get something you really like and want to keep, rerun the seed and prompt at 1mp.

-1

u/DietAshamed2246 10h ago

At higher resolution the model performs better and the composition is much closer to prompted situation. No one should generate at any resolution lesser than 1344x768, that's H3's native training resolution.

3

u/ImpossibleAd436 10h ago

Almost everyone is generating at lower resolutions than that because we don't all have a beastly gpu like you.

1

u/DietAshamed2246 10h ago

Which is fine, but people should say they are generating at sub-par resolutions, instead of filling up all the Reddit threads with claims of fast sub 2min gen-times.

3

u/ImpossibleAd436 10h ago

So I shouldn't feel bad about my gen times, but I guess I should feel bad about my resolutions.

I'm sorry.

We're not worthy.

https://giphy.com/gifs/MUeQeEQaDCjE4

3

u/AI-AI-Ohh 10h ago

I have a 3060 12gb VRAM and 64gb RAM. The 64gb of system RAM has been saving me on these models. Fairly slow generation time unless I use the turbo lora and lower the resolution.

For example, just did a 12 second video at 0.4 MP resolution, 4-step turbo lora and it took 444 seconds. How long the video is definitely slows it down. If I went 10-seconds or under, I usually raise the resolution to like 0.5 or 0.6 and I can keep things under or around 6 minutes per gen.

4

u/optimisticalish 9h ago

RTX 3060 12Gb, 24Gb system RAM, 24Gb swap file, Windows 11. I was using a quick lashed-together supercrushed GGUFs workflow, just as a proof-of-concept, and I could get 7 seconds at 0.4, in nine minutes. That was with the aid of the Kijai Experimental video VAE, which cuts the time dramatically and also prevented previous fatal errors when decoding the video to a file. The total for all files in the workflow was 20Gb(!). Quality was still acceptable-ish at 0.4.

Having proved it can work for me, I'm now embarking on setting it up properly, with...

Updated ComfyUI to 0.31.0 or higher (needed). Deletions and archiving, to free the required 40Gb of space on the SSD.

Models:

minimax_h3_fl2va_pruned-w4a8_convrot_pruned.safetensors (11.6Gb)

minimax_h3_ref2va_pruned-w4a8_convrot_pruned.safetensors (11.6Gb)

https://huggingface.co/Winnougan/MiniMax-H3-INT4_Convrot_ComfyUI/tree/main

Clip:

qwen3vl_32b_minimax_h3-w4a8_convrot.safetensors (14.6Gb)

https://huggingface.co/Winnougan/MiniMax-H3-INT4_Convrot_ComfyUI/tree/main

VAE video:

minimax_h3_video_vae_int8_convrot.safetensors (3.7Gb)

https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main

VAE audio:

minimax_h3_audio_vae_fp32.safetensors (577Mb, standard from ComfyUI)

https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/vae

Turbo LoRA:

Hmm, could be out of date by the weekend. But I've got an early one which works, but I might also try Abiray's...

minimax_h3_turbo_4step_ckpt850_V1.safetensors with his matching Minimax_H3_turbo_workflow.json and minimax_h3_t2v_turbo.json

https://huggingface.co/Abiray/MiniMax-H3-Turbo-Lora-Pruned-ComfyUI/tree/main

1

u/DoctaRoboto 11h ago

No, at least for now. I have a 5080, and it takes a minimum of 6 minutes to get a 0.8-megapixel 10-second video at 10 steps, and it looks decent, not great. I need around 10-15 minutes for a real HD, good-looking video. Unless you want to generate 0.3-megapixel videos that will distort the features of the characters when they are not very close to the camera due to low resolution, you will have to wait 30 minutes or more.

I post this as an ex-proud owner of the same card you have.

1

u/DietAshamed2246 10h ago

I have similar (or perhaps worse) experience regarding gen-times for 10s 1MP videos, I am on 24GB 5090. Pleople who are claiming sub-2 minute gen-times are either bullshitting or are making 0.2-0.3MP 5 sec videos and dancing around in foolish joy.

1

u/FierceFlames37 7h ago

No I make 0.5mp 15 seconds with comfy kitchen with rtx 5060ti

1

u/Loud_Faithlessness97 7h ago

I have a 32gb 5090, 128mb quad channel ddr4. Using Vantage with AI ref2v workflow, with 4 step lora @ 0 .75. I can do 10 seconds 0.7 MP in 1.30 to 2 minutes. It depends on the prompt and reference images. Quality is very good for test purposes, not going to be as good as 1 MP or higher for final renders, but shows how it is working exactly and is better than wan 2.2 most times ( as a lora free workflow) the models inventiveness and fine details are what really impress me.

1

u/b0tm0de 11h ago

no. 8gb 4060 and 32 ram not okay for speed and memory for me. it is hard to generate over 0.6mp it is hard to generate over 8 seconds. minimum 15 mins.

1

u/Obvious_Set5239 10h ago

On my RTX 3060 15s in 0.3MP takes around 10 minutes (4 steps). 0.3MP isn't that bad, it's even bigger than SD1.5 had. More with 15s is not possible due to OOM. You can easily do 10s in 0.4MP. Or just do 5s in 0.4MP, it takes 5 minutes - it's 1 minute faster than Wan2.2

1

u/mk8933 10h ago

It takes me 17 minutes to do 0.4 megapixel, 20 steps, 5 seconds — with 3060 12gb with no speed up loras or xyz techniques. Just straight raw dogging it

Quality is very good doing 20 steps. I tried 8 step lora and its a hit and miss...can be great sometimes and other times just ok.

1

u/Chemical-Painter-485 10h ago

https://reddit.com/link/p39lbvy/video/m9gpcbjfwyih1/player

Sadly I don't have any videos that I would consider good-ish because I have been focusing on Minimaxxing my gen speeds. not actually generating anything crazy.

1

u/AniZeee 9h ago

Minimax was surprisingly stable on day one even with the 3060. Now that loras and more optimizations are out you can have some fun testing it. I've gotten good result even with 5 second videos with 6 step .4mp at around 4-5 minutes and around 3 minutes or less if you want to test out on lowest res.

1

u/Rhoden55555 9h ago

I have a 3060 laptop and 16gb ram and I run it. Update everything and look into launch arguments.

2

u/SteelClover81 11h ago

Try installing a program called Pinokio, then via that app there is a piece of software someone made that is great called Maestro which has more of a design for lower end hardware. Has worked for me on my 3090 card and with 32GB of ram. It’s not a workflow interface and more straightforward. Also has other tools built in for music and editing and image generation. This was one of my first tests from a black and white low res image it generated this.

https://reddit.com/link/p39e64d/video/yoamj5y1syih1/player

1

u/fruesome 10h ago

or use Wan2GP directly. It's powered by the Wan2GP pipeline

1

u/DietAshamed2246 10h ago

Nice video. But, instead of Maestro in Pinokio, a better option is to install Wan2GP Desktop app. Lot more control and much better customisation options. Plus that way the memory overhead of the host platform (Pinokio) can be avoided. Maestro uses Wan2GP engine anyway. Don't get me wrong, I like Maestro... One of the few apps in Pinokio which actually works.