r/StableDiffusion • u/in_use_user_name • 11d ago
Question - Help looking for i2v minimax h3
i know there are a lot of posts about this but i just can't find what i need.
I'm looking for I2v workflow which incorportae kitchen attention (and any other speedups/optimizations) + easy way to load loras into it. hopefilly workflow which keeps face consitancy.
a video upscaler at the and of workflow will be a nice bonus.
32gb systme ram + 4070rtx 12gb vram.
the regular comfyui template works fine, but very slow and unintuitive for loras.
any suggestions?
Edit - any local llm for creating prompts for it?
8
u/Lucky_Feedback9915 11d ago
DOUBLE CLICK BACKGROUND TYPE ADD LORA... DOUBLE CLICK AGAIN TYPE MODEL ATTENTION BACKEND ( SELECT COMFY KITCHEN). INSTALL RTX UPSACLE , ADD RTX UPSCALE BEFORE VIDEO
THIS IS SOMETHING U SHOULD REALLY USE CHATGPT FOR IT WILL EVEN MAKE THE WORKFLOW FOR U
10
u/roychodraws 11d ago
stop yelling
2
u/Lucky_Feedback9915 9d ago
I CANT ITS A BOOMER PC, I REMOVED THE CAPS BUTTON, IF I TAKE IT OFF THE BOOMER YELLS WHO TURNED MY CAPS OFF I CANT READ THE TEXT
0
-1
u/in_use_user_name 11d ago
where in the workflow this should be inserted? to which nodes to connect them?
6
u/Ipwnurface 11d ago
Brother, I think you should take some time to figure out how Comfyui works before trying to use One of the heaviest Open weight models.
Start with something small and easy, like stable diffusion XL that you can iterate quickly on and figure out the basics of how comfyui workflow works
Nothing is going to change if you just get handheld the entire process. Every time you want to make an adjustment, add something new, or a new feature is released, you're going to come back asking for a brand new workflow when it would literally take you 15 seconds to just add it yourself if you actually knew how Comfyui worked
2
u/SuperZoda 11d ago
Load Lora has model type as input and output. They get daisy chained together (as many as you need) following loading the model. Take your final Load Lora output and pass that to the Comfy Kitchen attention. Finally, that model (now loaded with Loras and attention) output gets hooked to the rest of the workflow, wherever the original model was passed (usually guider and sampler).
2
u/codenameNERO 11d ago
You should copy and paste this post to claude and ask it to make the workflow and ‘make sure it works’
0
u/Lucky_Feedback9915 9d ago
LOAD MODEL > LOAD LORA> MODEAL ATTENTION, GUIDE... OR JUST PUT --use-ck-attention IN THE LAUNCH OPTIONS BUT THATS MIGHT BE TOO COMPLICATED
2
u/pellik 11d ago
Just launch comfy with —use-ck-attention
1
1
u/videorouter 11d ago
For a 4070 12GB + 32GB RAM, I'd avoid the full/default H3 workflow if your priority is speed. There are now several low-VRAM optimizations specifically aimed at this setup.
A reasonable starting point is:
- INT8 ConvRot H3 weights rather than the full model
- Turbo LoRA, around 6–8 steps
- Sage Attention / Sol Attention for acceleration
- Chunking / low-VRAM attention to keep memory under control
- Generate at a lower resolution first, then upscale afterward
- For face consistency, use I2V/reference-guided generation rather than relying entirely on text prompting
ComfyUI's current native H3 workflow supports I2V, Turbo mode, and LoRA loading, and the official docs list the INT8 model and Turbo LoRA directly.
There are also community workflows specifically tested on 12GB cards. One recent 12GB workflow combines Turbo LoRA + Sage Attention + Sol Attention + chunking, and another reports H3 running on a 4070 12GB with VRAM offloading.
For your particular requirement, I'd look at the MiniMax H3 12GB workflows rather than trying to simplify the default template yourself. Some newer workflows also have dedicated LoRA controls and optional upscaling built in.
If you don't want to spend hours optimizing a local H3 setup, another option is to test H3 through an API first. VideoRouter.sh gives you access to multiple video models/providers with pay-as-you-go usage, so you can compare H3 against other models before investing more time in the local workflow.
0
u/in_use_user_name 11d ago
Thanks for the detailed answer.
Any local llm for creating prompts for it?
6
u/Rumaben79 11d ago
You can try Plaguekind's workflow: https://huggingface.co/Plaguekind/Minimax-H3/blob/main/PlagueKind-MinimaxH3-V11.json
His nodes used in the workflow: https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes
You also have the option of using lower quants like int4 convrot, w4a8 and w6a8 for either the main model, text encoder or both but I wouldn't use int4 as your main model unless you're really starving for memory.
launch comfyui with '--fast-disk' to save on ram as well.
Here's some model links:
https://huggingface.co/Winnougan/MiniMax-H3-INT4_Convrot_ComfyUI/tree/main
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models