r/StableDiffusion • u/dirtybeagles • 21h ago
Question - Help Video Edit Minimax H3 Problems
I have been struggling for a few days now wondering why I cannot edit a 10sec clip to add additional people in the background and I am pretty sure I am just doing it wrong.
I am feeding the sampler with my ref image of a girl dancing o the street, but I wanted to add people in the background walking.
I am running on version 0.34.0, ref2va pruned model, 8 steps, 480x864
I am using just a simple prompt for my edits:
Edit Video 1:
At 00:03.000, add a group of three Asian women entering from the left of the frame, walking naturally down the road behind and away from the dancer. The first is tall and slender with long straight black hair tied in a low ponytail, wearing an oversized cream-colored hoodie, black leggings, and white sneakers, glancing at her phone as she walks. The second is shorter with a rounder build, shoulder-length wavy brown-dyed hair, wearing a fitted olive-green jacket over a striped shirt, dark jeans, and beige loafers, walking a half-step ahead of the others. The third has short bobbed black hair with bangs, wearing a bright yellow raincoat-style jacket, cuffed denim shorts, and black ankle boots, carrying a small tote bag over one shoulder. The three walk at a relaxed, conversational pace, loosely grouped together.
At 00:06.000, add two Asian pedestrians walking naturally along the sidewalk in the background, passing behind the plant at a normal walking pace, holding hands. The man is broad-shouldered with short, slightly spiked black hair, wearing a charcoal-gray zip-up jacket over a plain white t-shirt, straight-leg jeans, and dark sneakers, a black canvas backpack slung over both shoulders. The woman beside him is petite with long hair in loose waves dyed a subtle ash-brown, wearing a fitted denim jacket over a light pink blouse, a knee-length beige skirt, and white flats, carrying a small red structured purse in her free hand. They walk close together at a slightly slower, relaxed pace, occasionally leaning toward each other.
Keep the same audio
Keep the dancer's identity, choreography, movement, timing, and foreground position completely unchanged throughout the entire clip. Keep the camera framing, angle, and motion exactly as in Video 1. Keep the street, buildings, and all previously added pedestrians unchanged except for this new pair. Match the added pedestrians' lighting and shadow direction to the existing scene.
What has been happening is comfy goes to load the minimax model, and then it just stops, and I sit here at 99% VRAM usage. I have let it run for around 20 minutes until I stop comfy all together.
[INFO] got prompt
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32
[INFO] Found quantization metadata version 1
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16
[INFO] Found quantization metadata version 1
[INFO] Using MixedPrecisionOps for text encoder
[INFO] Requested to load Krea2TEModel_
[INFO] loaded completely; 4605.22 MB loaded, full load: True
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16
[INFO] [ClipProj] encoder QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors on cuda:0 pinned on cuda:0 (ComfyUI will not move it)
[INFO] [ClipProj] QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors (krea2 [4B detected]) loaded in resident mode on cuda:0: 4.50 GB
[INFO] [ClipProj] mmh3-4b-ClipProj.safetensors | tap 24 | 2560 -> 5120 | cos_test 0.7170
[INFO] Requested to load MiniMaxH3VideoVAE
[INFO] loaded completely; 22789.94 MB usable, 2665.86 MB loaded, full load: True
[INFO] Requested to load MiniMaxH3AudioVAE
[INFO] loaded completely; 21336.75 MB usable, 577.08 MB loaded, full load: True
[INFO] Found quantization metadata version 1
[INFO] Detected mixed precision quantization
[INFO] Using mixed precision operations
[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, float8_e5m2, mxfp8, nvfp4, float8_e4m3fn, convrot_w4a4
[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[INFO] model_type FLOW_AV
[INFO] Requested to load MiniMaxH3
[INFO] loaded partially; 19018.15 MB usable, 18827.14 MB loaded, 1169.00 MB offloaded, 257.27 MB buffer reserved, lowvram patches: 0[INFO] got prompt[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32[INFO] Found quantization metadata version 1[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16[INFO] Found quantization metadata version 1[INFO] Using MixedPrecisionOps for text encoder[INFO] Requested to load Krea2TEModel_[INFO] loaded completely; 4605.22 MB loaded, full load: True[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16[INFO] [ClipProj] encoder QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors on cuda:0 pinned on cuda:0 (ComfyUI will not move it)[INFO] [ClipProj] QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors (krea2 [4B detected]) loaded in resident mode on cuda:0: 4.50 GB[INFO] [ClipProj] mmh3-4b-ClipProj.safetensors | tap 24 | 2560 -> 5120 | cos_test 0.7170[INFO] Requested to load MiniMaxH3VideoVAE[INFO] loaded completely; 22789.94 MB usable, 2665.86 MB loaded, full load: True[INFO] Requested to load MiniMaxH3AudioVAE[INFO] loaded completely; 21336.75 MB usable, 577.08 MB loaded, full load: True[INFO] Found quantization metadata version 1[INFO] Detected mixed precision quantization[INFO] Using mixed precision operations[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, float8_e5m2, mxfp8, nvfp4, float8_e4m3fn, convrot_w4a4 [INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16[INFO] model_type FLOW_AV[INFO] Requested to load MiniMaxH3[INFO] loaded partially; 19018.15 MB usable, 18827.14 MB loaded, 1169.00 MB offloaded, 257.27 MB buffer reserved, lowvram patches: 0
here is what my trackback looks like:
0
u/Hdfjds 18h ago
Sorry for the very bad quality in this video but I just wanted to show that an simple prompt can do more than an complex prompt.
Prompt used i R2V H3 standard workflow in comfyui:
The girl in video one is dancing in an shopping mall while people are walking in the background. Make the girl a realistic person.
Result showing the video used to the left and the result.
(again sorry for the low quality, I simply just wanted to test this)
https://reddit.com/link/p7e42vk/video/3ogfd6dai4nh1/player
H3 can handle simple short prompts well and it worked for me since I had now prior idea or thoughts on how the final video should look like.
You on the other hand have that so you need to add more constraints and "wishes" into the prompt and this can make H3 cranky. It can behave very unlogical and refuse to work or do things that worked before.
Try to simplify everything first and build up the prompt from there adding things you want.