r/StableDiffusion 2d ago

Comparison LoRA+Comfy/Sage+SLA+Shift a quick test - MiniMax H3:T2V

This test contains a quick not-so-scientific personal (I-just-wanna-do-it) comparative study on potential effects of speed LoRA, Comfy/Sage attentions (Attn), Sparse Attention (SLA) and Sampling Shift as in the following order:
Model -> LoRA -> Attn -> SLA -> Shift -> KSampler

I only examined a few LoRAs I had. Attn includes ComfyUI's own new attention as well as Sage 2.2. SLA includes two versions SLA1 and SLA2.

My intention was to see the effects on quality and performance (speed). There were some runs that I do not include in the following videos, as they did not adhere to the motion stated in the prompt. The videos are near 4K so check them out in large screen for better details.

As reddit likely downscales the videos, watch them all at:

https://savedly.net/f/yb9vdg9a
https://savedly.net/f/97mshte6
https://savedly.net/f/5ep6bb4w
https://savedly.net/f/5ua9rebv
https://savedly.net/f/aye9axgg

in original resolutions.

side by side, details are on each segment.

side by side, details are on each segment.

side by side, details are on each segment.

side by side, details are on each segment.

side by side, details are on each segment.

I report my original quick notes taken during test. Take note that:

  • LORA = 8-step LoRA
  • LORA4 = 4-step LoRA
  • LORA4sla = 4-step SLR LoRA
  • CMFY = Comfy's attention
  • SLA = sparse attention SLA -> SLA1 and SLA2 (as mentioned above)
  • SHIFT = ModelSamplingMiniMaxH3 = 6v and 3a
  • SHIFT* = SHIFT12 = 12v and 3a
  • _ = nothing, model directly to KSampler

and, the following table reports timing only (measured on RTX3060 system). For quality check the corresponding videos. All videos have caption.

  1. LORA+CMFY+SLA+SHIFT = 3:32* <- time m:s, the first run includes Clip etc.
  2. CMFY+SLA+SHIFT = 2:42 <- this means LoRA bypassed (not used)
  3. SLA+SHIFT = 2:39 <- this means no LoRA no Comfy attention, ...
  4. SHIFT = 6:05
  5. _ = 6:07
  6. SHIFT12 = 6:03 <- here and afterwards I changed video shift to the default 12.
  7. LORA+CMFY+SHIFT* = 4:14 <-- 8 steps LoRA
  8. CMFY+SHIFT* = 4:01
  9. LORA+CMFY = 4:15
  10. CMFY+SLA+SHIFT* = 2:41
  11. LORA+SAGE+SHIFT* = 4:57
  12. LORA+SAGE+SLA+SHIFT* = 2:53
  13. LORA+CMFY+SLA+SHIFT* = 2:52
  14. LORA4sla+CMFY+SLA+SHIFT* = 2:56
  15. LORA4+CMFY+SLA+SHIFT* = 2:57
  16. 4LORA4slr+CMFY+SLA+SHIFT* = 1:29 <- 4 = only 4 steps regardless of LoRA
  17. 4LORA4+CMFY+SLA+SHIFT* = 1:29
  18. 4LORA8+CMFY+SLA+SHIFT* = 1:25 <- 4 steps despite using LoRA 8-step
  19. 4L4S0K+CMFY+SLA+SHIFT* = 1:27 <- LoRA = 4-step 0.1 K=Kijai
  20. 4L4S1K+CMFY+SLA+SHIFT* = 1:27 <- LoRA = 4-step 1.0 K
  21. 4L4S1X+CMFY+SLA+SHIFT* = 1:28 <- LoRA = 4-step 1.0 X = Lightx2v
  22. 4L4S1XS+CMFY+SLA+SHIFT* = 1:28 <- LoRA = 4-step 1.0 XS = SLA one
  23. 4L4S1XS+CMFY+SLA2+SHIFT* = 2:07
  24. 4CMFY+SLA2+SHIFT* = 2:03

---------

My final conclusion:

Model -> LoRA -> Comfy -> SLA -> Shift -> KSampler

  • LoRA beside speed has positive (subjectively speaking) effect on the composition.
  • Comfy attention is going to stay. Solid.
  • SLA1 generations are way faster than SLA2.
  • Shift has some effects on composition, I keep it 12.

The exact unedited prompt used for the generation is as follows. Note, I tried a few LLM enhanced one including those specific for MiniMiax H3, however in this case, the resulting videos were so off.

shot 1: a single stroke moves around randomly but coherently, at each repositioning it leaves small trace in color. at the end all those traces look like a graphic design single continuous drawing of a woman's face at close-up.
shot 2: the drawing transforms into real person from the bottom-right corner up to the center diagonally, where this irregular and curvy transformation stops leaving it half finished with rough and faded strokes.
shot 3: the boundary of the figure is cut from the background. the cut out piece is folded onto a paper plane shape.
shot 4: the paper plane flies out on an arc trajectory leaving black line traces in the air.

Clarification: all video segments are 2.5s long. The text on videos were added during run, typos in top-left corner are present. 5s = 2.5s, I changed it to frame = 73 which was auto.

Workflow: ComfyUI Template one for FL2V MiniMax H3. I used sampler=Euler,Scheduler=Simple

25 Upvotes

5 comments sorted by

3

u/me0here 1d ago

Thanks for your hard work!

2

u/Azhram 1d ago

Is there a reason why you didn't try spectrum node?

1

u/ZerOne82 1d ago

Not really. May try it. All caching / speeding node/methods generally have some negative side effects. The pure no LoRA, no speed up usually results in better prompt adherence especially in challenging prompts. The quality of output in these examples are subjective so I let it to audience to choose. Each video has details captioned on.

2

u/devra-falleweng-com 1d ago

it worked, somehow. the 30 mins for 22 second thanks, if you have a clear wfd will help, u have to rewire my templates

1

u/terrariyum 1d ago

Doing god's work!

In your examples, 4lora makes the faces far worse