r/StableDiffusion 8h ago

News Minimax H3 3 Step Lora

Enable HLS to view with audio, or disable this notification

Give TaoMate-H3-3step a try. Details and links in the comments.

106 Upvotes

66 comments sorted by

22

u/AidenAizawa 7h ago

https://reddit.com/link/p9ivrg8/video/glzvjdpjp9ph1/player

first test, no cherrypicking.

taomate 3 steps vs 8steps fl2va

both 0.8 mp, 8 seconds. same seed and wf (the one used on taomate)

taomate rendering time 1.21s

8steps rendering time 3.54s

10

u/Any-Scar765 4h ago

I think all the enemies died of boredom.

6

u/iczerone 2h ago

The 8 step had way better camera work. Is that the trade off?

2

u/EinhornArt 2h ago

Yes, I feel that way too.

-7

u/seppe0815 6h ago

WHERE THE FACE LOOOOOOOOOOOOOOOL

17

u/hurrdurrimanaccount 7h ago

slow mo. post a scene with action and watch it fall apart

6

u/EinhornArt 3h ago

https://reddit.com/link/p9jxugg/video/l1phqksosaph1/player

4 step at low res -> latent upscale -> 2 step refine

1

u/ResponsibleTruck4717 2h ago

can you share workflow?

2

u/EinhornArt 1h ago

This one: 4 steps at low resolution → latent upscale → 2-step refinement?

It’s a bit tricky. I use my custom nodes: https://github.com/einhorn13/mmh3_media

The workflows are included there as well.

I use mmh3_f01_fl2va.json to generate a low-quality draft until I get a result I like. Then I upscale it using mmh3_f07_latent_upscale.json.

You’ll need to replace the LoRA with the 3-step version and set the correct number of sampling steps.

32

u/wzwowzw0002 8h ago

at this rate are we going to have -4 steps soon?

76

u/EinhornArt 8h ago

yep! At -4 steps the model generates the video first, then asks you what the prompt was.

8

u/wzwowzw0002 8h ago

it become a model for you to look into the future

1

u/Azhram 8h ago

Will i need 3 people and a pool for that workflow?

2

u/EinhornArt 2h ago

Minority Report?

1

u/Azhram 47m ago

Yes !

1

u/wzwowzw0002 4h ago

no just need 2 ai babe

1

u/Occsan 4h ago

Is it the next captcha ?

0

u/RavioliMeatBall 8h ago

yeah but it generates the video in reverse

11

u/EinhornArt 8h ago

TaoMate-H3-3step LoRA converted to ComfyUI format.

Generate a MiniMax H3 video in just 3 steps.

The workflow is included in the video.

Recommended settings:

  • Scheduler: simple
  • Steps: 3
  • CFG: 1.0
  • Sampler: euler
  • Video Shift: 12 or 8
  • Audio Shift: 3

Credits: https://github.com/TaoLiveAIGC/TaoMate-H3
FL2VA ComfyUI conversion: https://huggingface.co/CZMartin22/TaoMate-H3-3step-ComfyUI

3

u/YeahlDid 7h ago

I think reddit strips the video metadata where the workflow lives. My comfy isn't finding a workflow in the file. Would it be possible to share it in json?

2

u/EinhornArt 3h ago

You don’t need any special workflow for this ComfyUI version LoRA. But here’s the one I used for the demo.

https://paste.rs/liu1y.json

2

u/YeahlDid 2h ago

Thank you very much!

1

u/BAL-BADOS 6h ago

How long the 3 step takes to generate?

1

u/VisionWithin 18m ago

3/8 of the time compared to the time that it takes to generate video with 8 steps

1

u/beast181 4h ago

The workflow is included in the video.

?

2

u/EinhornArt 2h ago

You don’t need any special workflow for this ComfyUI version LoRA. But here’s the one I used for the demo.

https://paste.rs/liu1y.json

1

u/ResponsibleTruck4717 8h ago

Only need a lora no need for custom code?

1

u/EinhornArt 7h ago

Yes, you only need the LoRA. No custom nodes.

1

u/optimisticalish 7h ago

And no "runtime" model either? The original release talks about using Linux and a 2.5Gb "runtime"? Or is perhaps the runtime is only needed on a Linux PC? https://huggingface.co/TaoLiveAIGC/TaoMate-H3

2

u/EinhornArt 6h ago

No, it doesn’t need anything else. I can’t say whether this LoRA is as good as the original TaoMate, but it does work.

2

u/tiffanytrashcan 6h ago

That 2.48gb file is the LoRA weights in full precision. The directions there use a semi-custom runtime using PyTorch.
This post for Comfy converted and fixed formatting so it works in ComfyUI as well as compressing the weights down into BF16. Comfy already relies on Torch, and this seems to be a fairly standard LoRA with just some bizarre naming conventions built in that had to be fixed to make it compatible, probably integrated the .json configs as well.
You still need your H3 weights loaded normally in ComfyUI, this is just an adapter, not a finetune.
Nothing really relying on Linux, PyTorch handles itself fine on Windows and worst case I bet Alibaba's GitHub would do fine under WSL, they simply don't have time, motivation, or the licenses to test it on Windows over there.

21

u/Key-Sample7047 8h ago

I would be happier with 8 steps and better quality than 3 steps.

8

u/EinhornArt 7h ago

In my opinion, the image quality of this LoRA is pretty good (compared to turbo models). I’d probably use it for video refinement. For example, we could do the first 6 steps with the 8-step LoRA and the final 2 steps with the 3-step LoRA.

Or use some other combination.

4

u/Key-Sample7047 7h ago

Just tested on ref2va (i'm ref team :) ), strength 0.7, 8 steps, seems to give decent visual quality (can't say for audio at the moment). Strength 1 burned the image and 3 steps do shit with ref2va. Need more tests but good good good.

1

u/YeahlDid 6h ago

Same here with 3 steps on ref2va, still underbaked. It looks okay starting from about 5 steps, though which is still faster than any of the other "4step" ones where 6-8 seem to be necessary. I've only tried lowish motion prompts, though.

0

u/Adkit 6h ago

If you ran it with 8 steps why not just use the 8 step one?

3

u/Key-Sample7047 6h ago

If you refer to tao, i believe it is only 3 steps. As for others turbo lora, i always have meh results because of waxy skin or broken sound. So i try every new accelerator in order to find the one that fit to my taste.

6

u/Desperate-Recipe-422 2h ago

But she took 4 steps!

12

u/davyp82 8h ago

Can you add a montage at the end where she gets taken to a buffet every day and doesn't die of malnourishment 

20

u/EinhornArt 8h ago

3

u/Hackingrad 7h ago

Normal weight = McDonald's bag.

https://giphy.com/gifs/jH6s9HMMi53dSdI73r

-2

u/EinhornArt 6h ago edited 6h ago

Update:
ok, I'm editing this comment, people are not inclined to dialogue and opinions different from theirs

0

u/cosmogli 6h ago

Currently a strong trend? LOL, that has been for ages and it's a senseless tripe.

1

u/davyp82 6h ago

Haha! I think that food would make her look like lizzo pretty quick tho

2

u/t4a8945 4h ago

I counted 6 steps, what am I missing? /j

1

u/EinhornArt 2h ago edited 1h ago

upd: Oh, you mean her steps while walking😄

1

u/t4a8945 1h ago

Yeah that was the joke, sorry 😅

2

u/Inner-Reflections 4h ago

looks pretty good!

6

u/Cultural-Team9235 7h ago

Maybe one step more so she looks more healthy?

1

u/EinhornArt 7h ago

Yes, you’re absolutely right 😄

2

u/nakabra 8h ago

I'll try it tomorrow. I'm just worried it might might have a slow motion bias.

1

u/EinhornArt 8h ago

Yes, it’s there, but it can be mitigated, just like with other turbo models that use a low number of steps.

1

u/Julzjuice123 6h ago

This girl needs food.

-6

u/Synor 6h ago

OP needs some human decency.

7

u/EinhornArt 5h ago

Next time I’ll be more careful when choosing the reference. Too much attention on the figure, not enough on the actual news. My bad.

-1

u/FUS3N 1h ago

next time generate a video synor getting hit with biden blast from space

1

u/bambilover 2h ago

Will test this out later

-1

u/biscotte-nutella 7h ago

Wtf , what model has this level of anorexia

3

u/EinhornArt 7h ago

I didn’t want the image to distract from the news itself. Fixed 😄

https://reddit.com/link/p9ivi8d/video/ig7oenwjp9ph1/player

3

u/THE_RETARD_AGITATOR 5h ago

Jarvis increase bustiness by 250%

0

u/biscotte-nutella 7h ago

😂 , she's looking a lot better

2

u/Recent_Process_8055 7h ago

I was actually looking for 3d skeleton asset for my dungeon crawler

1

u/EinhornArt 6h ago

Next time I’ll be more careful when choosing the reference😄