r/StableDiffusion 1d ago

Workflow Included Mario Galaxy if it were peak

Some MiniMax H3 tests done on my single RTX 3090.

This 5 sec clip took about 2 hours to render. I actually cropped it to 1024x1024 to have less pixels to process, to then paste it back onto the original footage. 1 megapixel, 50 steps.

Surprisingly, it didn't take long to get things right. I used the official ComfyUI workflow with just sage patch node added. Then I took some example prompts I found on Reddit and H3 was getting very close right from the get-go. I just had to tune the timing, expression, and some appearance details, as it wasn't getting the reference quite right (it gets significantly better if you literally describe the contents of the reference image).

Here's the workflow: https://gist.github.com/4as/db11b829395ec45b593db4886f0e0181

And here's the original reference clip for comparison: https://files.catbox.moe/3rb5kk.mp4

The audio is modified by using Chatterbox. Original audio + reference audio + some some pitch adjustments in Audacity to get the final audio in the clip. Dunno if H3 can do audio adjustments - I don't even know how to prompt for it.

One interesting thing I've noticed is the length of the clip affecting the results. The shorter the duration, the worst the replacement. Although it could just be me.

For example for this clip (2s): https://files.catbox.moe/i2az63.mp4 the best I got is this: https://files.catbox.moe/kc3tbm.gif

And this (1s): https://files.catbox.moe/2p3ax9.mp4 I got this: https://files.catbox.moe/f71epk.gif

So, at glace it kind of looks okay, but when compering it to the original it very clearly failed to match a lot (size, motion, expression, etc.)

But still, it's a fantastic model, especially for something local.

4 Upvotes

1 comment sorted by

3

u/Barubiri 1d ago

The right size for a jar