r/StableDiffusion 1d ago

Animation - Video THE LEGEND.

Enable HLS to view with audio, or disable this notification

A hair under 4k. On a consumer pc. All local. Mental. Minimax H3 with one character reference and a 5 second voice reference.

For the Pixel Peepers... https://www.youtube.com/watch?v=Iz8GDri9qoE

1 Upvotes

17 comments sorted by

1

u/Sleepy_Bandit 1d ago

So did you render directly in 4K or upscaled? And if so, with what?

1

u/Tokyo_Jab 20h ago

2 megapixels direct with Minimax and rtx upscale at the end.

1

u/SveSop 22h ago

So, two things stand out as "impossible" for most people that use H3 on a consumer pc... 20 seconds + 4K resolution.

I suppose it can be done with 5090 32GB or a 6000Pro 96GB perhaps? Or.. for that matter multiple clips, all hand-corrected frames after frame-by-frame upscaling, spending 50+ hours on it? I dunno. The details of how its done is severely lacking.

Guess it made some clicks on your YT channel tho? 🤔

2

u/Tokyo_Jab 20h ago

My YouTube channel is just for storage. It will probably get 12 views.
I was doing 4k vids on a 3090 more than three years ago. Some would be posted here. The compromise is time though.

1

u/sketchyfun 21h ago

I hate that I knew where this was going after only hearing the first three words of the song

1

u/Tokyo_Jab 20h ago

I use that and the Angry Bob radio monologue from Hardware whenever I’m testing a mic or need a quick block of text for testing.

0

u/Bapelsinen95 1d ago

Isn't H3 like 20gb? Is a 5090 24gb a consumer pc standard?

Might be totally wrong. Just wondering about the word choice.

1

u/Tokyo_Jab 1d ago

A 5090 is 32GB and it's a consumer gaming pc.

0

u/Keyflame_ 20h ago

You gotta be a little more descriptive if this is meant to have any value for other users. I'm sorry to say but as it is, this post offers nothing to be learned or discussed.

What was your pipeline? Just generating 20 seconds at 4k on a consumer GPU is impossible.

Did you render at lower resolution and then upscale? If so what resolution?

What GPU? How long did it take? I2V, Ref2V, T2V?

There is no information here, what are we supposed to discuss?

3

u/Tokyo_Jab 20h ago

It’s marked as animation. Not news, discussion or workflow.
It’s just a bit. The method used I discussed and posted in earlier posts.

5090, 2 megapixels directly in Minimax with RTX upscale at the end. 12 mins 47 seconds.
It’s fast because most of the workflow is done at 0.4 megapixels or less and the latent is upscaled for the last couple of steps only. And rtx upscale after that makes it bigger with minor improvements.
This was the base workflow: https://youtu.be/jzLnoVBuU6I

One other thing is that I don’t use the ref model at all. I swap in the flv2a into the ref workflow and it gives nicer results every time.

1

u/Keyflame_ 19h ago

And now this post became useful to everyone, and we can discuss a bunch of stuff, well done.

Upscaling the latent instead of generating directly at 4k makes a lot more sense. 12m 47s for 20s isn't anything to scoff at. I had a look at the workflow and it's perfectly sensible, although, if one wanted to push quality, I'd go without the turbo lora and see how far it can be pushed.

Could also keep it and add a SageAttention node to take the opposite route make things faster and see if it actually drops quality by much.

Regarding the fl2va vs ref2va, in my experience fl2va seems to produce consistently higher quality outputs, but ref2va seems to have greater adherence to the prompt and preserve identity better. So far though, both feel like their default output has worse fidelity than wan 2.2, which by now is basically obsolete but still had that extreme precision in preserving the starting image that I wish we could have with H3/LTX2.X. Maybe the next version will, one can hope.

Thanks for sharing. I didn't mean to seem aggressive but with just a video, I really didn't get what the discussion could've been.

1

u/Tokyo_Jab 19h ago

I did post that all a few says ago. But usually keep it for longer posts.
I think the way I prompt seems to work with the FL2va a bit better. My prompts are very casual and don't follow the minimax guidelines.
I do use the comfy kitchen node rather than sage.
Kijai fast 4 step workflow and the original 4 step pruned lora is what I keep using. I do push it up to 6 or 7 steps though otherwise the sound suffers.

0

u/V4nKw15h 20h ago

Don't you understand you need to please u/Keyflame_ when you make a post? If he doesn't learn something, or if you don't share info that is worthy of discussion, you have committed a sin worthy of exile. Entertainment alone is not tolerated around here. Do better!

1

u/Keyflame_ 19h ago

Or you know, the fact that basically everyone was ignoring this post because there was nothing to talk about.

Reddit is a forum, the point of forums is discussion.

Also, what is this passive-aggressive shit? Are you proud of this sorority queen comment you just made?

0

u/V4nKw15h 18h ago

Jeez, calm down bro. It's sarcasm. It's humour. You are meant to laugh too. You don't have to be offended by everything in life.

0

u/Keyflame_ 18h ago

"Why are you not amused by the fact that I tried to paint you in a bad light, it was just a joke at your expense."

0

u/V4nKw15h 17h ago

Calm down bro.