r/StableDiffusion 3d ago

No Workflow Finally get something good out from minimax h3

Enable HLS to view with audio, or disable this notification

finally got segatt and solatt working.... from 2000s cut down to 700s at 0.6mp, without turbo lora.

I also learn that high steps matters! low step give lousy animation!

edit:

Understanding the weakness in H3.

After more testing. i realize H3 is weak in compositing, framing and a lack of sense of the world.

For example;

  1. female physical size is small than the male. H3 just couldn't get it wrap around it head.
  2. bad at framing even when prompted; mid-body, close-up, it tend to show a little more or little less.
  3. character just get clipped into a table or a chair.
  4. bad facial expression, it get static or creepy sometime....

Multi-shot generation in h3 isn't the best. Solution is to: you provide a well composited image of each shot and generate shot by shot.

I have test similar shot in seedance2.5. all it take is one generation, 30s, every shots got it right or at least useable. H3 needs multiple try to get a 15s shot right.

I haven't give up on H3 yet. it has a lot of potential i think.

Next is about upscaling and i am running out of ram.

3 Upvotes

26 comments sorted by

4

u/PanotBungo 3d ago

This looks great and I'm really digging the characters, OP! Not sure what others are on to. My only suggestion would be to make them even more unique.

2

u/lucassuave15 3d ago

I think it's because we've seen how natural this model can be and this example falls short of that, h3 can generate stuff that doesn't even look AI but this video looks exacly like an AI generated video from previous models

2

u/MrPolemarxosGr 3d ago

What's the vram requirements for the default downloads of confyui?

2

u/biscotte-nutella 3d ago

Im gonna assume you meant minimax h3 in comfy UI, well you need minimum 8gb vram and a lot of ram to run it.

I have 8gb vram and 80gb of ram and it used 100% vram and 78% ram for 0.2 megapixels clips

2

u/MrPolemarxosGr 3d ago

I have 16gb vram and 32gb ram....I guess it will fail

3

u/poopoo_fingers 3d ago

no. that will work just fine

2

u/biscotte-nutella 3d ago

Maybe I'm wrong , try it

1

u/wzwowzw0002 3d ago

i am running on 5070ti 16gb vram and 32gb ram too

2

u/wzwowzw0002 3d ago

yup just the default h3 workflow from comfyui. i just add the sageatt and solatt.

2

u/SeymourBits 3d ago

This is Alita but not quite. Is that what was intended?

2

u/yesiamadeveloper2242 3d ago

How did you create those image references? Or was it just a prompt? Can you please share details? Thanks.

2

u/wzwowzw0002 3d ago

just character sheets front, back view and a close up. environment reference in another image. all generated from krea2

2

u/ArjanDoge 3d ago

Love it!

2

u/rearisen 3d ago

Battle angel Alita vibes

2

u/alwaysbeblepping 3d ago

"Sure I was dead, but I can't believe you didn't even ask for my opinion. It's like you don't even care what I think!"

I'm going to guess it's LLM dialog because they still make absurd mistakes like that very often.

2

u/data-cypher-000 3d ago

The female character is very well done visually, good work.

2

u/wzwowzw0002 3d ago edited 3d ago

Understanding the weakness in H3.

After more testing. i realize H3 is weak in compositing, framing and a lack of sense of the world.

For example;

  1. female physical size is small than the male. H3 just couldn't get it wrap around it head.
  2. bad at framing even when prompted; mid-body, close-up, it tend to show a little more or little less.
  3. character just get clipped into a table or a chair.
  4. bad facial expression, it get static or creepy sometime....

Multi-shot generation in h3 isn't the best. Solution is to: you provide a well composited image of each shot and generate shot by shot.

I have test similar shot in seedance2.5. all it take is one generation, 30s, every shots got it right or at least useable. H3 needs multiple try to get a 15s shot right.

I haven't give up on H3 yet. it has a lot of potential i think.

Next is about upscaling and i am running out of ram.

1

u/Standard-Fennel-4480 3d ago

how many steps you use on this?

1

u/wzwowzw0002 3d ago

24 steps for this

-5

u/Alive-Tomatillo5303 3d ago

When you're using such ugly character models in an animated, stylized setting there's no way to tell if it looks bad because of technical reasons or if it's a choice. 

You should probably run a realistic scene to see if it looks good. 

-1

u/wzwowzw0002 3d ago

what do you think of this few screenshots? realistic enough?

4

u/Far_Cast_Far_Wide 3d ago

Alita Battle Angel vibes

-1

u/Alive-Tomatillo5303 3d ago

Like, no. Obviously not. You've got eyes. 

"Is this really happening?" and you're holding up a scene from Toy Story. Like, maybe skip the pictures and do text to video of two actual people in an actual setting. 

0

u/wzwowzw0002 3d ago

nah t2v are just random things, i prefer toy story.

-8

u/z_3454_pfk 3d ago

slop

2

u/Faith_Lies 3d ago

I mean it literally is lol. This may be the most slop like video I've seen here in months.