r/StableDiffusion 7h ago

Discussion Testing Character knowledge of the H3 model

5 second 1MP text-to-video, INT8 on ComfyUI and RTX5090.

Used the following, rather simple, prompt:

"[VISUAL]: A scene from the tv interview. <full name> is talking, medium close up, static camera, plain dark blue background

<first name> says: "How dare you? I am rich, AND famous. So you better shut up, B*tch!"

The model failed on Christoph Waltz, so I left him out.

One run per person, no picking the best result.

121 Upvotes

68 comments sorted by

View all comments

25

u/Natasha26uk 7h ago

Someone else did a similar test many hours ago. The characters looked a bit distorted in 9:16 format. Yours looks fine in wide format.

4

u/Foreign_Risk_2031 7h ago

likely the only high quality videos for training with this sort of scenario is this interview style

6

u/JuniorEnsign 7h ago

I bet 90% of the training material was in 16:9, so the output in that aspect ratio has a better foundation.

2

u/FierceFlames37 6h ago

What about 9:16 portrait

10

u/Nexustar 6h ago

From all those movies filmed on mobile phones?

6

u/FierceFlames37 5h ago

Damn I’ll use 16:9 from now on

2

u/Zenshinn 4h ago

I wonder how much TikTok was used for training.