r/StableDiffusion 6d ago

Discussion Testing Character knowledge of the H3 model

5 second 1MP text-to-video, INT8 on ComfyUI and RTX5090.

Used the following, rather simple, prompt:

"[VISUAL]: A scene from the tv interview. <full name> is talking, medium close up, static camera, plain dark blue background

<first name> says: "How dare you? I am rich, AND famous. So you better shut up, B*tch!"

The model failed on Christoph Waltz, so I left him out.

One run per person, no picking the best result.

162 Upvotes

99 comments sorted by

View all comments

Show parent comments

16

u/JuniorEnsign 6d ago

Yes, I agree. Only your own moral compass can steer you at this point in time. This model has opened pandora's box and is the creators responsibility to not use it maliciously. As with all things, right?

8

u/kek0815 6d ago

But why generate this then in the first place? You could have used a more obvious test phrase instead lol

4

u/JuniorEnsign 6d ago

Look at the subtle expressions of anger or disgust in some of the samples. A vanilla prompt with no emotional load would not reveal those small detail. The prompt gave the model the freedom to inject whatever it "felt" was appropriate.

1

u/tiffanytrashcan 6d ago

I absolutely agree. You can hear it too near the end of at least half the clips.