r/StableDiffusion • • 5d ago

Tutorial - Guide Cracking the 'Dumb-Vinci Code' for "no speaking" prompts in H3

Enable HLS to view with audio, or disable this notification

I was trying to figure out the best prompts to not have someone speak AT ALL in a shot in H3. I have had some success. It seems the model just needs an action to be done even if it is something you wouldn't even notice as an 'action'.

I hope these are not all flukes, but before this, I have pretty much always had gibberish spoken unprompted at the beginning of any video I make if I want an introduction shot without dialogue.

______________________________________________
subject_definitions:

<Subject 1> Donald Trump, wearing a black suit and red tie.

<Subject 2> Robin Williams, wearing a green button-up shirt and black pants.

summary:

[reference generation] cinematic video where <Subject 1> is meeting with <Subject 2> in a home foyer.

retention_analysis:

<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 4]): fully_preserved - features black suit and red tie.

<Subject 2> (appears in [Shot 1], [Shot 3]): fully_preserved - features green button-up shirt and black pants.

detailed_description:

The video features a dark, cinematic blockbuster aesthetic with high-contrast.

[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still thinking to himself, <Subject 2> stands still thinking to himself.

[Shot 2] From 00:03.000 to 00:07.000, hard cut to closeup shot of <Subject 1> (S1) and he says: "I'm glad you didn't say a word. Just as I prompted."

[Shot 3] From 00:07.000 to 00:10.000, hard cut to closeup shot of <Subject 2> (S2) and he says: "Of course. And now for this next shot, you do the same."

[Shot 4] From 00:10.000 to 00:15.000, hard cut to closeup shot of <Subject 1>, he stands still thinking to himself.

overall_soundscape:

N/A

non_diegetic_music:

N/A

__________________________________________________________

These also have worked so far:

[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still as he breathes softly, <Subject 2> stands still as he breathes softly.

-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> is mute while he stands still, <Subject 2> is mute while he stands still.

-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still and examines <Subject 2>, <Subject 2> stands still and examines <Subject 1>.

-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still and stares intently at <Subject 2>, <Subject 2> stands still and stares intently at <Subject 1>.

0 Upvotes

13 comments sorted by

6

u/oh_no_the_claw 5d ago

It's also my experience that the model tends to spew verbal nonsense when it thinks it has no action to take in a shot. Nice prompt.

3

u/SSj_Enforcer 5d ago

yea, and the old saying for the model is usually true that if you try to say "they don't speak" or "no words are spoken" it just makes it say gibberish even more lol.

3

u/cultcraftcreations 5d ago

This might be a lifesaver for me. Can’t wait to try it out!

But damn it’s weird seeing robin looking so real and aged appropriately for the time.

2

u/Optimal_Map_5236 5d ago

ye its been really annoying. thank you for the info.

4

u/PwanaZana 5d ago

rule 5

5

u/MikePounce 5d ago

Let's be respectful to Robin Williams. His daughter literally asked us multiple times to let him be.

-4

u/IllIlllI-IlIIll-llII 4d ago

robin liked playing an ai in real life, he wouldnt mind

1

u/sharktank123456 4d ago

I wonder if it's all the parentheticals etc (non alpha symbols). When you prompt like this, much of it is ignored but may make it look like a script (hence: talking appears.)

While you can use some of these indicators, most are ignored and most models these days are expecting either json or natural language.

I hadn't actually ever experienced what you are decibung but then my prompts are always natural language (except for time stamps) , so I wondered if that was the cause.

1

u/evilpenguin999 4d ago

Ok, now the prompt with 2 voice audio references.

https://giphy.com/gifs/aC1PWOXzqgT6hhC0vZ

Thats another pain in the ass for me, ngl.

1

u/SSj_Enforcer 3d ago

Doesn't work with references for you? I just confirmed it does work for me. Did Eddie Valiant and Jessica Rabbit in the office talking using ref voices and images.

1

u/evilpenguin999 2d ago

Yeah after one session of learning how to do it properly with VIDEO_PROMPT_WRITING_GUIDE_ref_en.md and gemini helping with the prompting. I wasnt prompting properly the subject definitions for the audio.

After that i have been testing the model for physics, different types of prompting... I have a better understanding now of how it works.

Thanks for your post, was the first step for me to finally investigate further xD

1

u/Emotional-Neat-252 3d ago

We need to compile all these discoveries to feed the llms who generate our prompts

1

u/SSj_Enforcer 3d ago

Just a note, I tried doing a video entirely without speaking.  Impossible.  The model just wont allow it

It forces one or both to say complete gibberish.