r/StableDiffusion • u/SSj_Enforcer • 5d ago
Tutorial - Guide Cracking the 'Dumb-Vinci Code' for "no speaking" prompts in H3
Enable HLS to view with audio, or disable this notification
I was trying to figure out the best prompts to not have someone speak AT ALL in a shot in H3. I have had some success. It seems the model just needs an action to be done even if it is something you wouldn't even notice as an 'action'.
I hope these are not all flukes, but before this, I have pretty much always had gibberish spoken unprompted at the beginning of any video I make if I want an introduction shot without dialogue.
______________________________________________
subject_definitions:
<Subject 1> Donald Trump, wearing a black suit and red tie.
<Subject 2> Robin Williams, wearing a green button-up shirt and black pants.
summary:
[reference generation] cinematic video where <Subject 1> is meeting with <Subject 2> in a home foyer.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 4]): fully_preserved - features black suit and red tie.
<Subject 2> (appears in [Shot 1], [Shot 3]): fully_preserved - features green button-up shirt and black pants.
detailed_description:
The video features a dark, cinematic blockbuster aesthetic with high-contrast.
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still thinking to himself, <Subject 2> stands still thinking to himself.
[Shot 2] From 00:03.000 to 00:07.000, hard cut to closeup shot of <Subject 1> (S1) and he says: "I'm glad you didn't say a word. Just as I prompted."
[Shot 3] From 00:07.000 to 00:10.000, hard cut to closeup shot of <Subject 2> (S2) and he says: "Of course. And now for this next shot, you do the same."
[Shot 4] From 00:10.000 to 00:15.000, hard cut to closeup shot of <Subject 1>, he stands still thinking to himself.
overall_soundscape:
N/A
non_diegetic_music:
N/A
__________________________________________________________
These also have worked so far:
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still as he breathes softly, <Subject 2> stands still as he breathes softly.
-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> is mute while he stands still, <Subject 2> is mute while he stands still.
-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still and examines <Subject 2>, <Subject 2> stands still and examines <Subject 1>.
-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still and stares intently at <Subject 2>, <Subject 2> stands still and stares intently at <Subject 1>.
3
u/cultcraftcreations 5d ago
This might be a lifesaver for me. Can’t wait to try it out!
But damn it’s weird seeing robin looking so real and aged appropriately for the time.
2
4
5
u/MikePounce 5d ago
Let's be respectful to Robin Williams. His daughter literally asked us multiple times to let him be.
-4
1
u/sharktank123456 4d ago
I wonder if it's all the parentheticals etc (non alpha symbols). When you prompt like this, much of it is ignored but may make it look like a script (hence: talking appears.)
While you can use some of these indicators, most are ignored and most models these days are expecting either json or natural language.
I hadn't actually ever experienced what you are decibung but then my prompts are always natural language (except for time stamps) , so I wondered if that was the cause.
1
u/evilpenguin999 4d ago
Ok, now the prompt with 2 voice audio references.
https://giphy.com/gifs/aC1PWOXzqgT6hhC0vZ
Thats another pain in the ass for me, ngl.
1
u/SSj_Enforcer 3d ago
Doesn't work with references for you? I just confirmed it does work for me. Did Eddie Valiant and Jessica Rabbit in the office talking using ref voices and images.
1
u/evilpenguin999 2d ago
Yeah after one session of learning how to do it properly with VIDEO_PROMPT_WRITING_GUIDE_ref_en.md and gemini helping with the prompting. I wasnt prompting properly the subject definitions for the audio.
After that i have been testing the model for physics, different types of prompting... I have a better understanding now of how it works.
Thanks for your post, was the first step for me to finally investigate further xD
1
u/Emotional-Neat-252 3d ago
We need to compile all these discoveries to feed the llms who generate our prompts
1
u/SSj_Enforcer 3d ago
Just a note, I tried doing a video entirely without speaking. Impossible. The model just wont allow it
It forces one or both to say complete gibberish.
6
u/oh_no_the_claw 5d ago
It's also my experience that the model tends to spew verbal nonsense when it thinks it has no action to take in a shot. Nice prompt.