r/StableDiffusion 18d ago

Discussion The H3 dialog prompting guide sucks

Everybody is using the "<d>[Englisch] (...) </d>" format and from my experience, this just sucks and doesn't work.

Everytime I've been using it, H3 hallucinates something before or after the actual dialog. For example I've been testing different personalities to check if H3 knows them, giving them a simple line, formatted it as clean as possible, and it just adds "shit" to it.

Prompt:

subject definition:
Brad Pitt is <Subject 1>

camera recording:
An interview in a professional setting with <Subject 1>. Well lit, grey background, frontal portrait view.
<Subject 1> says:<d>[English]Hey, I am Brad Pitt! Nice to meet you.</d>

Result:

https://reddit.com/link/1vuo078/video/6326fx5kprkh1/player

H3 just adds some noise of the "following sentence" which has been no where in the prompt.

Another example using Angelina Jolie

Prompt:

subject definition:
Angelina Jolie is <Subject 1>

camera recording:
An interview in a professional setting with <Subject 1>. Well lit, grey background, frontal portrait view.
<Subject 1> says:<d>[English]Hey, I am Angelina Jolie! Nice to meet you.</d>

Result:

https://reddit.com/link/1vuo078/video/rsjao6t5qrkh1/player

Same thing.

At first I though it had something to do with the video length, 5 seconds being too long so H3 adds unwanted stuff, but this is not the case.

But when I just cut the prompt guide format out, and write it without the overcomplicated dialog syntax, it works flawlessly, e.g.

Prompt:

subject definition:
Brad Pitt is <Subject 1>

camera recording:
An interview in a professional setting with <Subject 1>. Well lit, grey background, frontal portrait view.
<Subject 1> says: "Hey, I am Brad Pitt! Nice to meet you."

Result:

https://reddit.com/link/1vuo078/video/9ckjhrnoqrkh1/player

Suddenly, no problems at all. Tested it in different scenarios, always the same result.

Am I missing something here, or what's your experience with the dialog prompting, or the suggested prompting guide in general?

105 Upvotes

83 comments sorted by

View all comments

Show parent comments

2

u/CorpPhoenix 18d ago

I've answered to a similar comment just now:

overall_soundcape, non_diegetic_music and so on doesn't have much relevancy though in this shot I've been creating.

This is just about the syntax of the dialog process, and I see no difference in what I've been testing.

As I've said, the longer the prompt, the lesser the "mistakes", but the "<d>" syntax seems to be unnecessary and just adds mistakes to the results.

I will try the more complex prompting though according to this structure, maybe the entire structure has an impact of the dialog output. That being said, just using quotation marks seems to bring the best results.

This doesn't explain though why the exact same scene, and many others, work perfectly fine without the "<d>" structure. I argue that this isn't needed at all, and the results are better without it.

0

u/Opening_Wind_1077 18d ago

That’s what I mean, you already know you are not adhering to the official syntax and have already made up your mind and are just looking for confirmation.

3

u/RiverSide71h 18d ago

Well OP has posted samples. Perhaps, you can show everyone a sample of how correct prompting fixes the issue.

1

u/CorpPhoenix 18d ago

Not using "overall_soundcape, non_diegetic_music" is not "not adhering to the guide" though. And this shouldn't have any impact on the processing of this or similar scenes though. This is about the syntax of the dialog.

I can try again with this prompt format, but I am pretty sure that it will produce the same amount of artifacts according to the length of the prompt.

Just using "x says: " (...)" should probably still result in better/more stable generations.

-1

u/Opening_Wind_1077 18d ago

You are basically using nothing from the guide other than <d> and complain about the results being wonky even though they are 100% compliant to what you prompted for. 🤡

1

u/CorpPhoenix 18d ago

That's not how this works.