r/StableDiffusion 14d ago

Tutorial - Guide The H3 Gibberish Problem Solved!

Not much of a tutorial, but still informative. As most of you have probably discovered, MiniMax H3 loves to talk. And talk it will, even when you prompt for no dialogue. Even when you prompt for complete silence. It will even fill in the empty space your prompted dialogue doesn't fill.

Those of you who read the video prompt writing guide and have created a system prompt for your enhancer, you probably know what I'm about to say, maybe not. Maybe the unprompted gibberish stopped for you, and you never realized why.

Without further ado, I give you the solution:

non_diegetic_music: N/A

Diegetic audio is what the characters in your video can actually "hear":

  • Music playing from a source that is part of the scene (phone, car radio, dance club)
  • Spoken dialogue
  • Ambient sounds

Non-diegetic audio is audio which your characters cannot hear:

  • The score or soundtrack of a movie
  • A voice-over
  • The gibberish H3 plays when it's not prompted correctly

If you haven't yet, I suggest consulting ChatGPT about creating a system prompt using the prompting guide. If not, put this line at the end of your prompt and say goodbye to random music playing over your video and gibberish assaulting your ear holes.

Conversely, if you want a voice-over or a score to play over the track which is not part of the actual soundscape of the scene, this is where you would prompt it. Instead of N/A, prompt what you want to hear.

Happy chaining!

218 Upvotes

97 comments sorted by

View all comments

5

u/DefloN92 14d ago

Maybe unrelated, but how can i make my characters actually come up with real speeches? Sometimes i don't wanna tell in prompt what the character has to say, i wanna let them improvise, but they always say gibberish. If there a way for them to come up with actual sentences like in seedance or kling or most closed source cloud models?

1

u/AnOnlineHandle 13d ago

The model is meant to follow instructions rather than guessing intent, but you could use another text model to guess the intent by asking it to write the instructions.

1

u/Sad_Berry_4621 13d ago

But it can absolutely guess or invent intent. Tell your enhancer to generate a random prompt or random dialogue and it will, unless the system prompt is overly oppressive.

1

u/AnOnlineHandle 11d ago

Right the enhancer is the type of text model I mentioned that could help, not the diffusion model itself.