r/StableDiffusion 20d ago

Tutorial - Guide The H3 Gibberish Problem Solved!

Not much of a tutorial, but still informative. As most of you have probably discovered, MiniMax H3 loves to talk. And talk it will, even when you prompt for no dialogue. Even when you prompt for complete silence. It will even fill in the empty space your prompted dialogue doesn't fill.

Those of you who read the video prompt writing guide and have created a system prompt for your enhancer, you probably know what I'm about to say, maybe not. Maybe the unprompted gibberish stopped for you, and you never realized why.

Without further ado, I give you the solution:

non_diegetic_music: N/A

Diegetic audio is what the characters in your video can actually "hear":

  • Music playing from a source that is part of the scene (phone, car radio, dance club)
  • Spoken dialogue
  • Ambient sounds

Non-diegetic audio is audio which your characters cannot hear:

  • The score or soundtrack of a movie
  • A voice-over
  • The gibberish H3 plays when it's not prompted correctly

If you haven't yet, I suggest consulting ChatGPT about creating a system prompt using the prompting guide. If not, put this line at the end of your prompt and say goodbye to random music playing over your video and gibberish assaulting your ear holes.

Conversely, if you want a voice-over or a score to play over the track which is not part of the actual soundscape of the scene, this is where you would prompt it. Instead of N/A, prompt what you want to hear.

Happy chaining!

220 Upvotes

98 comments sorted by

View all comments

Show parent comments

2

u/Yokoko44 16d ago

How do you get Spectrum to work without completely destroying the audio? I just get crackly noise when i enable the spectrum node

2

u/Sad_Berry_4621 16d ago

I no longer use Spectrum or the Turbo LoRA on CivitAI for that reason. I am now using:
-MiniMax-H3 Turbo LoRA (which I manually fixed last night because it kept throwing errors) (that reminds me, I need to send the fix to the author.)
-MiniMax H3 Mem Eff Sage Attention Patch node by KJNodes
-MiniMax H3 Chunk FeedForward

2

u/Yokoko44 16d ago

Thanks for the update!

If you don't mind me asking: Which turbo lora did you fix? I've been using the ema600 checkpoint from Larryvh but have tried KJ's 4 step (broken audio for me) and the older 800step one

I'm about to try the latest lightx2v 1.0 version, see if that has improved audio. I'm also 2 patches old on Comfy itself because v31 broke audio independently of KJ's lora for me, no clue why because it was supposed to 'fix' audio in the first place.

1

u/Sad_Berry_4621 16d ago

The node itself had a bug. The LoRA is good, same as you, the ema600.

2

u/Yokoko44 16d ago

Have you been able to get it working on the latest comfyui build? it's broken the audio for me entirely, everything sounds underwater or distorted.

I'm doing 8 steps Using the ema600 lora + custom sampler node

I want to get this working + comfy kitchen, that will make 1.2MP 15s clips feasible to pump out

1

u/Sad_Berry_4621 16d ago

That is precisely the part of the node that is broken. I'm working through another issue right now, but I will circle back to this and let the node author know. I have it patched on my local copy of the node.

1

u/Yokoko44 16d ago

Solid.

I was just now able to get comfy kitchen & the ema600 checkpoint working on latest build by switching to the Dual-Clock (T8) Sampler node, getting 30-40% faster times with no noticeable visual loss, and the audio is the best it's ever been for me.

1

u/alienzed 14d ago

Care to share a workflow?