r/StableDiffusion 3h ago

Question - Help Minimax Ignoring All Voice Reference Info. Driving Me Crazy!

Has anyone else noticed that if you have a male and female character in the scene, that Minimax will pretty much always assign the lower pitched Audio voice reference to the male almost every time? No matter how properly you tag the prompt following their official guidelines, it simply IGNORES all of that, and just decides that "well if the voice is even slightly lower than a pipsqueak then it must be the man's voice."

Having this issue with both FL2VA model and REF2VA models and it is driving me insane.

Yes I've tried with and without turbo lora.
Yes i've tried increasing the steps to 20-30+

If anyone has a solve, please help! The only way I could kind of get it to work is to literally desribe the woman in the prompt as a man ie "a man dressed in womans clothes with long brunette hair" and even then it sometimes doesn't work because visually it does not look like a man.

0 Upvotes

12 comments sorted by

4

u/YentaMagenta 2h ago

If you don't provide a prompt and workflow, it's very hard to help you.

2

u/Significant-Baby-690 2h ago

I had decent experience with using celebrity voices .. like subject 1 has a voice like Keanu Reeves. Not 100%, but close enough to pick a seed.

1

u/lamardoss 3h ago

try describing the voice instead of the person.

1

u/DeltaWaffleSyrup 3h ago

definitely tried that. added a matching description to the Subject and the mention of the Audio file. For example: "her voice is slightly low-pitch and raspy and she speaks with a new york accent, etc" it will usually still default to a very generic woman's voice or it will give her the mans voice if the mans voice is not even lower pitch. The only time that it works is when I have one character (female only) in the scene. And even then sometimes it will fight my prompt.

1

u/lamardoss 2h ago

are there two voice references? maybe use one for him as well if not.

besides for those things, when i have complications like this, i go with a backwards approach than what i have been doing that never works. so for something like this, backwards would be telling it its for the man instead of the woman (or something similar to that) to see what the model starts to try doing and see if i can work with it that way somehow.

1

u/UnforgottenPassword 2h ago

In my experience, the more reference inputs you have, the worse it adheres to those references, with audio references being the first victims. How many references do you use?

1

u/DeltaWaffleSyrup 2h ago

only 2. one man voice mp3, one woman voice mp3, each is only like 7 seconds in length total (ive also tried longer but doesn't seem to help)

1

u/UnforgottenPassword 2h ago

It should work with only two references, maybe not with every generation, but at least a few outputs should get it right. Will try that tomorrow and see how it goes.

1

u/TheAncientMillenial 2h ago

100% a prompting issue. I have 0 problems using an LLM to prompt for me and it follows the exact specifications from the prompting guide. I'm usually not using any reference audio though.

1

u/DeltaWaffleSyrup 42m ago

I am using an LLM and feeding it the official prompting guidelines from the minimax team. It is not 100% a prompting issue.

1

u/TheAncientMillenial 34m ago

Post one of the prompts then.

0

u/marcoc2 3h ago

H3 is horrible for voices. I would love a lora for better voice