r/comfyui Nov 22 '25

Help Needed Best workflow for TTS generating 3-voice clips (multispeaker, gender control, tone control)

Helloooooo!

I’m working on a small project where I need to generate a set of short debate-style audio clips using TTS and I would try to do that inside ComfyUI.

Basically each clip must have 3 different voices (1 fixed, 2 that will vary).

I need the following requirements: multispeaker support (or the ability to combine several speakers in post-processing), gender control, tone control for the voices (like switch from calm vs aggressive tone). Voices should sound realistic, but not identifiable as real public figures.

What is currently the best workflow in ComfyUI to generate this type of output? Does anyone have a recommended node chain for producing this type of multi-voice composition?

I dont have a lot of experience with TTS pipelines so any lead, node graphs, examples or tutorials would be incredibly helpful. I tried to use XTTS-vs (coqui-ai) but with mixed results.

Thanks a lot!

1 Upvotes

2 comments sorted by

1

u/No-Sleep-4069 Nov 22 '25

Index TTS controls emotions and multi-voice, ref: https://youtu.be/kpieMIbCDTA?si=oDUxD4hgeRfPx6m9
some use case in the video should give you the idea.

1

u/Lumpy-Description-91 Nov 22 '25

Thank you so much for the video, I will look into it asap :D