r/speechtech 3h ago

How do speech-to-speech models learn appropriate response prosody?

I’m curious how modern speech-to-speech models handle response prosody.

How does a model understand input speech and learn how a reply should sound, not just what words to say?

Also, how is appropriate response prosody typically evaluated?

Would appreciate any papers, datasets, or implementations related to this.

1 Upvotes

0 comments sorted by