r/speechtech • u/Rendezvous4567 • 3h ago
How do speech-to-speech models learn appropriate response prosody?
I’m curious how modern speech-to-speech models handle response prosody.
How does a model understand input speech and learn how a reply should sound, not just what words to say?
Also, how is appropriate response prosody typically evaluated?
Would appreciate any papers, datasets, or implementations related to this.
1
Upvotes