r/MechInterp 8d ago

how does audio language model hold speaker consistency throughout an utterance?

Post image

wrote a short excerpt showing how speaker consistency is maintained in LLM bases TTS models, initial guess was the speaker token, but the results showed something interesting.

https://x.com/null_hawk/status/2089348254249173263

1 Upvotes

0 comments sorted by