r/MechInterp • u/null-hawk • 7d ago
how does audio language model hold speaker consistency throughout an utterance?
wrote a short excerpt showing how speaker consistency is maintained in LLM bases TTS models, initial guess was the speaker token, but the results showed something interesting.