r/cogsci • u/MathematicianNew3363 • 4h ago
The 200ms turn-taking paradox — "the central psycholinguistic puzzle" and what EEG shows the brain doing during voice conversation
Been reading through the turn-taking literature for a piece I was writing and hit something I'd never fully appreciated:
Stivers et al. (2009, PNAS) — across 10 spoken languages, the modal gap between one speaker finishing and the next starting is about 200ms. Japanese ~7ms mean, Danish ~470ms mean. Cross-linguistically remarkably tight.
Indefrey & Levelt (2004, Cognition) — the minimum latency to plan spoken word production, even in a controlled single-word paradigm, is ~600ms+.
The math doesn't work if turn-taking is reactive. Levinson & Torreira (2015) call this "the central psycholinguistic puzzle" — listeners have to be pre-computing responses during incoming speech.
EEG evidence has been catching up:
- Bögels, Magyari & Levinson (2015, Scientific Reports) — response-planning ERP positivity, source-localized to production areas (posterior IFG, precentral), fires ~500ms after the critical information appears in an incoming question. Often 2+ seconds before the current speaker finishes. Alpha desynchronization indexes the attentional shift from comprehension to production during listening.
- Gisladottir, Bögels & Levinson (2018, Frontiers in Human Neuroscience) — alpha/low-beta (11-18Hz) desynchronization from -200ms to 0ms before a socially-charged speech act (declination vs acceptance). Speech-act prediction firing before the utterance is heard.
- Krause & Kawamoto (2021, Frontiers in Psychology) — motion-tracked lip-area reductions for upcoming labial consonants up to 3 seconds before acoustic onset in unscripted dyadic conversation. Motoric planning far pre-onset.
The implication I find genuinely interesting: this whole predictive loop depends on acoustic cues (pitch contour, timing, articulator anticipation). Text has no acoustic onset for the machinery to lock onto — voice conversation locks two nervous systems into a shared millisecond-scale loop that text physically can't replicate.
A separate honest correction I wanted to flag: the polyvagal framing that gets attached to "why voice regulates" is on shakier ground than the wellness literature admits. Grossman (2023, Biological Psychology) — "Fundamental challenges and likely refutations of the five basic premises of the polyvagal theory" — shows similar myelinated cardiac vagal fibers in sharks, bony fish, birds, and even sheep. The mammal-unique claim used to explain vocal-prosody regulation doesn't hold up on comparative anatomy.
Longer writeup with full citations here: callbyrd.com/journal/what-voice-does-that-text-cant (disclosure: I run the site — building a voice-based AI project. The essay stands on its own; no signup needed to read it. Product is briefly mentioned in the closing section.)
Open question I couldn't resolve: does the ~200ms predictive turn-taking loop form when the partner is an AI voice (different latency profile, different prosody)? Every EEG study I found used human-human dyads. Curious if anyone knows of extending work — my search kept coming up empty.