There is a special kind of awkwardness in saying “Actually—” to a voice assistant while it continues reading paragraph three of an answer nobody wants anymore.
The transcription may be perfect. The generated answer may be correct. The voice may be warm, expressive, and expensive. The conversation still feels rude.
Voice interfaces have spent years getting better at words. Humans notice the gaps.
A transcript is not a conversation
Written chat has explicit turns. I press Send; then the assistant responds. Speech has no send button. People trail off, restart, pause to remember a name, say “mm-hmm” without requesting the floor, and finish each other's sentences with varying levels of permission.
Research such as TurnGPT treats turn-taking as its own prediction problem, using linguistic and conversational context to estimate when a speaker has completed a turn. That distinction matters. Silence is audio; completion is meaning.
The expensive half-second
Respond too quickly and the assistant cuts people off. Wait too long and every exchange feels like a satellite call. A fixed silence threshold cannot win: half a second after “What is the weather...” is different from half a second after “What is the weather in Phoenix tomorrow?”
The system needs several signals at once—acoustics, syntax, semantics, gaze or button state when available, and the cost of being wrong. A medical intake assistant should tolerate thoughtful pauses. A drive-through ordering system may prefer a brisk confirmation.
Let people interrupt
Barge-in support is not an advanced feature. It is the voice equivalent of a Stop button. When the user begins speaking, playback should yield quickly, preserve what was already heard, and understand that “No, Tuesday” is probably a correction—not a fascinating new topic.
Interruptions also reveal answer quality. Long responses feel longer out loud. A screen lets the eye skip; audio charges one second per second. Good voice answers lead with the useful bit and offer depth instead of performing it.
Design the listening
Test conversations, not transcripts. Include background acknowledgments, self-corrections, long pauses, two speakers, a noisy room, and a user changing the goal halfway through. Measure cutoffs, awkward gaps, successful interruptions, and how often people must repeat themselves.
The uncanny valley of voice AI is not only how the voice sounds. It is the moment the machine demonstrates that, despite hearing every word, it has no idea how to listen.