Yeah, this one caught us early. Barge-in cancels the in flight generation and flushes the playout buffer in the same step, so nothing that was already queued keeps playing. The flush is the part that's easy to miss. If you only stop feeding the buffer, the tail still goes out over the caller and it sounds exactly like you said, the system not listening. The other half of it was on the model side. Audio the caller never actually heard has to come out of the conversation history too, otherwise the agent thinks it said something nobody heard and the next turn answers a question that was never asked. And you're right about the missing visual cue. That's why we count barge in response time inside the same 300ms budget instead of treating it as its own thing.
