Yeah, this one caught us early. Barge-in cancels the in flight generation and flushes the playout buffer in the same step, so nothing that was already queued keeps playing.
The flush is the part that's easy to miss. If you only stop feeding the buffer, the tail still goes out over the caller and it sounds exactly like you said, the system not listening.
The other half of it was on the model side. Audio the caller never actually heard has to come out of the conversation history too, otherwise the agent thinks it said something nobody heard and the next turn answers a question that was never asked.
And you're right about the missing visual cue. That's why we count barge in response time inside the same 300ms budget instead of treating it as its own thing.
The 300ms framing matches what I've seen too, and it forces streaming synthesis into a day-one architecture decision rather than an optimization pass, exactly as you describe. One thing worth adding for the SIP path specifically: real phone callers interrupt mid-sentence far more than a browser-mic session does, so the overlapped generation-and-synthesis pipeline needs a fast, reliable barge-in path that cancels an in-flight synthesis job and stops the audio stream the instant caller speech is detected, not just at the end of the current turn. Getting that wrong on a phone line is worse than on a demo, because there's no visual turn-taking cue for the caller to lean on, and a stray tail of already-buffered audio playing over the caller's own speech reads as the system not listening at all.