The 300ms framing matches what I've seen too, and it forces streaming synthesis into a day-one architecture decision rather than an optimization pass, exactly as you describe. One thing worth adding for the SIP path specifically: real phone callers interrupt mid-sentence far more than a browser-mic session does, so the overlapped generation-and-synthesis pipeline needs a fast, reliable barge-in path that cancels an in-flight synthesis job and stops the audio stream the instant caller speech is detected, not just at the end of the current turn. Getting that wrong on a phone line is worse than on a demo, because there's no visual turn-taking cue for the caller to lean on, and a stray tail of already-buffered audio playing over the caller's own speech reads as the system not listening at all.