The separation between generation latency and perceived latency is useful, and the moderation-buffering example shows that delaying delivery can have a real purpose. I would distinguish transport packet loss from a broken application stream in the failure section: HTTP over TCP already recovers ordinary packet loss while the connection survives.
The harder UI case is a reconnect after the response was partly displayed. A test could cut the connection after a complete SSE event, then check whether the client marks the answer incomplete, restarts it or resumes from an application event ID. That makes the reliability tradeoff explicit without treating every dropped packet as permanently lost generated text.