Great explainer — the distinction between the model being sequential and the transport being chunked is one most people conflate. The part about partial markdown rendering resonated: tracking which structures are still open instead of naively re-parsing after every chunk is exactly the kind of detail that separates polished chat UIs from broken ones. One related wrinkle worth mentioning: token healing at stream boundaries. When a chunk boundary splits a multi-byte character or a token mid-word, naive concatenation can produce subtly wrong text even after all bytes arrive. Same class of problem as your UTF-8 example, just one layer up. Really enjoyed the Eloquent mention too — nice to see research on resumable streams getting attention.