The late-chunking flip is the part I keep coming back to, embedding first and cutting after preserves neighbor context that vanilla chunk-then-embed silently drops. I also like that you kept the latency numbers honest (0.14 to 0.8s per chunk on self-hosted GCP vs 1 to 3s on Spaces), the "always-on cost" caveat is what most write-ups skip. Curious whether you saw retrieval quality shift on tables and code blocks specifically, those tend to break mean-pooled token embeddings for us.