DK
That's a great point. You're right that showing only input_tokens can hide a huge part of the actual input usage when prompt caching is involved. I'll look into splitting cache_read_input_tokens and cache_creation_input_tokens in the per-model breakdown. And yes, the SSE parsing is probably one of the more interesting parts of the proxy — thanks for the feedback!
