Good correction. Caching changes the billing floor, not the context floor.
Once the cache is warm, the input cost drops sharply, but those tokens still occupy the window on every request. So 1,138 tokens is both a cold-request cost and an every-request context footprint, and those two numbers age differently. I should distinguish them in the post.
The tool-schema half is the part nobody else measures. "Per-request floor" frames it better than "tokens per turn" would too. Worth flagging for anyone sizing their own harness against this: on Anthropic's API, tool definitions are cacheable same as the system prompt. So 1,138 is a real per-request cost only until the cache warms. After that it's a fraction of the price on every later turn in the same session, though it still counts against context length either way. Doesn't undercut your point about the number nobody quotes. Just means the cost and context halves of that floor age differently once caching kicks in.