A production budget also needs headroom for the next observation, not just the history already present. I would reserve tokens by category - instructions, evidence, tool schemas, next tool result, and final answer - then let compression operate only on the eligible evidence portion. Every summary should retain source IDs and a hash of the covered records so the agent can reopen raw evidence when a later contradiction appears. That makes compression lossy but auditable, and prevents a large tool response from turning a healthy trajectory into emergency truncation.
Ahmet Özel
AI Engineer. Computer Vision, RAG and LLM agents.