Checking exact values after one compaction is a good start. I would also test successive compactions, because the second summarizer may only see the first summary rather than the original turns. A detail preserved once can disappear after several rounds even if each summary looks reasonable on its own.
A small fixture could carry an exact filename, a rejected alternative, and an unresolved request through three compact-and-continue cycles. Compare those fields after every cycle, including whether the rejected alternative accidentally becomes the selected plan. That would measure cumulative semantic loss separately from the illustrative four-turn or twenty-three-turn cost break-even.
Checking exact values after one compaction is a good start. I would also test successive compactions, because the second summarizer may only see the first summary rather than the original turns. A detail preserved once can disappear after several rounds even if each summary looks reasonable on its own.
A small fixture could carry an exact filename, a rejected alternative, and an unresolved request through three compact-and-continue cycles. Compare those fields after every cycle, including whether the rejected alternative accidentally becomes the selected plan. That would measure cumulative semantic loss separately from the illustrative four-turn or twenty-three-turn cost break-even.