The structured tool-result example would benefit from an explicit completeness field. A log query can succeed while returning only the first page, and a command can run successfully while its captured output is truncated. A status of success alone does not tell the agent whether it has enough evidence to conclude that no other errors occurred.
I would return returned_count, a continuation token or has_more flag, and a truncation indicator where relevant. Then a fixture with the decisive error just beyond the first page can check whether the agent fetches more evidence instead of treating a partial result as the whole system state.
The structured tool-result example would benefit from an explicit completeness field. A log query can succeed while returning only the first page, and a command can run successfully while its captured output is truncated. A status of success alone does not tell the agent whether it has enough evidence to conclude that no other errors occurred.
I would return returned_count, a continuation token or has_more flag, and a truncation indicator where relevant. Then a fixture with the decisive error just beyond the first page can check whether the agent fetches more evidence instead of treating a partial result as the whole system state.