The part that stands out is that delegationMode: "prefer" is described as prompt guidance only, it doesn't enforce the handoff and doesn't change tool policy. So the actual proof that a task went through the researcher/writer/reviewer contracts rather than the coordinator just doing it inline isn't in the config at all, it has to come from somewhere else. The acceptance test in the piece checks the shape of the returned artifact, a path, sources, checks, uncertainty, but a coordinator holding tools similar to a specialist's could in principle produce something that looks the same without ever calling sessions_spawn, especially under time pressure where delegating costs more turns than just answering directly. Is there an actual trace of spawn calls per task somewhere, so the split can be audited after the fact instead of trusted from the artifact's shape alone?
The part that stands out is that delegationMode: "prefer" is described as prompt guidance only, it doesn't enforce the handoff and doesn't change tool policy. So the actual proof that a task went through the researcher/writer/reviewer contracts rather than the coordinator just doing it inline isn't in the config at all, it has to come from somewhere else. The acceptance test in the piece checks the shape of the returned artifact, a path, sources, checks, uncertainty, but a coordinator holding tools similar to a specialist's could in principle produce something that looks the same without ever calling sessions_spawn, especially under time pressure where delegating costs more turns than just answering directly. Is there an actual trace of spawn calls per task somewhere, so the split can be audited after the fact instead of trusted from the artifact's shape alone?