It's measurement, not judgment. In my case platform is queryable, so the agent asks it the same question the app asks, then asks again with one thing changed — same record, different user, or the same data without the filter the screen applies. If it's already wrong before our code touches it, the answer is obvious. That's what stops a misdirected ticket, not the diagnosis itself so it is basically re-try loop. Unfortunetely it isn't bulletproof. I hit a case where the platform accepted a filter and quietly ignored it, so both queries matched and "not our code" looked proven. So a probe has to show it could actually see what it claims to check. And agreed on can't-reproduce — it's usually a data problem, which is why the agent has to try arranging data before it's allowed to say it.