You are right that inspecting reasoning and tool calls still does not tell you the outcome was correct. I score the outcome against user intent, because an agent can take every step cleanly and still land on the wrong page. For browser agents, how are you defining a correct outcome in a way you can check automatically?