A static list goes stale fast. The module behind last quarter's incident often isn't on it yet. Scoring each change against what has actually broken before keeps that list current, and tells the human reviewer which PRs deserve a real comment instead of a click.
The audit point lands too. "A check ran" is a timestamp. A verdict that records why a change was judged safe or not is something a person can stand behind. That's the gap Tomosu works on: a reliability score and merge verdict on every PR with the reasoning attached, using incident history once it's connected, and a human still making the call.
A static list goes stale fast. The module behind last quarter's incident often isn't on it yet. Scoring each change against what has actually broken before keeps that list current, and tells the human reviewer which PRs deserve a real comment instead of a click.
The audit point lands too. "A check ran" is a timestamp. A verdict that records why a change was judged safe or not is something a person can stand behind. That's the gap Tomosu works on: a reliability score and merge verdict on every PR with the reasoning attached, using incident history once it's connected, and a human still making the call.