The stage-by-stage budget is useful, but summing each stage's 99th percentile is not guaranteed to upper-bound the end-to-end 99th percentile. Different requests can encounter different stages' tails, so more than one percent of requests may exceed that sum even if each stage meets its own p99 budget.
I would keep per-stage percentiles for diagnosis and enforce the final objective on joined end-to-end traces, as your measurement section proposes. A synthetic test with non-overlapping slow-stage events would demonstrate the distinction. Correlation between stages determines how their tails combine; the budget table alone cannot establish that coverage.
The stage-by-stage budget is useful, but summing each stage's 99th percentile is not guaranteed to upper-bound the end-to-end 99th percentile. Different requests can encounter different stages' tails, so more than one percent of requests may exceed that sum even if each stage meets its own p99 budget.
I would keep per-stage percentiles for diagnosis and enforce the final objective on joined end-to-end traces, as your measurement section proposes. A synthetic test with non-overlapping slow-stage events would demonstrate the distinction. Correlation between stages determines how their tails combine; the budget table alone cannot establish that coverage.