The framing that the retry loop, not Lambda duration, drives cost lands hard, since a worker-evaluator that bounces back five times quietly multiplies token spend per request. I now log the attempt count per invocation as a first-class metric, because a rising average predicts a bad prompt long before the bill does. Do you cap retries by count alone, or on an evaluator confidence delta too?