The framing that the retry loop, not Lambda duration, drives cost lands hard, since a worker-evaluator that bounces back five times quietly multiplies token spend per request. I now log the attempt count per invocation as a first-class metric, because a rising average predicts a bad prompt long before the bill does. Do you cap retries by count alone, or on an evaluator confidence delta too?
The framing that the retry loop, not Lambda duration, drives cost lands hard, since a worker-evaluator that bounces back five times quietly multiplies token spend per request. I now log the attempt count per invocation as a first-class metric, because a rising average predicts a bad prompt long before the bill does. Do you cap retries by count alone, or on an evaluator confidence delta too?