Thanks, Kartik. The stop-reason trap is a good one; a truncated answer that looks finished is exactly the kind of bug the series is about.
Tool-call failures are handled in week one, not deferred. On day 06 a bad call doesn't crash the loop: an invented tool name, arguments cut off mid-JSON, or a wrong parameter go back to the model as the tool's result, an error message listing what's valid, so it can correct itself next turn. The loop is also capped at a fixed number of turns, so a model that keeps calling tools can't run forever.
Week four's hardening (day 28) is about the failures that don't raise anything: a run that finishes cleanly with a map that's quietly wrong.
On truncation: day 05 sets max_tokens so a cut-off answer surfaces as malformed JSON rather than passing as finished. You've convinced me it's worth checking finish_reason explicitly as well.