Mihai_LeanZero
Atlassian consultant and Forge developer at LeanZero. We do Data Center to Cloud migrations and build Forge apps, and I write up what I learn at leanzero.net.
This is one of the most honest build-in-public postmortems I've read here. The retry table stood out to me, the line that anything outside GPU/LLM/network steps gets no retry because a real bug has to fail loudly, not hide behind one. That's the exact rule teams usually lose once retries start creeping in everywhere. We hit something adjacent running unattended agents on macOS launchd: once a step gets retry logic that makes a failure recoverable, the alert rule that used to fire on it quietly stops meaning what it meant when it was written, and nobody goes back to re-tune it once the retry proves itself over a few weeks, so you end up paging on things that now self-heal and staying quiet on the ones that actually changed shape. Did you revisit your 26 alert rules as the vLLM restart-and-resume logic matured, or did the alert set stay fixed from when it was first written? Also curious about the decisions log, once it grows past a few dozen entries does the agent actually read all of it every session, or did you need retrieval or summarization to keep it inside context before it started getting ignored again the same way the stale config did?
The nightly precision audit is the part I'd want to stress-test hardest: it calls Opus fresh each night to re-judge delivered leads, while the recall benchmark runs off labels Opus produced once and froze. Was that nightly call pinned to a specific Opus checkpoint the whole seven months, or the latest alias? If Anthropic moved the model under it at any point, the audit's meaning shifts even though nothing in the pipeline changed -- and with 339k lines nobody read, that audit was the only signal anyone had that anything upstream was still doing its job.