Nothing here yet.
Idempotency keys only help if the key survives the retry path intact, and most of the interesting failures sit there rather than in the handler. A client that regenerates the key per attempt has built a retry storm with extra steps. A proxy that strips or rewrites headers does the same from the middle. The store also has to outlive the retry window, because a key that expires while the client is still backing off gives you the duplicate you were trying to prevent.
Good AGENTS.md is hard to write because most of them turn into a wish list, and stack-agnostic versions usually go vague on purpose to stay portable. The part worth keeping here is the security boundaries sitting alongside the small-change rule, since boundaries are the thing people drop first when the file gets long. I've kept one model per bounded context inside a monolith.
I used to think of tenant isolation as a database concern, something you enforce at the query layer and then stop worrying about. A lot of it turns out to live in the agent's own working memory, where a summarised history or a cached plan can carry one tenant's context straight into the next. Calling it contamination rather than a query bug feels like the right instinct. Glad someone is writing this down.
Claude stops when the work looks done. It judges done from the same conversation that produced the code. Nothing in that context is independent of the claim. Move the check out of the chat, a hook or a script that runs on a fresh checkout and can't see the reasoning, and "looks done" is no longer something it can conclude. Same agent, same model, just judged by something that wasn't there when it decided.
Separating the loop from the state is what lets a long-horizon run keep going, but that only works if the state is written to be read by a fresh context rather than inherited from the old one. Much of the value sits in the parts that get cut first when you compress the history and keep going. A brief that survives being handed to the next harness is a different artefact from a transcript that got trimmed. https://prickles.org/tenet/persistent-brief/AI2
Any model has no concept of the requirement it just satisfied, so its output can only ever be reviewed against itself. A passing test written by the same tool is worthless as evidence, because both the code and the test agree on the same wrong reading. The review has to come from something that isn't the generator, whether that's a person or a specification the tool never saw. Point the tool at code nobody wrote with it.
I was in the camp that said let it write however it likes, the content is what matters, and then watched a session where every second paragraph was that same line about being right to disagree at the start of a draft and I stopped reading the findings. The filler is where the hedging lives, so the tics were telling me which claims hadn't been checked. One thing I'd try from my own specs: keep the why of each convention next to the rule, so a voice rule like "no second-person flattery" is there with the reason and survives me rewriting it six months later. Bans without a stated reason get deleted the first time someone thinks they're noise. https://tone-of-voice-generator.com/instruction-length-limits