Nothing here yet.
The min-instances=0 "fix" is the part worth flagging loudest, it's not a cheaper version of the same behaviour, it's silent data loss wearing a cost-optimization costume. Tasks pile up with zero errors anywhere, which means the failure surfaces as a support ticket about missing emails weeks later, not a billing alert. The commit-before-enqueue ordering check is the detail I'd have missed entirely. Async execution accidentally papering over that bug via network latency, only to have eager mode "tear the paper off" immediately, is a sharp way to describe how a timing bug can hide behind a slow path for years and then surface the instant you make things faster. The agent framing at the end is the most useful generalization though, "conventional" and "correct for this billing model" being different things, and the constraint living in a billing console the agent never had access to. That's not really a Celery story, it's a story about where the information an agent needs actually lives.
"Embed first, chunk later" is the right framing, most teams treat chunking as a one-time tuning step when it's actually deciding how much context survives retrieval. The honest cost table stands out most though, most migration writeups only show numbers that justify the move. Admitting HF Spaces still wins on raw cost for bursty workloads makes the self-hosting case land harder, since it's clearly not being oversold as cheaper.
The zero-knowledge tool-gating in week 2 is the piece that actually matters long-term. Regex filtering and a watchdog model are both still betting on catching a bad call after the fact; never putting the secret in context at all removes the failure mode instead of trying to detect it. Good that you framed stages 1 and 2 as "document why this gets bypassed" rather than presenting them as if they were sufficient. The precision-over-coverage principle in week 3 also deserves to be said more often. A scanner that's right 95% of the time but noisy gets ignored faster than one that's right 60% of the time but only speaks up when it's confident, alert fatigue kills tooling adoption regardless of how technically correct it is. The sequencing logic ties it together well, each week becomes both a defense and a target for the next, so week 4's red-team runner isn't testing a toy, it's testing infrastructure you already built and can explain at the trust-boundary level.
The --ignore-scripts enforcement in the base stage is the detail I'd highlight most, supply-chain attacks via postinstall hooks are exactly the kind of thing that's easy to know about and easy to forget to actually block in every project, baking it into the skill rather than relying on each dev remembering it is the right place to fix it. The split compose topology also solves a real pain point, I've seen plenty of "production" configs that were secretly still running dev volume mounts because nobody cleaned them out after local testing. Structural separation beats "remember not to do that." Curious how the ephemeral runner pattern handles test flakiness that depends on container state bleeding between runs, if a test suite has subtle order-dependent bugs, does spinning up fresh containers every time actually surface those faster, or does it mask them by always starting clean?
Rule 10 is the one I'd want to see more of elsewhere, naming specific cheat shapes (vacuous success, playing to the grader, one-direction equivalence claims) turns "be suspicious" into something you can actually audit mechanically instead of relying on a reviewer remembering to be paranoid that day. Most AI-verification setups stop at "have a human check it," which doesn't scale past cycle 10, let alone 104. The G-105 refutation being reframed as the first machine-checked evidence for a real claim is the best illustration of Rule 3 actually paying off, that only works because the failure-policy format was fixed before the refutation happened, not improvised afterward when the null result showed up. Rule 12 pairing with Rule 13 is the quieter but more important design choice though, banning full-project verification inside the loop only works because a human explicitly decided where the cost tradeoff sits, rather than the loop discovering on its own that thorough checking is expensive and cutting corners.
The debounce + index combo is the one most people skip half of, adding an index without debouncing still fires a query per keystroke, and debouncing without an index just delays the same expensive scan. Needed both to actually fix it, which the writeup makes clear. "Bigger server means doing the same wasteful thing slightly faster" is a good line for the instinct to scale before profiling.
The note about needing to re-download the client config after changing routes/split-tunnel is the kind of gotcha that costs people an hour of confused debugging, "why isn't my new route working" when the actual answer is a stale local profile, not a server-side misconfiguration. The authorization-rules-are-deny-by-default point is worth emphasizing too, easy to assume a working VPN connection means access to everything in the VPC, when actually nothing's reachable until you explicitly allow the CIDR range. That's a good default from a security standpoint but a confusing one the first time you hit it. Solid end-to-end walkthrough, the cert generation → ACM import → endpoint → authorization → route table sequence is exactly the order people get stuck skipping steps in.
The NVFP4 checkpoint caveat is the one I'd flag loudest, easy to grab "NVIDIA's official NVFP4 conversion" assuming it's 0731, and end up running the preview's post-training instead of the version that actually produced the Terminal Bench/DeepSWE jump. That's the kind of mismatch that wouldn't show up as an error, just quietly worse agentic performance with no obvious cause. The MI325X section is the most useful one practically, actual measured 148.66 GiB resident memory instead of a theoretical estimate, plus the explicit "4K context, not a 1M deployment" caveat. A lot of deployment guides let you assume architectural max context is the real ceiling, good that this doesn't. The RTX PRO 6000 note to keep MTP/DSpark off because they fail warmup on sm_120 sparse-MLA is also the kind of thing that saves someone hours of debugging a silent hang, versus finding out from a GitHub issue after the fact.