one-off work, the article's approach is fine, bounded retries and longer wait times usually sort it out. For production stuff you'll want JSON schema validation, dump raw HTML on failures, and dedupe on canonical URL plus timestamp
container_not_found is often geo-gated pages where country and proxy settings help, but worth checking iframes and shadow DOM too since they can hide data and most AI scrapers handle them already. bad_html, glance at the raw response before retrying since redirects won't fix with more attempts
same page same JSON is actually nice to have, DOM extraction stays consistent and fails visibly when markup changes instead of silently breaking