The training vs retrieval split is useful, and the CDN layer is the one people skip. I’ve seen the same pattern: robots.txt looks open, but the edge is quietly blocking OAI-SearchBot or PerplexityBot, so the allow never matters. Auditing both layers, then checking Security Events for actual fetch attempts, is the right order before anyone worries about schema or “AI-ready” content.