the core insighr is solid, but i'd recommend you to dig into some extra stuff like hard numbers on what each endpoint actually costs you like scrape vs. search vs. interact, success rates broken down by domain type, and crucially - where this pipeline actually breaks down in the wild , sites that redesign overnight, crawl scope spiraling, rate-limit cascades wreaking havoc. The agent instructions are smartly conservative , like 'stick to official docs, cap your depth, log the failures", but they'd hit different if they explained why models keep trying to dodge these guardrails and how you'd actually catch it happening. One note, they're not claiming retrieval-as-a-layer is some brand-new RAG invention -- the real win is just stitching all your scattered tools into one visible boundary. For teams caught between do we buy firecrawl or build with playwright in house, honestly a quick cost control compliance matrix would make that build-vs-buy call way less fuzzy