I'd add one more thing that trips people up early on, placement inference before you even get to extraction. Cheap classify on rendered bytes first then extract then run field level gates against your baseline, doing extraction before that check just wastes cycles on pages you were going to quarantine anyway. Also worth saying loud, defended sites rarely bother with retry after, so your backoff needs to run off a typed failure class instead of waiting on a header that isn't coming. On the buy side the question that actually separates vendors is who owns classify and validate after the fetch, most of them will dodge that and you realize you're paying for infrastructure, not the architecture
I'd add one more thing that trips people up early on, placement inference before you even get to extraction. Cheap classify on rendered bytes first then extract then run field level gates against your baseline, doing extraction before that check just wastes cycles on pages you were going to quarantine anyway. Also worth saying loud, defended sites rarely bother with retry after, so your backoff needs to run off a typed failure class instead of waiting on a header that isn't coming. On the buy side the question that actually separates vendors is who owns classify and validate after the fetch, most of them will dodge that and you realize you're paying for infrastructure, not the architecture