Yeah, fair on all of these.
The 403s, you're right, that's just the server blocking my request, not an actual robots rule. A 4xx there basically means "no rules, go ahead," so lumping them in with real blocks was sloppy of me. Will reword it.
WSJ is the one that actually bugs me. I only counted a block when a site named the bot and did Disallow: /. WSJ flips it: blocks everyone with User-agent: * and then allowlists a few. So my script saw "no named blocks" and called it open, when it's really the opposite. Means I'm undercounting all the default-deny sites, not just WSJ. Gonna rerun it to catch that and fix the numbers.
8 vs 9, yeah that's just a mistake. Text says eight, table has nine because I threw Amazonbot in and forgot to line them up. Will fix.
News number is 20 out of 25 sites, so small sample, fair to want that called out. I'll add it.
Good catches. Reworking it.
OnlineProxy
The 403s are probably just bot blocking and not a crawl policy, since a 4xx on robots.txt normally means no rules were found
I checked the WSJ file today and it looks pretty strict, it disallows everything by default and only allowlists a few agents, so counting only named blocks might undercount closed sites.
Small thing -- the text says 8 crawlers but the table has 9, and a sample size for the news numbers would be nice to see