Your 20% matches haschek's any modification rate, which is the fair comparison for a hash check since even harmless edits fail it. Malicious-only studies land near 5%, so 4x that is expected
Worth adding a clean streak counter too, some proxies only inject sometimes per the NDSS slides, a single canary fetch misses that. Cert validation stops rewriting but https only will not catch blocking or SNI leaks, your hash check already handles rewriting.
The shodan data mostly points to misconfigured mikrotik boxes rather than malice, tho real malicious proxies exist too, your honeypot check cannot tell you which one you are seeing