That "read tool that quietly writes" case is exactly why I wouldn't trust the labels. Annotations like a read-only hint are self-reported by the server, so I treat them as UI hints, not a security boundary.
My approach: require human approval by default for anything that can mutate state or reach external systems, auto-approve only an allowlist of tools I've reviewed, and enforce the real boundary below the protocol, with read-only credentials for read use cases, so a mislabeled tool can't write even if it tries. Resources vs tools is a useful design signal, but permissions have to be enforced where the data lives.
Have you found anything that catches mislabeled tools before they run, or is it mostly review and logging?