Extending your framing one step: "The model chooses the way a user chooses a button in an app they've never used before: by reading the label." Labels also have a total weight, not just individual clarity. Every tool definition rides along in context on every turn, so a 40-tool MCP server with beautifully written descriptions can still degrade selection simply by consuming the budget the model needs for the actual task, which is why several clients now lazy-load tool schemas behind a search step instead of presenting the full palette. That suggests a measurable check in toolens: total token weight of the toolset per common tokenizer, with a warning threshold, alongside the per-tool lints. Over-tooling by count is in there already; over-tooling by tokens is the sibling failure and it hits exactly the teams who followed the advice to write rich descriptions. Would you take a PR adding that?