The design choice that sells this is holding the model constant and only swapping the tool surface, so the 49 vs 60 gap can only come from the contracts. The six unsupported-request prompts are the sharper measurement though. The vague surface reached for a generic lookup all six times, meaning the model never treated the absence of a fitting tool as a signal to stop and ask. That confident wrong call is the failure mode that actually bites in production, and capping it at directional evidence instead of overselling one model makes it easier to trust