The $40 seat with an AI assistant inside is the scenario I see most often, and the invoice surprise usually comes from the same place: the plan was priced on the average user, and the margin is decided by the heaviest ten percent.
Two things have helped on the builds we have done. First, metering per tenant and per feature from day one, even when nothing is billed on it yet. You cannot price what you never measured, and adding attribution after three months of usage means guessing about the period that matters most. Second, deciding in product terms what happens at the limit before launch: a hard stop, a cheaper model or a queue. Silently continuing is the default when nobody makes that call, and it is the one that hurts.
I also think the unit you price on should be something the customer recognises, such as a report produced or a ticket resolved, rather than tokens. It is easier to explain, and it forces the team to define exactly what one unit of finished work is, which tends to clarify the feature itself.