The decision-versus-document distinction gives a concrete way to revisit a product cost model. One cost I would include beside the classifier request is the input passed downstream: if routing requires sending the same large context to both the decision model and the selected generator, the extra reading can offset part of the savings.
That suggests benchmarking a smaller routing payload containing only the fields needed for the decision, then checking disagreements against the full-context version. The useful product metric would be cost per successfully handled request, including unnecessary escalations and missed escalations, rather than savings calculated from classifier pricing alone.
The decision-versus-document distinction gives a concrete way to revisit a product cost model. One cost I would include beside the classifier request is the input passed downstream: if routing requires sending the same large context to both the decision model and the selected generator, the extra reading can offset part of the savings.
That suggests benchmarking a smaller routing payload containing only the fields needed for the decision, then checking disagreements against the full-context version. The useful product metric would be cost per successfully handled request, including unnecessary escalations and missed escalations, rather than savings calculated from classifier pricing alone.