Separating time to first token from total generation time is the distinction I wish more procurement conversations started with, since a model that starts fast but crawls through a long answer fails a batch job while feeling perfectly responsive in interactive chat. The cost-per-workflow point is the other one people skip, because per-token pricing hides what retries and tool calls actually do to the bill. When you evaluate a provider, do you load-test P99 under a full queue before committing, or is that something teams only find out after the migration is already locked in?