The restart-on-below-0.7 behavior is the most interesting stage because it turns retrieval quality into an explicit abstention policy. That threshold only remains meaningful if scores are calibrated and versioned with the embedding and reranker models; a model update can shift the distribution without changing user-visible relevance. I would monitor restart and empty-result rates by query class, since a spike there would reveal drift long before citation counts do.