Versioning the prompt is only useful for debugging if each run records which revision was actually rendered. A repository can contain the right template while a worker still serves an older cached one, so the deployed identifier needs to travel with the request trace.
I would record the template revision, model configuration, and adapter version together, then verify that a rollout changes those identifiers on live requests. That gives the production regression a reproducible configuration instead of merely proving that someone committed an updated prompt file.
Putting prompt versioning and user input is also data next to each other is the right instinct, because those two sections carry most of the operational risk in this list. Unversioned prompts are the reason nobody can explain a behaviour change: the model is pinned, the code is in git, and the one artifact that actually determines the output was edited in a dashboard by someone in a hurry. On fallback routing I would add a caution - a fallback to a different provider is not a transparent substitution, since output shape, refusal behaviour and token accounting all shift, so anything downstream that parses the response needs to survive the swap. That is where structured output plus validation earns its place: the validator is what turns a silent degradation into a caught error.