The 95% per-target-token threshold and 11 consecutive passes make the acceptance rule unusually transparent. To separate memorization from generalization, I would hold out entire question templates and token orders rather than only examples, then evaluate those frozen paraphrases after adaptive SFT. The catastrophic-forgetting result would also be useful as a curve: rehearsal share versus retained accuracy on the original 14 relations.