The "golden dataset first, metrics fall out of it" ordering is the part I wish I had internalized earlier, because I spent a while picking metrics off a list and measuring things that did not matter. The one habit I would add to your four grading methods: version the dataset and log which examples each change fixed or broke, since an LLM-as-a-judge score that moves is useless if you cannot point at the specific cases that regressed. I have found even 20 well-chosen adversarial examples beat a large but uncurated set.