Your LLM-as-judge disagrees with itself between runs
Same outputs, same judge, two runs, two scores. The gate flickered red then green on a branch with zero code changes, and that flapping cost me more trust than any real regression.
The flap
I had a fa
ethanwalkerwrites.hashnode.dev6 min read