A noisy judge does not just add error bars. It shrinks the effect you are trying to measure.
Most of what I have written about LLM judges argues in one direction: your improvement is probably not real. Small eval sets, selection across many experiments, position bias, clustered examples. All
llmasajudge.hashnode.dev15 min read