An LLM judge is a biased instrument, not a measurement
Last month I shipped an eval that ranked two prompt variants. Variant A won by four points. A teammate reran the same eval the next morning and Variant B won. Same model, same judge, same test set. Th
llmasajudge.hashnode.dev8 min read