Catching your own judge lying is the plot twist everyone using LLM-as-judge needs to sit with. I have seen a judge confidently pass answers it should have failed, which means you have to evaluate the evaluator before you trust its scores. Here is how I A/B test prompts without fooling myself in the same way: kartiknvjk.hashnode.dev/how-i-a-b-test-llm-prompt…
Kartik N V J K
Catching your own judge lying is the plot twist everyone using LLM-as-judge needs to sit with. I have seen a judge confidently pass answers it should have failed, which means you have to evaluate the evaluator before you trust its scores. Here is how I A/B test prompts without fooling myself in the same way: kartiknvjk.hashnode.dev/how-i-a-b-test-llm-prompt…