PCThanks, glad it landed. I used a stronger judge, but even that can't really be trusted without evaluating it on its own. I see you've been working on RAG evals including LLM judges, so I'm curious how your experience has been with evaluating the judge itself.Reply·Article·Jul 11·Building a Kubernetes docs assistant that refuses to guess