An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrow…
机构:牛津
来源:arXiv 2609.29709 | AI4Papers 论文推荐平台