Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support…
机构:Dell Technologies
来源:arXiv 2609.30471 | AI4Papers 论文推荐平台