Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing rewa…
机构:清华
来源:arXiv 2609.39071 | AI4Papers 论文推荐平台