How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already …
机构:NVIDIA
来源:arXiv 2610.10478 | AI4Papers 论文推荐平台