Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all ed…
机构:阿里
来源:arXiv 2609.35228 | AI4Papers 论文推荐平台