Parameter-efficient reinforcement learning aims to improve reasoning with a compact trainable interface to a pretrained model. We introduce the Thalamic Rout…
机构:清华
来源:arXiv 2609.39109 | AI4Papers 论文推荐平台