A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor ca…
机构:Google
来源:arXiv 2609.38142 | AI4Papers 论文推荐平台