Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token effi…
机构:复旦
来源:arXiv 2609.34718 | AI4Papers 论文推荐平台