Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior. Looped language models (LoopLMs) repeatedly apply s…
机构:UIUC
来源:arXiv 2610.02741 | AI4Papers 论文推荐平台