Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at…
机构:牛津
来源:arXiv 2609.35699 | AI4Papers 论文推荐平台