Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict m…
机构:浙大
来源:arXiv 2610.06670 | AI4Papers 论文推荐平台