Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffu…
机构:KlingAI Research
来源:arXiv 2608.20161 | AI4Papers 论文推荐平台