Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the pres…
机构:西安交大
来源:arXiv 2610.06399 | AI4Papers 论文推荐平台