Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under…
机构:Amazon
来源:arXiv 2608.19739 | AI4Papers 论文推荐平台