Evaluating WebUI code generation at scale is difficult: outputs may compile and look plausible yet fail under user interaction, and prior benchmarks largely …
机构:Meta
来源:arXiv 2610.02617 | AI4Papers 论文推荐平台