Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignor…
机构:阿里
来源:arXiv 2609.28560 | AI4Papers 论文推荐平台