Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory…
机构:Stanford
来源:arXiv 2609.38095 | AI4Papers 论文推荐平台