mL-BFGS: A Momentum-based L-BFGS for Distributed Large-Scale Neural Network Optimization
Transactions on Machine Learning Research, 2023
This work proposes mL-BFGS, a momentum-based, block-wise Hessian approximation to L-BFGS that stabilizes stochastic convergence and distributes compute/memory across nodes, achieving both iteration-wise and wall-clock speedups over SGD, Adam, and other quasi-Newton methods in large-scale distributed DNN training.
BibTeX
@article{niu2023ml,
title={Ml-bfgs: A momentum-based l-bfgs for distributed large-scale neural network optimization},
author={Niu, Yue and Fabian, Zalan and Lee, Sunwoo and Soltanolkotabi, Mahdi and Avestimehr, Salman},
journal={Transactions on machine learning research},
volume={2023},
pages={967},
year={2023}
}