Communication-Efficient Parallelization Strategy for Deep Convolutional Neural Network Training
SC Workshop, 2018
This paper proposes a communication-efficient gradient-averaging algorithm for synchronous SGD that maximizes computation-communication overlap and is shown analytically to outperform traditional allreduce, achieving up to 2516x (VGG-16) and 2734x (ResNet-50) speedups on ImageNet with up to 8192 cores on NERSC’s Cori supercomputer.
BibTeX
@inproceedings{lee2018communication,
title={Communication-efficient parallelization strategy for deep convolutional neural network training},
author={Lee, Sunwoo and Agrawal, Ankit and Balaprakash, Prasanna and Choudhary, Alok and Liao, Wei-Keng},
booktitle={2018 IEEE/ACM Machine Learning in HPC Environments (MLHPC)},
pages={47--56},
year={2018},
organization={IEEE}
}