Communication-Efficient Parallelization Strategy for Deep Convolutional Neural Network Training

Sunwoo Lee, Ankit Agrawal, Prasanna Balaprakash, Alok Choudhary, Wei-Keng Liao

SC Workshop, 2018

This paper proposes a communication-efficient gradient-averaging algorithm for synchronous SGD that maximizes computation-communication overlap and is shown analytically to outperform traditional allreduce, achieving up to 2516x (VGG-16) and 2734x (ResNet-50) speedups on ImageNet with up to 8192 cores on NERSC’s Cori supercomputer.

BibTeX

@inproceedings{lee2018communication,
  title={Communication-efficient parallelization strategy for deep convolutional neural network training},
  author={Lee, Sunwoo and Agrawal, Ankit and Balaprakash, Prasanna and Choudhary, Alok and Liao, Wei-Keng},
  booktitle={2018 IEEE/ACM Machine Learning in HPC Environments (MLHPC)},
  pages={47--56},
  year={2018},
  organization={IEEE}
}