Passer à la navigation principale Passer à la recherche Passer au contenu principal

Communication trade-offs for Local-SGD with large step size

  • ENAC-IIC-GEL
  • TTIC-Toyota Technological Institute Chicago

Résultats de recherche: Contribution à un journalArticle de conférenceRevue par des pairs

Résumé

Synchronous mini-batch SGD is state-of-the-art for large-scale distributed machine learning. However, in practice, its convergence is bottlenecked by slow communication rounds between worker nodes. A natural solution to reduce communication is to use the “local-SGD” model in which the workers train their model independently and synchronize every once in a while. This algorithm improves the computation-communication trade-off but its convergence is not understood very well. We propose a non-asymptotic error analysis, which enables comparison to one-shot averaging i.e., a single communication round among independent workers, and mini-batch averaging i.e., communicating at every step. We also provide adaptive lower bounds on the communication frequency for large step-sizes (t-a, a ? (1/2, 1)) and show that local-SGD reduces communication by a factor of O (Pv3T/2 ), with T the total number of gradients and P machines.

langue originaleAnglais
journalAdvances in Neural Information Processing Systems
Volume32
étatPublié - 1 janv. 2019
Evénement33rd Annual Conference on Neural Information Processing Systems, NeurIPS 2019 - Vancouver, Canada
Durée: 8 déc. 201914 déc. 2019

Empreinte digitale

Examiner les sujets de recherche de « Communication trade-offs for Local-SGD with large step size ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation