QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding

Alistarh, Dan; Grubic, Demjan; Li, Jerry; Tomioka, Ryota; Vojnović, Milan

doi:10.48550/arxiv.1610.02132

preprintarXiv (Cornell University)Oct 7, 2016GREEN OA

QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding

DADan Alistarh DGDemjan Grubic JLJerry Li RTRyota Tomioka MVMilan Vojnović

Indexed inarxivdatacite

Abstract

Parallel implementations of stochastic gradient descent (SGD) have received significant research attention, thanks to excellent scalability properties of this algorithm, and to its efficiency in the context of training deep neural networks. A fundamental barrier for parallelizing large-scale SGD is the fact that the cost of communicating the gradient updates between nodes can be very large. Consequently, lossy compression heuristics have been proposed, by which nodes only communicate quantized gradients. Although effective in practice, these heuristics do not always provably converge, and it is not clear whether they are optimal. In this paper, we propose Quantized SGD (QSGD), a family of compression schemes…

Citation impact

910

total citations

FWCI: —
Percentile: —
References: 0

Citations per year

Authors

5

Topics & keywords

Topics

Keywords

Computer science
Stochastic gradient descent
Scalability
Quantization (signal processing)
Heuristics
Lossy compression
Gradient descent
Artificial neural network

No related works found for this paper.