Layer-wise Adaptive Gradient Sparsification for Distributed Deep Learning with Convergence Guarantees

Open in new window