Goto

Collaborating Authors

 Country



b6af2c9703f203a2794be03d443af2e3-Paper.pdf

Neural Information Processing Systems

In this work, we combine these observations to assess whether such trainable, transferrable subnetworks exist in pre-trained BERT models. For a range of downstream tasks, we indeed find matching subnetworks at 40% to 90% sparsity.


Towards Efficient Pre-Trained Language Model via Feature Correlation Distillation

Neural Information Processing Systems

Therefore, a series of attempts Chung et al. [2020], Wu et al. [2020], Wang et al. [2020c], Gordon et al. [2020a], Tang et al. [2019], Aguilar et al. [2019] have been made to review the techniques for effective




Lower Bounds and Nearly Optimal Algorithms in Distributed Learning with Communication Compression Xinmeng Huang 1 Yiming Chen 2,3 Wotao Yin

Neural Information Processing Systems

Recent advances in distributed optimization and learning have shown that communication compression is one of the most effective means of reducing communication. While there have been many results for convergence rates with compressed communication, a lower bound is still missing. Analyses of algorithms with communication compression have identified two abstract properties that guarantee convergence: the unbiased property or the contrac-tive property.