R-Drop: Regularized Dropout for Neural Networks

Oct-10-2024, 14:09:34 GMT–Neural Information Processing Systems

Dropout is a powerful and widely used technique to regularize the training of deep neural networks. Though effective and performing well, the randomness introduced by dropout causes unnegligible inconsistency between training and inference. In this paper, we introduce a simple consistency training strategy to regularize dropout, namely R-Drop, which forces the output distributions of different sub models generated by dropout to be consistent with each other. Specifically, for each training sample, R-Drop minimizes the bidirectional KL-divergence between the output distributions of two sub models sampled by dropout. Theoretical analysis reveals that R-Drop reduces the above inconsistency. Experiments on \bf{5} widely used deep learning tasks ( \bf{18} datasets in total), including neural machine translation, abstractive summarization, language understanding, language modeling, and image classification, show that R-Drop is universally effective.

neural network, r-drop, regularized dropout, (6 more...)

Neural Information Processing Systems

Oct-10-2024, 14:09:34 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language > Machine Translation (0.99)
  - Machine Learning > Neural Networks
    - Deep Learning (0.61)