Learning to Combat Noisy Labels via Classification Margins

Lin, Jason Z., Bradic, Jelena

arXiv.org Machine Learning 

In recent years, deep neural networks have emerged as the state-of-the-art method for classification in numerous domains, in particular computer vision (Krizhevsky et al., 2012; Zeiler & Fergus, 2014; Simonyan & Zisserman, 2015; Szegedy et al., 2015; He et al., 2016). However, a large-scale dataset with clean annotations, such as ImageNet (Deng et al., 2009), is usually required for the networks to show superior performance. In practice, obtaining such a large amount of data with manually verified labels is prohibitively laborious or expensive. On the other hand, various crowdsourcing platforms make it possible (and cheaper) to collect large amounts of annotated data, even though the labels may be incorrect. This calls for robust methods that can still effectively learn from the copious data in the presence of noisy labels. When the training set contains noisy labels, it has been observed in Zhang et al. (2017); Arpit et al. (2017) that deep neural networks exhibit the early learning phenomenon (Liu et al., 2020), in which the clean instances are fit to the network before the noisy ones are (over-)fit to the network. On the theoretical front, it is proved in Liu et al. (2020) that even simple linear models in high dimensions exhibit such behavior as well. This intriguing phenomenon, as a double-edged sword, begs the question: Is there a way to make use of the discriminating power the network learns during the early stage of the training, in order to prevent itself from memorizing the noisy labels at the later stage of the training? In this paper, we propose MARVEL (MARgins Via Early Learning), a method that attempts to exclude noisy instances from future training by leveraging this early learning phenomenon.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found