Goto

Collaborating Authors

 Technology


GeneralizedJensen-ShannonDivergenceLoss forLearningwithNoisyLabels

Neural Information Processing Systems

Based on this observation, we adopt ageneralized version ofthe JensenShannon divergence for multiple distributions to encourage consistency around data points. Using this loss function, we show state-of-the-art results on both synthetic(CIFAR),andreal-world(e.g.WebVision)noisewithvaryingnoiserates.



46f76a4bda9a9579eab38a8f6eabcda1-AuthorFeedback.pdf

Neural Information Processing Systems

For`, we construct the systems of hyperrectangles by4 first precomputing anapproximate k-NN distance estimate using Ball Trees foreachdata point, andthen clustering5 the top-q densest data points intoT partitions using the k-means algorithm, where we binary search for the optimal6 parameterq.






d0ffb35aaa7faa894afe5060c694d674-Paper-Conference.pdf

Neural Information Processing Systems

For example, in image object recognition and language models building[4,5,6,7,8,9], the number of classes scales as the number of possible objects or the dictionary size respectively.