Goto

Collaborating Authors

 Country


DistributedDeepLearningInOpenCollaborations

Neural Information Processing Systems

Wedemonstratetheeffectiveness of our approach for SwAV and ALBERT pretraining in realisticconditions and achieve performance comparable to traditional setups at a fraction of the cost.




ProfileEntropy: AFundamentalMeasureforthe LearnabilityandCompressibilityofDistributions

Neural Information Processing Systems

The profile of a sample is the multiset of its symbol frequencies. We show that for samples of discrete distributions, profile entropy is a fundamental measure unifying the concepts of estimation, inference, and compression.




Initialization-Dependent Sample Complexity of Linear Predictors and Neural Networks

Neural Information Processing Systems

Clearly, in order for learning to be possible, we must impose some constraints on the size of the function class. One possibility is to bound the number of parameters (i.e., the dimensions of the matrix W), in which case learnability follows from standard VC-dimension or covering number arguments (see Anthony and Bartlett [1999]).