Goto

Collaborating Authors

 Africa


Sub-LinearMemory: HowtoMakePerformersSLiM

Neural Information Processing Systems

Recent works proposed various linear self-attention mechanisms, scaling only asO(L)for serial computation. We conduct a thorough complexity analysis of Performers,aclass which includes most recent linear Transformer mechanisms.






Network-to-NetworkRegularization: Enforcing Occam'sRazortoImproveGeneralization

Neural Information Processing Systems

What makes a classifier have the ability to generalize? There have been a lot of important attempts to address this question, but a clear answer is still elusive.





31784d9fc1fa0d25d04eae50ac9bf787-Paper.pdf

Neural Information Processing Systems

Indeedin learning applications, where symmetric tensors areformed from statistical moments (higher-order covariances) or multivariate derivatives (higher-order Hessians), CP decomposition has enabled parameter estimation for mixtures of Gaussians [20, 35], generalized linear models [34], shallow neuralnetworks[19,24,42],deepernetworks[17,18,30],hiddenMarkovmodels[5],amongothers.