Goto

Collaborating Authors

 Statistical Learning




Stein Variational Gradient Descent with Matrix-Valued Kernels

Neural Information Processing Systems

On the other hand, standard SVGD only uses the first order gradient information, and can not leverage the advantage of the second order methods, such as Newton's method and natural gradient, to achieve better performance on challenging problems with complex loss landscapes or domains.






Copula Multi-label Learning

Neural Information Processing Systems

Unfortunately, the statistical properties of existing multi-label dependency modelings are still not well understood.



Implicit Regularization for Optimal Sparse Recovery

Neural Information Processing Systems

Hence the total running cost is O(nd), which is the cost to store/read the data in/from memory. These results attest that there are regimes where optimal methods for sparse linear regression exist.