Goto

Collaborating Authors

 Statistical Learning






Export Reviews, Discussions, Author Feedback and Meta-Reviews

Neural Information Processing Systems

First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. One of the most common reasons for using Markov chain Monte Carlo (MCMC) is to estimate the value of an otherwise intractable integral. Typically MCMC algorithms will give an exact answer as the number of iterations increases to infinity. However, this gives little assurance about the precision of the estimate in finite samples. This paper addresses this important issue from a theoretical point of view for the Hamiltonian Monte Carlo (HMC) algorithm, an algorithm which has been receiving a fair amount of recent attention.


Scalable Kernel Methods via Doubly Stochastic Gradients

Neural Information Processing Systems

The general perception is that kernel methods are not scalable, so neural nets become the choice for large-scale nonlinear learning problems. Have we tried hard enough for kernel methods? In this paper, we propose an approach that scales up kernel methods using a novel concept called doubly stochastic functional gradients''. Based on the fact that many kernel methods can be expressed as convex optimization problems, our approach solves the optimization problems by making two unbiased stochastic approximations to the functional gradient---one using random training points and another using random features associated with the kernel---and performing descent steps with this noisy functional gradient. Our algorithm is simple, need no commit to a preset number of random features, and allows the flexibility of the function class to grow as we see more incoming data in the streaming setting.





Supplementary Material for Submission ID8000: Learnability with Indirect Supervision Signals

Neural Information Processing Systems

T H is weak VC-major with dimension d < . Then, H is T -learnable. We need several intermediate results to prove this. The first lemma bounds the empirical risk via the averaged Rademacher complexity. Work done while at the Allen Institute for AI and at the University of Illinois at Urbana-Champaign.