Goto

Collaborating Authors

 Statistical Learning



Beating SGD Saturation with T ail-A veraging and Minibatching

Neural Information Processing Systems

Stochastic gradient descent (SGD) provides a simple and yet stunningly efficient way to solve a broad range of machine learning problems.



Convolutional Networks on Graphs for Learning Molecular Fingerprints

Neural Information Processing Systems

We introduce a convolutional neural network that operates directly on graphs. These networks allow end-to-end learning of prediction pipelines whose inputs are graphs of arbitrary size and shape. The architecture we present generalizes standard molecular feature extraction methods based on circular fingerprints. We show that these data-driven features are more interpretable, and have better predictive performance on a variety of tasks.




p(A | Z) = ฮพ null

Neural Information Processing Systems

We sincerely thank all reviewers for their valuable comments. Our responses to the comments are listed below. If we understand it correctly, "the approximate posterior in If so, we do consider the effects of the KL term in the proof. Does Claim 4.1 rely on a specific distribution (q1 under As for Gaussian variables (the case in VGAE), we are not sure if the claim still holds. We agree sometimes it might be hard to follow and we shall add the details in the revision.