Goto

Collaborating Authors

 Statistical Learning


Deep Model Transferability from Attribution Maps

Neural Information Processing Systems

Unlike the seminal work of taskonomy that relies on a large number of annotations as supervision and is thus computationally cumbersome, the proposed approach requires no human annotations and imposes no constraints on the architectures of the networks.



Incremental Few-Shot Learning with Attention Attractor Networks

Neural Information Processing Systems

After learning the novel classes, the model is then evaluated on the overall classification performance on both base and novel classes. To this end, we propose a meta-learning model, the Attention Attractor Network, which regularizes the learning of novel classes. In each episode, we train a set of new weights to recognize novel classes until they converge, and we show that the technique of recurrent back-propagation can back-propagate through the optimization process and facilitate the learning of these parameters.


Copula-like Variational Inference

Neural Information Processing Systems

This paper considers a new family of variational distributions motivated by Sklar's theorem. This family is based on new copula-like densities on the hypercube with non-uniform marginals which can be sampled efficiently, i.e. with a complexity linear in the dimension d of the state space. Then, the proposed variational densities that we suggest can be seen as arising from these copula-like densities used as base distributions on the hypercube with Gaussian quantile functions and sparse rotation matrices as normalizing flows. The latter correspond to a rotation of the marginals with complexity O (d log d) . We provide some empirical evidence that such a variational family can also approximate non-Gaussian posteriors and can be beneficial compared to Gaussian approximations. Our method performs largely comparably to state-of-the-art variational approximations on standard regression and classification benchmarks for Bayesian Neural Networks.



Predicting the Politics of an Image Using Webly Supervised Data

Neural Information Processing Systems

We collect a dataset of over one million unique images and associated news articles from left-and right-leaning news sources, and develop a method to predict the image's political leaning. This problem is particularly challenging because of the enormous intra-class visual and semantic diversity of our data. We propose a two-stage method to tackle this problem. In the first stage, the model is forced to learn relevant visual concepts that, when joined with document embeddings computed from articles paired with the images, enable the model to predict bias. In the second stage, we remove the requirement of the text domain and train a visual classifier from the features of the former model. We show this two-stage approach facilitates learning and outperforms several strong baselines.




Sampled Softmax with Random Fourier Features

Neural Information Processing Systems

Motivated by our analysis and the work on kernel-based sampling, we propose the Random F ourier Softmax (RF-softmax) method that utilizes the powerful Random Fourier Features to enable more efficient and accurate sampling from an approximate softmax distribution. We show that RF-softmax leads to low bias in estimation in terms of both the full softmax distribution and the full softmax gradient.


Hierarchical Optimal Transport for Multimodal Distribution Alignment

Neural Information Processing Systems

OT and constrain the solution space. Here, we leverage the fact that heterogeneous datasets often admit clustered or multi-subspace structure to improve OT -based distribution alignment.