Genre
Does Neural Machine Translation Benefit from Larger Context?
Jean, Sebastien, Lauly, Stanislas, Firat, Orhan, Cho, Kyunghyun
We propose a neural machine translation architecture that models the surrounding text in addition to the source sentence. These models lead to better performance, both in terms of general translation quality and pronoun prediction, when trained on small corpora, although this improvement largely disappears when trained with a larger corpus. We also discover that attention-based neural machine translation is well suited for pronoun prediction and compares favorably with other approaches that were specifically designed for this task.
Statistical inference for high dimensional regression via Constrained Lasso
In this paper, we propose a new method for estimation and constructing confidence intervals for low-dimensional components in a high-dimensional model. The proposed estimator, called Constrained Lasso (CLasso) estimator, is obtained by simultaneously solving two estimating equations---one imposing a zero-bias constraint for the low-dimensional parameter and the other forming an $\ell_1$-penalized procedure for the high-dimensional nuisance parameter. By carefully choosing the zero-bias constraint, the resulting estimator of the low dimensional parameter is shown to admit an asymptotically normal limit attaining the Cram\'{e}r-Rao lower bound in a semiparametric sense. We propose a tuning-free iterative algorithm for implementing the CLasso. We show that when the algorithm is initialized at the Lasso estimator, the de-sparsified estimator proposed in van de Geer et al. [\emph{Ann. Statist.} {\bf 42} (2014) 1166--1202] is asymptotically equivalent to the first iterate of the algorithm. We analyse the asymptotic properties of the CLasso estimator and show the globally linear convergence of the algorithm. We also demonstrate encouraging empirical performance of the CLasso through numerical studies.
Visual-Inertial-Semantic Scene Representation for 3-D Object Detection
Dong, Jingming, Fei, Xiaohan, Soatto, Stefano
We describe a system to detect objects in three-dimensional space using video and inertial sensors (accelerometer and gyrometer), ubiquitous in modern mobile platforms from phones to drones. Inertials afford the ability to impose class-specific scale priors for objects, and provide a global orientation reference. A minimal sufficient representation, the posterior of semantic (identity) and syntactic (pose) attributes of objects in space, can be decomposed into a geometric term, which can be maintained by a localization-and-mapping filter, and a likelihood function, which can be approximated by a discriminatively-trained convolutional neural network. The resulting system can process the video stream causally in real time, and provides a representation of objects in the scene that is persistent: Confidence in the presence of objects grows with evidence, and objects previously seen are kept in memory even when temporarily occluded, with their return into view automatically predicted to prime re-detection.
A Joint Speaker-Listener-Reinforcer Model for Referring Expressions
Yu, Licheng, Tan, Hao, Bansal, Mohit, Berg, Tamara L.
Referring expressions are natural language constructions used to identify particular objects within a scene. In this paper, we propose a unified framework for the tasks of referring expression comprehension and generation. Our model is composed of three modules: speaker, listener, and reinforcer. The speaker generates referring expressions, the listener comprehends referring expressions, and the reinforcer introduces a reward function to guide sampling of more discriminative expressions. The listener-speaker modules are trained jointly in an end-to-end learning framework, allowing the modules to be aware of one another during learning while also benefiting from the discriminative reinforcer's feedback. We demonstrate that this unified framework and training achieves state-of-the-art results for both comprehension and generation on three referring expression datasets. Project and demo page: https://vision.cs.unc.edu/refer
Error Asymmetry in Causal and Anticausal Regression
Blöbaum, Patrick, Washio, Takashi, Shimizu, Shohei
It is generally difficult to make any statements about the expected prediction error in an univariate setting without further knowledge about how the data were generated. Recent work showed that knowledge about the real underlying causal structure of a data generation process has implications for various machine learning settings. Assuming an additive noise and an independence between data generating mechanism and its input, we draw a novel connection between the intrinsic causal relationship of two variables and the expected prediction error. We formulate the theorem that the expected error of the true data generating function as prediction model is generally smaller when the effect is predicted from its cause and, on the contrary, greater when the cause is predicted from its effect. The theorem implies an asymmetry in the error depending on the prediction direction. This is further corroborated with empirical evaluations in artificial and real-world data sets.
Variational Hamiltonian Monte Carlo via Score Matching
Zhang, Cheng, Shahbaba, Babak, Zhao, Hongkai
Traditionally, the field of computational Bayesian statistics has been divided into two main subfields: variational methods and Markov chain Monte Carlo (MCMC). In recent years, however, several methods have been proposed based on combining variational Bayesian inference and MCMC simulation in order to improve their overall accuracy and computational efficiency. This marriage of fast evaluation and flexible approximation provides a promising means of designing scalable Bayesian inference methods. In this paper, we explore the possibility of incorporating variational approximation into a state-of-the-art MCMC method, Hamiltonian Monte Carlo (HMC), to reduce the required gradient computation in the simulation of Hamiltonian flow, which is the bottleneck for many applications of HMC in big data problems. To this end, we use a {\it free-form} approximation induced by a fast and flexible surrogate function based on single-hidden layer feedforward neural networks. The surrogate provides sufficiently accurate approximation while allowing for fast exploration of parameter space, resulting in an efficient approximate inference algorithm. We demonstrate the advantages of our method on both synthetic and real data problems.
Attend, Adapt and Transfer: Attentive Deep Architecture for Adaptive Transfer from multiple sources in the same domain
Rajendran, Janarthanan, Lakshminarayanan, Aravind S., Khapra, Mitesh M., Prasanna, P, Ravindran, Balaraman
Transferring knowledge from prior source tasks in solving a new target task can be useful in several learning applications. The application of transfer poses two serious challenges which have not been adequately addressed. First, the agent should be able to avoid negative transfer, which happens when the transfer hampers or slows down the learning instead of helping it. Second, the agent should be able to selectively transfer, which is the ability to select and transfer from different and multiple source tasks for different parts of the state space of the target task. We propose A2T (Attend, Adapt and Transfer), an attentive deep architecture which adapts and transfers from these source tasks. Our model is generic enough to effect transfer of either policies or value functions. Empirical evaluations on different learning algorithms show that A2T is an effective architecture for transfer by being able to avoid negative transfer while transferring selectively from multiple source tasks in the same domain.
Robert Taylor, A Pioneer Of Modern Computing And The Internet, Dies At 85
Nearly 50 years ago, computer visionary Robert Taylor helped lay the foundations for what we know today as the internet. Taylor, who had Parkinson's disease, died Thursday at his home in Woodside, Calif., his son Kurt Taylor tells NPR. Like many of his peers who helped build the internet, Bob Taylor, as he was known, wasn't a computer scientist. The University of Texas at Austin graduate had a background in psychology and mathematics. Taylor was inspired by the idea of expanding human interaction using computer technology, Guy Raz noted in an interview profiling Taylor in 2009.
Machine-Learning with Renthop
Contributed by David Letzler, Kyle Gallatin and Christopher Capozzola They enrolled in the NYC Data Science Academy 12-week full time Data Science Bootcamp program taking place between January 9th, 2017 and March 31st, 2017. The original article can be found here. For this project, we took on the Two Sigma Connect: Rental Listing Inquiries Challenge on Kaggle. The rental website Renthop provided us with a csv of data from 120,000 listings and asked us to produce a model to predict whether a given listing would receive "low," "medium," or "high" interest. The model would be judged by predicting a test set, with the log-loss formula determining its effectiveness.
Computer pioneer Robert W. Taylor dies at 85
WOODSIDE, CALIFORNIA – Robert W. Taylor, who was instrumental in creating the internet and the modern personal computer, has died. Taylor, who had Parkinson's disease, died Thursday at his home in the San Francisco Peninsula community of Woodside, his son, Kurt Taylor, told the Los Angeles Times and the New York Times. In 1961, Taylor was a project manager for NASA when he directed funding to Douglas Engelbart at the Stanford Research Institute, who helped develop the modern computer mouse. Taylor was working for the Pentagon's Advanced Research Projects Agency in 1966 when he shepherded the creation of a single computer network to link ARPA-sponsored researchers at companies and institutions around the country. Taylor was frustrated that he had to use three separate terminals to communicate with the researchers through their computer systems.