Goto

Collaborating Authors

 Genre


Gated Multimodal Units for Information Fusion

arXiv.org Machine Learning

This paper presents a novel model for multimodal learning based on gated neural networks. The Gated Multimodal Unit (GMU) model is intended to be used as an internal unit in a neural network architecture whose purpose is to find an intermediate representation based on a combination of data from different modalities. The GMU learns to decide how modalities influence the activation of the unit using multiplicative gates. It was evaluated on a multilabel scenario for genre classification of movies using the plot and the poster. The GMU improved the macro f-score performance of single-modality approaches and outperformed other fusion strategies, including mixture of experts models. Along with this work, the MM-IMDb dataset is released which, to the best of our knowledge, is the largest publicly available multimodal dataset for genre prediction on movies.


On SGD's Failure in Practice: Characterizing and Overcoming Stalling

arXiv.org Machine Learning

Stochastic Gradient Descent (SGD) is widely used in machine learning problems to efficiently perform empirical risk minimization, yet, in practice, SGD is known to stall before reaching the actual minimizer of the empirical risk. SGD stalling has often been attributed to its sensitivity to the conditioning of the problem; however, as we demonstrate, SGD will stall even when applied to a simple linear regression problem with unity condition number for standard learning rates. Thus, in this work, we numerically demonstrate and mathematically argue that stalling is a crippling and generic limitation of SGD and its variants in practice. Once we have established the problem of stalling, we generalize an existing framework for hedging against its effects, which (1) deters SGD and its variants from stalling, (2) still provides convergence guarantees, and (3) makes SGD and its variants more practical methods for minimization.


Model-based Classification and Novelty Detection For Point Pattern Data

arXiv.org Machine Learning

Point patterns are sets or multi-sets of unordered elements that can be found in numerous data sources. However, in data analysis tasks such as classification and novelty detection, appropriate statistical models for point pattern data have not received much attention. This paper proposes the modelling of point pattern data via random finite sets (RFS). In particular, we propose appropriate likelihood functions, and a maximum likelihood estimator for learning a tractable family of RFS models. In novelty detection, we propose novel ranking functions based on RFS models, which substantially improve performance.


Latent Sequence Decompositions

arXiv.org Machine Learning

Sequence-to-sequence models rely on a fixed decomposition of the target sequences into a sequence of tokens that may be words, word-pieces or characters. The choice of these tokens and the decomposition of the target sequences into a sequence of tokens is often static, and independent of the input, output data domains. This can potentially lead to a sub-optimal choice of token dictionaries, as the decomposition is not informed by the particular problem being solved. In this paper we present Latent Sequence Decompositions (LSD), a framework in which the decomposition of sequences into constituent tokens is learnt during the training of the model. The decomposition depends both on the input sequence and on the output sequence. In LSD, during training, the model samples decompositions incrementally, from left to right by locally sampling between valid extensions.


Lower Bounds on Active Learning for Graphical Model Selection

arXiv.org Machine Learning

We consider the problem of estimating the underlying graph associated with a Markov random field, with the added twist that the decoding algorithm can iteratively choose which subsets of nodes to sample based on the previous samples, resulting in an active learning setting. Considering both Ising and Gaussian models, we provide algorithm-independent lower bounds for high-probability recovery within the class of degree-bounded graphs. Our main results are minimax lower bounds for the active setting that match the best known lower bounds for the passive setting, which in turn are known to be tight in several cases of interest. Our analysis is based on Fano's inequality, along with novel mutual information bounds for the active learning setting, and the application of restricted graph ensembles. While we consider ensembles that are similar or identical to those used in the passive setting, we require different analysis techniques, with a key challenge being bounding a mutual information quantity associated with observed subsets of nodes, as opposed to full observations.


Learning sign language could give you super vision

Daily Mail - Science & tech

Researchers have found that learning sign language can be beneficial for hearing adults too, giving them faster reaction times in their peripheral vision. Improved peripheral vision is useful in many sports and for driving, making you more alert to changes in your peripheral field of vision. The research also found that deaf adults have far better peripheral vision and reaction times than both hearing adults and hearing adults who use sign language. Researchers at the University of Sheffield have found that learning sign language can be beneficial for hearing adults too, giving them faster reaction times in their peripheral vision. The research, conducted at the University of Sheffield's Academic Unit of Opthalmology, found that adults learning a visual-spatial language such as British sign language (BSL) had a positive impact on their visual field response.


When Things Go Missing

The New Yorker

A couple of years ago, I spent the summer in Portland, Oregon, losing things. I normally live on the East Coast, but that year, unable to face another sweltering August, I decided to temporarily decamp to the West. This turned out to be strangely easy. I'd lived in Portland for a while after college, and some acquaintances there needed a house sitter. Another friend was away for the summer and happy to loan me her pickup truck. Someone on Craigslist sold me a bike for next to nothing. In very short order, and with very little effort, everything fell into place. And then, mystifyingly, everything fell out of place. My first day in town, I left the keys to the truck on the counter of a coffee shop. The next day, I left the keys to the house in the front door. A few days after that, warming up in the midday sun at an outdoor café, I took off the long-sleeved shirt I'd been wearing, only to leave it hanging over the back of the chair when I headed home. When I returned to claim it, I discovered that I'd left my wallet behind as well. Prior to that summer, I should note, I had lost a wallet exactly once in my adult life: at gunpoint. Yet later that afternoon I stopped by a sporting-goods store to buy a lock for my new bike and left my wallet sitting next to the cash register.


How do you model that?

#artificialintelligence

Attend Multilevel Modeling of Hierarchical and Longitudinal Data Using SAS and learn how to identify complex and dynamic patterns within your multilevel data. This advanced class provides a conceptual understanding of multilevel linear models (MLM) and multilevel generalized linear models (MGLM). Meet the Presenters Catherine Truxillo and Chris Daman discuss what you can expect to learn in this class. Attend a public course or enjoy the classroom experience right at your desktop, the choice is yours!


Machine Learning Crash Course: Part 3 · ML@B

#artificialintelligence

How someone might identify a dog. Important inputs that are given a lot of weight are highlighted in red. Notice how the neurons are organized into layers, where the further right the neurons are, the more abstract the input? In other words, the neurons on the left ask questions about general shapes and lines, whereas the neurons on the right ask questions about objects such as eyes or fur. Trained neural networks function in a very similar way, although they arrive at this conclusion after training with a lot of data.


Wearable AI Detects Tone Of Conversation To Make It Navigable (And Nicer) For All

Forbes - Tech

A Samsung Simband displays real-time results on conversational narrative and tone. In the past few years, wearables have offered to track many things, to predict illness, and even to give us rudimentary advice on staying healthy. The makers of a new device hope to expand the role that wearables can play in helping with day-to-day life, however, by adding'conversational wing-person' and'social coach' to their list of skills. Researchers from MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) and Institute of Medical Engineering and Science (IMES) have developed programming to help those for whom conversation is difficult to navigate it with ease, and they've put it all in a wearable for real-time assistance. According to the team, the results of their study, "Predicting Latent Narrative Mood using Audio and Physiologic Data" [PDF], suggest that using such technology to pin down the tone of conversation as it happens is nearly within our reach--a potential boon for persons who experience anxiety, aspects of autism spectrum disorder, or other conditions that can make chewing the fat an intimidating prospect.