Goto

Collaborating Authors

 Deep Learning


Multi-task learning of daily work and study round-trips from survey data

arXiv.org Machine Learning

In this study, we present a machine learning approach to infer the worker and student mobility flows on daily basis from static censuses. The rapid urbanization has made the estimation of the human mobility flows a critical task for transportation and urban planners. The primary objective of this paper is to complete individuals' census data with working and studying trips, allowing its merging with other mobility data to better estimate the complete origin-destination matrices. Worker and student mobility flows are among the most weekly regular displacements and consequently generate road congestion problems. Estimating their round-trips eases the decision-making processes for local authorities. Worker and student censuses often contain home location, work places and educational institutions. We thus propose a neural network model that learns the temporal distribution of displacements from other mobility sources and tries to predict them on new censuses data. The inclusion of multi-task learning in our neural network results in a significant error rate control in comparison to single task learning.


Learning to Decompose and Disentangle Representations for Video Prediction

arXiv.org Machine Learning

Our goal is to predict future video frames given a sequence of input frames. Despite large amounts of video data, this remains a challenging task because of the high-dimensionality of video frames. We address this challenge by proposing the Decompositional Disentangled Predictive Auto-Encoder (DDPAE), a framework that combines structured probabilistic models and deep networks to automatically (i) decompose the high-dimensional video that we aim to predict into components, and (ii) disentangle each component to have low-dimensional temporal dynamics that are easier to predict. Crucially, with an appropriately specified generative model of video frames, our DDPAE is able to learn both the latent decomposition and disentanglement without explicit supervision. For the Moving MNIST dataset, we show that DDPAE is able to recover the underlying components (individual digits) and disentanglement (appearance and location) as we would intuitively do. We further demonstrate that DDPAE can be applied to the Bouncing Balls dataset involving complex interactions between multiple objects to predict the video frame directly from the pixels and recover physical states without explicit supervision.


Pressure Predictions of Turbine Blades with Deep Learning

arXiv.org Machine Learning

Deep learning has been used in many areas, such as feature detections in images and the game of go. This paper presents a study that attempts to use the deep learning method to predict turbomachinery performance. Three different deep neural networks are built and trained to predict the pressure distributions of turbine airfoils. The performance of a library of turbine airfoils were firstly predicted using methods based on Euler equations, which were then used to train and validate the deep learning neural networks. The results show that network with four layers of convolutional neural network and two layers of fully connected neural network provides the best predictions. For the best neural network architecture, the pressure prediction on more than 99% locations are better than 3% and 90% locations are better than 1%.


Deep Learning for Classification Tasks on Geospatial Vector Polygons

arXiv.org Machine Learning

The ability to analyse vector shapes of geospatial objects is useful for many tasks, such as quality assessment or enrichment of map data (Fan et al, 2014) or the classification of topographical objects (Keyes and Winstanley, 1999). An increasingly more common method for shape analysis is through machine learning. For example, machine learning can be applied to assess correct building types (Xu et al, 2017) or classify road sections (Andrรกลกik and Bรญl, 2016). The prediction of house prices (Montero et al, 2018) and the estimation of pedestrian side walk widths (Brezina et al, 2017) are tasks that could possibly also benefit from the application of machine learning analysis on geometric shapes. Current machine learning methods applied to geospatial vector data rely on extracting information from a geometry that characterizes its shape. This preprocessing step is known in machine learning as feature extraction (LeCun et al, 2015, 438) or feature engineering (Domingos, 2012, 84).


DropBack: Continuous Pruning During Training

arXiv.org Machine Learning

We introduce a technique that compresses deep neural networks both during and after training by constraining the total number of weights updated during backpropagation to those with the highest total gradients. The remaining weights are forgotten and their initial value is regenerated at every access to avoid storing them in memory. This dramatically reduces the number of off-chip memory accesses during both training and inference, a key component of the energy needs of DNN accelerators. By ensuring that the total weight diffusion remains close to that of baseline unpruned SGD, networks pruned using DropBack are able to maintain high accuracy across network architectures. We observe weight compression of 25x with LeNet-300-100 on MNIST while maintaining accuracy. On CIFAR-10, we see an approximately 5x weight compression on 3 models: an already 9x-reduced VGG-16, Densenet, and WRN-28-10 - all with zero or negligible accuracy loss. On Densenet and WRN, which are particularly challenging to compress, Both Densenet and WRN improve on the state of the art, achieving higher compression with better accuracy than prior pruning techniques.


Straight to the Tree: Constituency Parsing with Neural Syntactic Distance

arXiv.org Artificial Intelligence

In this work, we propose a novel constituency parsing scheme. The model predicts a vector of real-valued scalars, named syntactic distances, for each split position in the input sentence. The syntactic distances specify the order in which the split points will be selected, recursively partitioning the input, in a top-down fashion. Compared to traditional shift-reduce parsing schemes, our approach is free from the potential problem of compounding errors, while being faster and easier to parallelize. Our model achieves competitive performance amongst single model, discriminative parsers in the PTB dataset and outperforms previous models in the CTB dataset.


Meta Continual Learning

arXiv.org Artificial Intelligence

Using neural networks in practical settings would benefit from the ability of the networks to learn new tasks throughout their lifetimes without forgetting the previous tasks. This ability is limited in the current deep neural networks by a problem called catastrophic forgetting, where training on new tasks tends to severely degrade performance on previous tasks. One way to lessen the impact of the forgetting problem is to constrain parameters that are important to previous tasks to stay close to the optimal parameters. Recently, multiple competitive approaches for computing the importance of the parameters with respect to the previous tasks have been presented. In this paper, we propose a learning to optimize algorithm for mitigating catastrophic forgetting. Instead of trying to formulate a new constraint function ourselves, we propose to train another neural network to predict parameter update steps that respect the importance of parameters to the previous tasks. In the proposed meta-training scheme, the update predictor is trained to minimize loss on a combination of current and past tasks. We show experimentally that the proposed approach works in the continual learning setting.


Navigating with Graph Representations for Fast and Scalable Decoding of Neural Language Models

arXiv.org Artificial Intelligence

Neural language models (NLMs) have recently gained a renewed interest by achieving state-of-the-art performance across many natural language processing (NLP) tasks. However, NLMs are very computationally demanding largely due to the computational cost of the softmax layer over a large vocabulary. We observe that, in decoding of many NLP tasks, only the probabilities of the top-K hypotheses need to be calculated preciously and K is often much smaller than the vocabulary size. This paper proposes a novel softmax layer approximation algorithm, called Fast Graph Decoder (FGD), which quickly identifies, for a given context, a set of K words that are most likely to occur according to a NLM. We demonstrate that FGD reduces the decoding time by an order of magnitude while attaining close to the full softmax baseline accuracy on neural machine translation and language modeling tasks. We also prove the theoretical guarantee on the softmax approximation quality.


Accurate and Robust Neural Networks for Security Related Applications Exampled by Face Morphing Attacks

arXiv.org Artificial Intelligence

Artificial neural networks tend to learn only what they need for a task. A manipulation of the training data can counter this phenomenon. In this paper, we study the effect of different alterations of the training data, which limit the amount and position of information that is available for the decision making. We analyze the accuracy and robustness against semantic and black box attacks on the networks that were trained on different training data modifications for the particular example of morphing attacks. A morphing attack is an attack on a biometric facial recognition system where the system is fooled to match two different individuals with the same synthetic face image. Such a synthetic image can be created by aligning and blending images of the two individuals that should be matched with this image.


Defense Against the Dark Arts: An overview of adversarial example security research and future research directions

arXiv.org Artificial Intelligence

This article presents a summary of a keynote lecture at the Deep Learning Security workshop at IEEE Security and Privacy 2018. This lecture summarizes the state of the art in defenses against adversarial examples and provides recommendations for future research directions on this topic. "I.I.D." stands for "independent and identically distributed". It means that all of the examples in the training and test set are generated independently from each other, and are all drawn from the same data-generating distribution. This diagram illustrates this with an example training set and test set sampled for a classification problem with 2 input features (one plotted on horizontal axis, one plotted on vertical axis) and 2 classes (orange plus versus blue X). 2 ML reached "human-level performance" on many IID tasks circa 2013 (Goodfellow 2018) Figure 2: Until recently, machine learning was difficult, even in the I.I.D. setting. Adversarial examples were not interesting to most researchers because mistakes were the rule, not the exception. In about 2013, machine learning started to reach human-level performance on several benchmark tasks (here I highlight vision tasks because they have nice pictures to put on a slide).