Deep Learning
Global-to-local Memory Pointer Networks for Task-Oriented Dialogue
Wu, Chien-Sheng, Socher, Richard, Xiong, Caiming
End-to-end task-oriented dialogue is challenging since knowledge bases are usually large, dynamic and hard to incorporate into a learning framework. We propose the global-to-local memory pointer (GLMP) networks to address this issue. In our model, a global memory encoder and a local memory decoder are proposed to share external knowledge. The encoder encodes dialogue history, modifies global contextual representation, and generates a global memory pointer. The decoder first generates a sketch response with unfilled slots. Next, it passes the global memory pointer to filter the external knowledge for relevant information, then instantiates the slots via the local memory pointers. We empirically show that our model can improve copy accuracy and mitigate the common out-of-vocabulary problem. As a result, GLMP is able to improve over the previous state-of-the-art models in both simulated bAbI Dialogue dataset and human-human Stanford Multi-domain Dialogue dataset on automatic and human evaluation.
Sensorimotor learning for artificial body perception
Diez-Valencia, German, Ohashi, Takuya, Lanillos, Pablo, Cheng, Gordon
The great challenge was to generalize the reconstruction of the arm for any background without using segmentation. For that purpose, several background images were synthetically generated and were overlaid by automated labelled masks (i.e., boolean mask of the arm in the visual field) by means of background subtraction (Figure 1(b)). An example of the results of the generated arm given the a joint angle configuration is shown in Figure 1(c). The most right generated image shows difficulties of the model to properly reconstruct the robot arm when the majority of it is outside the field of view. Anyhow, the statistical evaluation of the network, over all experiments, showed an accuracy of 84.4% when comparing the matching between the original versus the generated image mask.
Exploiting Synchronized Lyrics And Vocal Features For Music Emotion Detection
Parisi, Loreto, Francia, Simone, Olivastri, Silvio, Tavella, Maria Stella
Support Vector Machines are employed engaging playlists according to sentiment and with good results also for multilabel classification [30], emotions. While previous works were mostly based more recently also Convolutional Neural Networks were on audio for music discovery and playlists generation, used in this field [45]. Lyrics-based approaches, on the we take advantage of our synchronized lyrics dataset other hand, make use of Recurrent Neural Networks architectures to combine text representations and music features in (like LSTM [13]) for performing text classification a novel way; we therefore introduce the Synchronized [46, 47]. The idea of using lyrics combined with Lyrics Emotion Dataset. Unlike other approaches that voice only audio signals is done in [29], where emotion randomly exploited the audio samples and the whole recognition is performed by using textual and speech data, text, our data is split according to the temporal information instead of visual ones. Measuring and assigning emotions provided by the synchronization between lyrics to music is not a straightforward task: the sentiment/mood and audio. This work shows a comparison between associated with a song can be derived by a combination of text-based and audio-based deep learning classification many features, moreover, emotions expressed by a musical models using different techniques from Natural Language excerpt and by its corresponding lyrics do not always Processing and Music Information Retrieval domains.
Energy-Efficient Thermal Comfort Control in Smart Buildings via Deep Reinforcement Learning
Gao, Guanyu, Li, Jie, Wen, Yonggang
Heating, Ventilation, and Air Conditioning (HVAC) is extremely energy-consuming, accounting for 40% of total building energy consumption. Therefore, it is crucial to design some energy-efficient building thermal control policies which can reduce the energy consumption of HVAC while maintaining the comfort of the occupants. However, implementing such a policy is challenging, because it involves various influencing factors in a building environment, which are usually hard to model and may be different from case to case. To address this challenge, we propose a deep reinforcement learning based framework for energy optimization and thermal comfort control in smart buildings. We formulate the building thermal control as a cost-minimization problem which jointly considers the energy consumption of HVAC and the thermal comfort of the occupants. To solve the problem, we first adopt a deep neural network based approach for predicting the occupants' thermal comfort, and then adopt Deep Deterministic Policy Gradients (DDPG) for learning the thermal control policy. To evaluate the performance, we implement a building thermal control simulation system and evaluate the performance under various settings. The experiment results show that our method can improve the thermal comfort prediction accuracy, and reduce the energy consumption of HVAC while improving the occupants' thermal comfort.
The Limitations of Adversarial Training and the Blind-Spot Attack
Zhang, Huan, Chen, Hongge, Song, Zhao, Boning, Duane, Dhillon, Inderjit S., Hsieh, Cho-Jui
The adversarial training procedure proposed by Madry et al. (2018) is one of the most effective methods to defend against adversarial examples in deep neural networks (DNNs). In our paper, we shed some lights on the practicality and the hardness of adversarial training by showing that the effectiveness (robustness on test set) of adversarial training has a strong correlation with the distance between a test point and the manifold of training data embedded by the network. Test examples that are relatively far away from this manifold are more likely to be vulnerable to adversarial attacks. Consequentially, an adversarial training based defense is susceptible to a new class of attacks, the "blind-spot attack", where the input images reside in "blind-spots" (low density regions) of the empirical distribution of training data but is still on the ground-truth data manifold. For MNIST, we found that these blind-spots can be easily found by simply scaling and shifting image pixel values. Most importantly, for large datasets with high dimensional and complex data manifold (CIFAR, ImageNet, etc), the existence of blind-spots in adversarial training makes defending on any valid test examples difficult due to the curse of dimensionality and the scarcity of training data. Additionally, we find that blind-spots also exist on provable defenses including (Wong & Kolter, 2018) and (Sinha et al., 2018) because these trainable robustness certificates can only be practically optimized on a limited set of training data.
Fast Deep Learning for Automatic Modulation Classification
Ramjee, Sharan, Ju, Shengtai, Yang, Diyu, Liu, Xiaoyu, Gamal, Aly El, Eldar, Yonina C.
In this work, we investigate the feasibility and effectiveness of employing deep learning algorithms for automatic recognition of the modulation type of received wireless communication signals from subsampled data. Recent work considered a GNU radio-based data set that mimics the imperfections in a real wireless channel and uses 10 different modulation types. A Convolutional Neural Network (CNN) architecture was then developed and shown to achieve performance that exceeds that of expert-based approaches. Here, we continue this line of work and investigate deep neural network architectures that deliver high classification accuracy. We identify three architectures - namely, a Convolutional Long Short-term Deep Neural Network (CLDNN), a Long Short-Term Memory neural network (LSTM), and a deep Residual Network (ResNet) - that lead to typical classification accuracy values around 90% at high SNR. We then study algorithms to reduce the training time by minimizing the size of the training data set, while incurring a minimal loss in classification accuracy. To this end, we demonstrate the performance of Principal Component Analysis in significantly reducing the training time, while maintaining good performance at low SNR. We also investigate subsampling techniques that further reduce the training time, and pave the way for online classification at high SNR. Finally, we identify representative SNR values for training each of the candidate architectures, and consequently, realize drastic reductions of the training time, with negligible loss in classification accuracy.
Memory Augmented Deep Generative models for Forecasting the Next Shot Location in Tennis
Fernando, Tharindu, Denman, Simon, Sridharan, Sridha, Fookes, Clinton
Considering the fact that present day ball speeds exceed 130mph, the time required by the receiver to make a decision regarding the opponents' intention, and initiate a response could exceed the flight time for the ball [1], [2], [3], [4]. Several studies have shown that this reactive ability is the product of pattern recognition skills that are obtained through a "biological probabilistic engine", that derives theories regardingopponents intentions with the partial information available[1], [5], [6]. For instance, it has been shown that expert tennis players are better at detecting events in advance [1], [7] and posses better knowledge/ expertise of situational probabilities [3]. Further investigation of human neurological structures have revealed that those capabilities occur due to a bottom-up computational process [1] within the human brain, from sensory memory to the experiences stored in episodic memory [8], [9] and knowledge derived in semantic memory [9], [10]. Despite the growing interest among researchers in the machine learning domain in better understanding factors influencing decision making in fastball sports, there have been very few studies transferring the observations of the underlying neural mechanisms to neural modelling in machine learning.Current state-of-the-art methodologies try to capture the underlying semantics through a handful of handcrafted features, without paying attention to essential mechanisms in the human brain, where the expertise and observations are stored and knowledge is derived.
AI Pipeline - bringing AI to you. End-to-end integration of data, algorithms and deployment tools
de Prado, Miguel, Su, Jing, Dahyot, Rozenn, Saeed, Rabia, Keller, Lorenzo, Vallez, Noelia
Next generation of embedded Information and Communication Technology (ICT) systems are interconnected collaborative intelligent systems able to perform autonomous tasks. Training and deployment of such systems on Edge devices however require a fine-grained integration of data and tools to achieve high accuracy and overcome functional and non-functional requirements. In this work, we present a modular AI pipeline as an integrating framework to bring data, algorithms and deployment tools together. By these means, we are able to interconnect the different entities or stages of particular systems and provide an end-to-end development of AI products. We demonstrate the effectiveness of the AI pipeline by solving an Automatic Speech Recognition challenge and we show that all the steps leading to an end-to-end development for Key-word Spotting tasks: importing, partitioning and pre-processing of speech data, training of different neural network architectures and their deployment on heterogeneous embedded platforms.
Deep Fusion: An Attention Guided Factorized Bilinear Pooling for Audio-video Emotion Recognition
Zhang, Yuanyuan, Wang, Zi-Rui, Du, Jun
Automatic emotion recognition (AER) is a challenging task due to the abstract concept and multiple expressions of emotion. Although there is no consensus on a definition, human emotional states usually can be apperceived by auditory and visual systems. Inspired by this cognitive process in human beings, it's natural to simultaneously utilize audio and visual information in AER. However, most traditional fusion approaches only build a linear paradigm, such as feature concatenation and multi-system fusion, which hardly captures complex association between audio and video. In this paper, we introduce factorized bilinear pooling (FBP) to deeply integrate the features of audio and video. Specifically, the features are selected through the embedded attention mechanism from respective modalities to obtain the emotion-related regions. The whole pipeline can be completed in a neural network. Validated on the AFEW database of the audio-video sub-challenge in EmotiW2018, the proposed approach achieves an accuracy of 62.48%, outperforming the state-of-the-art result.
Practical Lossless Compression with Latent Variables using Bits Back Coding
Townsend, James, Bird, Tom, Barber, David
Deep latent variable models have seen recent success in many data domains. Lossless compression is an application of these models which, despite having the potential to be highly useful, has yet to be implemented in a practical manner. We present `Bits Back with ANS' (BB-ANS), a scheme to perform lossless compression with latent variable models at a near optimal rate. We demonstrate this scheme by using it to compress the MNIST dataset with a variational auto-encoder model (VAE), achieving compression rates superior to standard methods with only a simple VAE. Given that the scheme is highly amenable to parallelization, we conclude that with a sufficiently high quality generative model this scheme could be used to achieve substantial improvements in compression rate with acceptable running time. We make our implementation available open source at https://github.com/bits-back/bits-back .