Deep Learning
Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation
Garbacea, Cristina, Carton, Samuel, Yan, Shiyan, Mei, Qiaozhu
Recent advances in deep learning have resulted in a resurgence in the popularity of natural language generation (NLG). Many deep learning based models, including recurrent neural networks and generative adversarial networks, have been proposed and applied to generating various types of text. Despite the fast development of methods, how to better evaluate the quality of these natural language generators remains a significant challenge. We conduct an in-depth empirical study to evaluate the existing evaluation methods for natural language generation. We compare human-based evaluators with a variety of automated evaluation procedures, including discriminative evaluators that measure how well the generated text can be distinguished from human-written text, as well as text overlap metrics that measure how similar the generated text is to human-written references. We measure to what extent these different evaluators agree on the ranking of a dozen of state-of-the-art generators for online product reviews. We find that human evaluators do not correlate well with discriminative evaluators, leaving a bigger question of whether adversarial accuracy is the correct objective for natural language generation. In general, distinguishing machine-generated text is a challenging task even for human evaluators, and their decisions tend to correlate better with text overlap metrics. We also find that diversity is an intriguing metric that is indicative of the assessments of different evaluators.
Elimination of All Bad Local Minima in Deep Learning
Kawaguchi, Kenji, Kaelbling, Leslie Pack
In this paper, we theoretically prove that we can eliminate all suboptimal local minima by adding one neuron per output unit to any deep neural network, for multi-class classification, binary classification, and regression with an arbitrary loss function. At every local minimum of any deep neural network with added neurons, the set of parameters of the original neural network (without added neurons) is guaranteed to be a global minimum of the original neural network. The effects of the added neurons are proven to automatically vanish at every local minimum. Unlike many related results in the literature, our theoretical results are directly applicable to common deep learning tasks because the results only rely on the assumptions that automatically hold in the common tasks. Moreover, we discuss several limitations in eliminating the suboptimal local minima in this manner by providing additional theoretical results and several examples.
Introducing Neuromodulation in Deep Neural Networks to Learn Adaptive Behaviours
Vecoven, Nicolas, Ernst, Damien, Wehenkel, Antoine, Drion, Guillaume
In this paper, we propose a new deep neural network architecture, called NMD net, that has been specifically designed to learn adaptive behaviours. This architecture exploits a biological mechanism called neuromodulation that sustains adaptation in biological organisms. This architecture has been introduced in a deep-reinforcement learning architecture for interacting with Markov decision processes in a meta-reinforcement learning setting where the action space is continuous. The deep-reinforcement learning architecture is trained using an advantage actor-critic algorithm. Experiments are carried on several test problems. Results show that the neural network architecture with neuromodulation provides significantly better results than state-of-the-art recurrent neural networks which do not exploit this mechanism.
Multi-Label Adversarial Perturbations
Song, Qingquan, Jin, Haifeng, Huang, Xiao, Hu, Xia
Adversarial examples are delicately perturbed inputs, which aim to mislead machine learning models towards incorrect outputs. While most of the existing work focuses on generating adversarial perturbations in multi-class classification problems, many real-world applications fall into the multi-label setting in which one instance could be associated with more than one label. For example, a spammer may generate adversarial spams with malicious advertising while maintaining the other labels such as topic labels unchanged. To analyze the vulnerability and robustness of multi-label learning models, we investigate the generation of multi-label adversarial perturbations. This is a challenging task due to the uncertain number of positive labels associated with one instance, as well as the fact that multiple labels are usually not mutually exclusive with each other. To bridge this gap, in this paper, we propose a general attacking framework targeting on multi-label classification problem and conduct a premier analysis on the perturbations for deep neural networks. Leveraging the ranking relationships among labels, we further design a ranking-based framework to attack multi-label ranking algorithms. We specify the connection between the two proposed frameworks and separately design two specific methods grounded on each of them to generate targeted multi-label perturbations. Experiments on real-world multi-label image classification and ranking problems demonstrate the effectiveness of our proposed frameworks and provide insights of the vulnerability of multi-label deep learning models under diverse targeted attacking strategies. Several interesting findings including an unpolished defensive strategy, which could potentially enhance the interpretability and robustness of multi-label deep learning models, are further presented and discussed at the end.
Evolutionary Construction of Convolutional Neural Networks
van Knippenberg, Marijn, Menkovski, Vlado, Consoli, Sergio
Neuro-Evolution is a field of study that has recently gained significantly increased traction in the deep learning community. It combines deep neural networks and evolutionary algorithms to improve and/or automate the construction of neural networks. Recent Neuro-Evolution approaches have shown promising results, rivaling hand-crafted neural networks in terms of accuracy. A two-step approach is introduced where a convolutional autoencoder is created that efficiently compresses the input data in the first step, and a convolutional neural network is created to classify the compressed data in the second step. The creation of networks in both steps is guided by by an evolutionary process, where new networks are constantly being generated by mutating members of a collection of existing networks. Additionally, a method is introduced that considers the trade-off between compression and information loss of different convolutional autoencoders. This is used to select the optimal convolutional autoencoder from among those evolved to compress the data for the second step. The complete framework is implemented, tested on the popular CIFAR-10 data set, and the results are discussed. Finally, a number of possible directions for future work with this particular framework in mind are considered, including opportunities to improve its efficiency and its application in particular areas.
Performance of Three Slim Variants of The Long Short-Term Memory (LSTM) Layer
The Long Short-Term Memory (LSTM) layer is an important advancement in the field of neural networks and machine learning, allowing for effective training and impressive inference performance. LSTM-based neural networks have been successfully employed in various applications such as speech processing and language translation. The LSTM layer can be simplified by removing certain components, potentially speeding up training and runtime with limited change in performance. In particular, the recently introduced variants, called SLIM LSTMs, have shown success in initial experiments to support this view. Here, we perform computational analysis of the validation accuracy of a convolutional plus recurrent neural network architecture using comparatively the standard LSTM and three SLIM LSTM layers. We have found that some realizations of the SLIM LSTM layers can potentially perform as well as the standard LSTM layer for our considered architecture.
How to get started with machine learning on graphs โ Octavian โ Medium
Since our talk at Connected Data London, I've spoken to a lot of research teams who have graph data and want to perform machine learning on it, but are not sure where to start. In this article, I'll share resources and approaches to get started with machine learning on graphs. From talking with research teams, it's really clear how broad and pervasive graph data is -- from disease detection, genetics and healthcare to banking and engineering, graphs are emerging as a powerful analysis paradigm for hard problems. Simply put, a graph is a collection of nodes (e.g. Fatima is a friend of Jacob).
The 'Godfather of Deep Learning' on Why We Need to Ensure AI Doesn't Just Benefit the Rich
Martin Ford made waves with his 2015 book, Rise of the Robots, which details the many accelerating trends in automation and how they're slated to impact business and, especially, employment. For his next book, Architects of Intelligence: The Truth About AI from the People Building It, he, well, attempts to hone in on precisely what that subtitle describes. It's stuffed with in-depth interviews with the biggest names in AI. One of those is Geoffrey Hinton. Currently a professor of computer science at the University of Toronto and a part of the Google Brain project, Hinton is considered by many in his field to be the'godfather of deep learning,' due to his pioneering work in artificial neural networks.
Early detection of pediatric epilepsy possible through 'deep learning' technique
Early detection of the most common form of epilepsy in children is possible through "deep learning," a new machine learning tool that teaches computers to learn by example, according to a new study that includes researchers from Georgia State University. Most BECT patients self-heal in puberty, but the disease can cause verbal dysfunction, attention deficit and language impairment in 18 to 25 percent of patients. Studies have shown that drug treatment could improve language skill and normalize centrotemporal spikes in electroencephalograph (EEG) tests, so it's important to distinguish epilepsy patients from healthy people. While studies have found that magnetic resonance imaging (MRI) and functional magnetic resonance imaging (fMRI) are promising for differentiating BECT patients from healthy people, these imaging techniques are mainly based on a doctor's knowledge and diagnostic ability, and hence also have limitations such as low accuracy. Few studies have focused on developing machine learning methods that can recognize BECT patients.
Creating your own style transfer mirror with Gradient and ml5.js
In this post, we will learn how to train a style transfer network with Paperspace's Gradient and use the model in ml5.js to create an interactive style transfer mirror. This post is the second on a series of blog posts dedicated to train machine learning models in Paperspace and then use them in ml5.js. You can read the first post in this series on how to train a LSTM network to generate text here. Style Transfer is the technique of recomposing images in the style of other images.1 It first appeared in September 2015, when Gatys et.