Deep Learning
OpenAI Releases Text Generator AI That Was Too "Dangerous" To Share
OpenAI, the AI research lab has finally published the GPT2 -- the text generating AI tool which the lab once said was too "dangerous" to share. In a blog post, OpenAI said that despite the arguments of GPT-2 potential in creating synthetic propaganda, fake news, and online phishing campaigns, "we've seen no strong evidence of misuse so far" Back in February, OpenAI announced the GPT2, a language model based upon 1.5 billion parameters and trained by analyzing over 8 million web pages. The main objective of GPT2 is to create coherent text from a few words. The text generating AI tool can be used for many tasks such as translation, chatbots, coming up with unprecedented answers and more. But citing concerns that it could be used for malicious intent, the company withheld the release of the full version.
Understanding Artificial Intelligence, Machine Learning, and Deep Learning
Technological change is the only constant in today's business world, disrupting everything from large organizations to small start-ups. Disruption affects everyone, but will you be the disruptor or the disrupted? You must pay close attention to the Hard Trends shaping the future of your industry, your business, and the outside world to identify opportunities used to innovate and grow rapidly, additionally using those Hard Trends to solve any problems your organization and customers might have before they occur. The shared definition and understanding of the words we use is an issue in business. While several companies are on course to use artificial intelligence (AI), machine learning (ML), and deep learning (DL), others hardly understand the fundamental differences between these powerful technologies. How can one be successful, much less disruptive, when they themselves do not differentiate between AI, ML, and DL? Recently, technology company Sage conducted surveys pertaining to AI and individuals' understanding of it.
Periodic Spectral Ergodicity: A Complexity Measure for Deep Neural Networks and Neural Architecture Search
Sรผzen, Mehmet, Cerdร , J. J., Weber, Cornelius
Establishing associations between the structure and the learning ability of deep neural networks (DNNs) is a challenging task in modern machine learning. Producing solutions to this challenge will bring progress both in the theoretical understanding of DNNs and in building new architectures efficiently. In this work, we address this challenge by developing a new simple complexity measure based on another new measure called Periodic Spectral Ergodicity (PSE) originating from quantum statistical mechanics. Based on this measure a framework is devised in quantifying the complexity of deep neural network from its learned weights and traversing network connectivity in a sequential manner, hence the term cascading PSE (cPSE) as an empirical complexity measure. Because of this cascading approach, i.e., a symmetric divergence of PSE on the consecutive layers, it is possible to use this measure in addition for Neural Architecture Search (NAS). We demonstrate the usefulness of this measure in practice on two sets of vision models, ResNet and VGG and sketch the computation of cPSE for more complex network structures.
Syntax-Infused Transformer and BERT models for Machine Translation and Natural Language Understanding
Sundararaman, Dhanasekar, Subramanian, Vivek, Wang, Guoyin, Si, Shijing, Shen, Dinghan, Wang, Dong, Carin, Lawrence
Attention-based models have shown significant improvement over traditional algorithms in several NLP tasks. The Transformer, for instance, is an illustrative example that generates abstract representations of tokens inputted to an encoder based on their relationships to all tokens in a sequence. Recent studies have shown that although such models are capable of learning syntactic features purely by seeing examples, explicitly feeding this information to deep learning models can significantly enhance their performance. Leveraging syntactic information like part of speech (POS) may be particularly beneficial in limited training data settings for complex models such as the Transformer. We show that the syntax-infused Transformer with multiple features achieves an improvement of 0.7 BLEU when trained on the full WMT '14 English to German translation dataset and a maximum improvement of 1.99 BLEU points when trained on a fraction of the dataset. In addition, we find that the incorporation of syntax into BERT fine-tuning outperforms baseline on a number of downstream tasks from the GLUE benchmark. Introduction Attention-based deep learning models for natural language processing (NLP) have shown promise for a variety of machine translation and natural language understanding tasks. For word-level, sequence-to-sequence tasks such as translation, paraphrasing, and text summarization, attention-based models allow a single token ( e.g., a word or subword) in a sequence to be represented as a combination of all tokens in the sequence (Luong, Pham, and Manning, 2015). The distributed context allows attention-based models to infer rich representations for tokens, leading to more robust performance.
Meta Label Correction for Learning with Weak Supervision
Zheng, Guoqing, Awadallah, Ahmed Hassan, Dumais, Susan
Leveraging weak or noisy supervision for building effective machine learning models has long been an important research problem. The growing need for large-scale datasets to train deep learning models has increased its importance. Weak or noisy supervision could originate from multiple sources including non-expert annotators or automatic labeling based on heuristics or user interaction signals. Previous work on modeling and correcting weak labels have been focused on various aspects, including loss correction, training instance re-weighting, etc. In this paper, we approach this problem from a novel perspective based on meta-learning. We view the label correction procedure as a meta-process and propose a new meta-learning based framework termed MLC for learning with weak supervision. Experiments with different label noise levels on multiple datasets show that MLC can achieve large improvement over previous methods incorporating weak labels for learning.
XceptionTime: A Novel Deep Architecture based on Depthwise Separable Convolutions for Hand Gesture Classification
Rahimian, Elahe, Zabihi, Soheil, Atashzar, Seyed Farokh, Asif, Amir, Mohammadi, Arash
Capitalizing on the need for addressing the existing challenges associated with gesture recognition via sparse multichannel surface Electromyography (sEMG) signals, the paper proposes a novel deep learning model, referred to as the XceptionTime architecture. The proposed innovative XceptionTime is designed by integration of depthwise separable convolutions, adaptive average pooling, and a novel non-linear normalization technique. At the heart of the proposed architecture is several XceptionTime modules concatenated in series fashion designed to capture both temporal and spatial information-bearing contents of the sparse multichannel sEMG signals without the need for data augmentation and/or manual design of feature extraction. In addition, through integration of adaptive average pooling, Conv1D, and the non-linear normalization approach, XceptionTime is less prone to overfitting, more robust to temporal translation of the input, and more importantly is independent from the input window size. Finally, by utilizing the depthwise separable convolutions, the XceptionTime network has far fewer parameters resulting in a less complex network. The performance of XceptionTime is tested on a sub Ninapro dataset, DB1, and the results showed a superior performance in comparison to any existing counterparts. In this regard, 5:71% accuracy improvement, on a window size 200ms, is reported in this paper, for the first time.
Adaptive versus Standard Descent Methods and Robustness Against Adversarial Examples
Since this phenomenon was first observed, researchers have attempted to develop methods which produce models that are robust to adversarial perturbations under specific attack models (Wong and Kolter (2018); Sinha et al. (2018); Raghunathan et al. (2018); Mirman et al. (2018); Madry et al. (2018); Zhang et al. (2019)). As machine learning proliferates into society, including security-critical settings like health care (Esteva et al. (2017)) or autonomous vehicles (Codevilla et al. (2018)), it is crucial to develop methods that allow us to understand the vulnerability of our models and design appropriate countermeasures. Additionally there is a growing literature on the theory of adversarial examples. Many of these results attempt to understand adversarial examples by constructing examples of learning problems for which it is difficult to construct a classifier that is robust to adversarial perturbations. This difficultly may arise due to sample complexity (Schmidt et al. (2018)), computational constraints (Bubeck et al. (2019); Degwekar et al. (2019)), or the high-dimensional geometry of the initial feature space (Shafahi et al. (2019); Khoury and Hadfield-Menell (2018)).
Empirical validation of network learning with taxi GPS data from Wuhan, China
Xu, Susan Jia, Xie, Qian, Chow, Joseph Y. J., Liu, Xintao
Many studies have illustrated the import ance to accurately and precisely measure the attributes of an urban transport system. Due to the rise of Big Data and Internet of Things, there are numerous machine learning methods to measur e attributes of the transport system . Chow ( 1) provides an overview of these techniques including several appl ications like Allahviranloo and Recker ( 2) for activity pattern prediction; Cai et al. ( 3) for short - term traffic forecasting; Luque - Baena et al. ( 4) for vehicle detection; Lv et al. ( 5) for t raffic flow prediction; and Ma et al. ( 6) for network congestion prediction. However, generic machine learning techniques are not specifically designed to exploit the unique structure of urban transport networks. As a result, in recent years a theory of inverse problems (see 7) have emerged to capture network structure, dubbed " inverse transportation problems " by Xu et al. ( 8).
Information Bottleneck Methods on Convolutional Neural Networks
Recent year, many researches attempt to open the black box of deep neural networks and propose a various of theories to understand it. Among them, information bottleneck theory (IB) claims that there are two distinct phases consisting of fitting phase and compression phase in the course of training. This statement attracts many attentions since its success in explaining the inner behavior of feedforward neural networks. In this paper, we employ IB theory to understand the dynamic behavior of convolutional neural networks (CNNs) and investigate how the fundamental features have impact on the performance of CNNs. In particular, through a series of experimental analysis on benchmark of MNIST and Fashion-MNIST, we demonstrate that the compression phase is not observed in all these cases. This show us the CNNs have a rather complicated behavior than feedforward neural networks.
L-FGADMM: Layer-Wise Federated Group ADMM for Communication Efficient Decentralized Deep Learning
Elgabli, Anis, Park, Jihong, Ahmed, Sabbir, Bennis, Mehdi
--This article proposes a communication-efficient decentralized deep learning algorithm, coined layer-wise federated group ADMM (L-FGADMM). T o minimize an empirical risk, every worker in L-FGADMM periodically communicates with two neighbors, in which the periods are separately adjusted for different layers of its deep neural network. A constrained optimization problem for this setting is formulated and solved using the stochastic version of GADMM proposed in our prior work. Numerical evaluations show that by less frequently exchanging the largest layer, L-FGADMM can significantly reduce the communication cost, without compromising the convergence speed. Surprisingly, despite less exchanged information and decentralized operations, intermittently skipping the largest layer consensus in L-FGADMM creates a regularizing effect, thereby achieving the test accuracy as high as federated learning (FL), a baseline method with the entire layer consensus by the aid of a central entity.