Goto

Collaborating Authors

 Deep Learning


Inferring Molecular Pathology and micro-RNA Transcriptome from mRNA Profiles of Cancer Biopsies through Deep Multi-Task Learning

arXiv.org Machine Learning

Despite great advances, molecular cancer pathology is often limited to use a small number of biomarkers rather than the whole transcriptome, partly due to the computational challenges. Here, we introduce a novel architecture of DNNs that is capable of simultaneous inference of various properties of biological samples, through multi-task and transfer learning. We employed this architecture on mRNA transcription profiles of 10787 clinical samples from 34 classes (one healthy and 33 different types of cancer) from 27 tissues. Our system significantly outperforms prior works and classical machine learning approaches in predicting tissue-of-origin, normal or disease state and cancer type of each sample. Furthermore, it can predict miRNA transcription profile of each sample, which enables performing miRNA expression research when only mRNA transcriptome data are available. We also show this system is very robust against noise and missing values. Collectively, our results highlight applications of artificial intelligence in molecular cancer pathology and oncological research.


Semi-Supervised Feature Learning for Off-Line Writer Identifications

arXiv.org Machine Learning

Conventional approaches used supervised learning to estimate off-line writer identifications. In this study, we improved the off-line writer identifica- tions by semi-supervised feature learning pipeline, which trained the extra unla- beled data and the original labeled data simultaneously. In specific, we proposed a weighted label smoothing regularization (WLSR) method, which assigned the weighted uniform label distribution to the extra unlabeled data. We regularized the convolutional neural network (CNN) baseline, which allows learning more discriminative features to represent the properties of different writing styles. Based on experiments on ICDAR2013, CVL and IAM benchmark datasets, our results showed that semi-supervised feature learning improved the baseline meas- urement and achieved better performance compared with existing writer identifications approaches.


Rethinking Numerical Representations for Deep Neural Networks

arXiv.org Machine Learning

With ever-increasing computational demand for deep learning, it is critical to investigate the implications of the numeric representation and precision of DNN model weights and activations on computational efficiency. In this work, we explore unconventional narrow-precision floating-point representations as it relates to inference accuracy and efficiency to steer the improved design of future DNN platforms. We show that inference using these custom numeric representations on production-grade DNNs, including GoogLeNet and VGG, achieves an average speedup of 7.6x with less than 1% degradation in inference accuracy relative to a state-of-the-art baseline platform representing the most sophisticated hardware using single-precision floating point. To facilitate the use of such customized precision, we also present a novel technique that drastically reduces the time required to derive the optimal precision configuration.


Importance of the Mathematical Foundations of Machine Learning Methods for Scientific and Engineering Applications

arXiv.org Machine Learning

There has been a lot of recent interest in adopting machine learning methods for scientific and engineering applications. This has in large part been inspired by recent successes and advances in the domains of Natural Language Processing (NLP) and Image Classification (IC). However, scientific and engineering problems have their own unique characteristics and requirements raising new challenges for effective design and deployment of machine learning approaches. There is a strong need for further mathematical developments on the foundations of machine learning methods to increase the level of rigor of employed methods and to ensure more reliable and interpretable results. Also as reported in the recent literature on state-of-the-art results and indicated by the No Free Lunch Theorems of statistical learning theory incorporating some form of inductive bias and domain knowledge is essential to success. Consequently, even for existing and widely used methods there is a strong need for further mathematical work to facilitate ways to incorporate prior scientific knowledge and related inductive biases into learning frameworks and algorithms. We briefly discuss these topics and discuss some ideas proceeding in this direction.


Data augmentation using synthetic data for time series classification with deep residual networks

arXiv.org Artificial Intelligence

Data augmentation in deep neural networks is the process of generating artificial data in order to reduce the variance of the classifier with the goal to reduce the number of errors. This idea has been shown to improve deep neural network's generalization capabilities in many computer vision tasks such as image recognition and object localization. Apart from these applications, deep Convolutional Neural Networks (CNNs) have also recently gained popularity in the Time Series Classification (TSC) community. However, unlike in image recognition problems, data augmentation techniques have not yet been investigated thoroughly for the TSC task. This is surprising as the accuracy of deep learning models for TSC could potentially be improved, especially for small datasets that exhibit overfitting, when a data augmentation method is adopted. In this paper, we fill this gap by investigating the application of a recently proposed data augmentation technique based on the Dynamic Time Warping distance, for a deep learning model for TSC. To evaluate the potential of augmenting the training set, we performed extensive experiments using the UCR TSC benchmark. Our preliminary experiments reveal that data augmentation can drastically increase deep CNN's accuracy on some datasets and significantly improve the deep model's accuracy when the method is used in an ensemble approach.


End-to-end Speech Recognition with Word-based RNN Language Models

arXiv.org Artificial Intelligence

ABSTRACT This paper investigates the impact of word-based RNN language models (RNN-LMs) on the performance of end-to-end automatic speech recognition (ASR). In our prior work, we have proposed a multilevel LM, in which character-based and word-based RNN-LMs are combined in hybrid CTC/attention-based ASR. Although this multilevel approach achieves significant error reduction in the Wall Street Journal (WSJ) task, two different LMs need to be trained and used for decoding, which increase the computational cost and memory usage. In this paper, we further propose a novel wordbased RNN-LM, which allows us to decode with only the wordbased LM, where it provides look-ahead word probabilities to predict next characters instead of the character-based LM, leading competitive accuracy with less computation compared to the multilevel LM. We demonstrate the efficacy of the word-based RNN-LMs using a larger corpus, LibriSpeech, in addition to WSJ we used in the prior work. Furthermore, we show that the proposed model achieves 5.1 %WER for WSJ Eval'92 test set when the vocabulary size is increased, which is the best WER reported for end-to-end ASR systems on this benchmark. Index Terms-- End-to-end speech recognition, language modeling, decoding, connectionist temporal classification, attention decoder 1. INTRODUCTION Automatic speech recognition (ASR) is currently a mature set of widely-deployed technologies that enable successful user interface applications such as voice search [1]. However, current systems lean heavily on the scaffolding of complicated legacy architectures that grew up around traditional techniques, including hidden Markov models (HMMs), Gaussian mixture models (GMMs), hybrid HMM/deep neural network (DNN) systems, and sequence discriminative training methods [2].


L-Shapley and C-Shapley: Efficient Model Interpretation for Structured Data

arXiv.org Machine Learning

We study instancewise feature importance scoring as a method for model interpretation. Any such method yields, for each predicted instance, a vector of importance scores associated with the feature vector. Methods based on the Shapley score have been proposed as a fair way of computing feature attributions of this kind, but incur an exponential complexity in the number of features. This combinatorial explosion arises from the definition of the Shapley value and prevents these methods from being scalable to large data sets and complex models. We focus on settings in which the data have a graph structure, and the contribution of features to the target variable is well-approximated by a graph-structured factorization. In such settings, we develop two algorithms with linear complexity for instancewise feature importance scoring. We establish the relationship of our methods to the Shapley value and another closely related concept known as the Myerson value from cooperative game theory. We demonstrate on both language and image data that our algorithms compare favorably with other methods for model interpretation.


OpenAI bots thrash team of Dota 2 semi-pros, set eyes on mega-tourney

#artificialintelligence

The human team โ€“ made up of popular Twitch streamers and former professionals ranked in the 99.95th percentile โ€“ hunkered down to play against the bots known as OpenAI Five in San Francisco on Sunday. OpenAI Five smashed its opponents, winning comfortably in two out of three games. It did lose one game, however, after spectators watching the match live and on Twitch were allowed to pick the pool of heroes โ€“ the playable characters in the game. Each hero comes with its own strengths and weaknesses and picking a balanced combination is paramount to winning. If you have too many characters for the same role, the other team will steamroll you.


Meet Fetch, the AI creation of DeepMind pioneers who refuse to succumb to Google

#artificialintelligence

Google battled Facebook to acquire artificial intelligence lab DeepMind in 2014 for a reported $660 million - but not everyone was happy with the move. "How in the world can you build a truly neutral artificial general intelligence under that company?" questions Humayun Sheikh, an early investor who helped London-based DeepMind commercialise the AI technology that Google now owns. Sheikh, who personally donated to DeepMind founder Demis Hassabis during the technology's very early stages, cut ties with the firm when it was sold to Google four years ago. DeepMind software lead Toby Simpson, creator of the "Creatures" game series in the 1990s, left the company. He believed the Google acquisition...


Google DeepMind's AI program learns human navigation skills

#artificialintelligence

Notch up another win for the robots: the latest program from Google's artificial intelligence group, DeepMind, has trounced experts at a maze game after it learned to find its way around like a human. Scientists noticed that when they trained the AI to move through a landscape, it spontaneously developed electrical activity akin to that seen in the specialised brain cells that underpin human navigational skills. So-called'grid cells' were only identified in animals in 2005 in work that earned researchers a Nobel prize. The latest breakthrough reveals the potential for human brain-like activity to emerge from scratch in AI systems. Beyond making smarter programs, it paves the way for computer engineers to build models that help neuroscientists better understand the human brain.