Goto

Collaborating Authors

 Deep Learning


Convolutional Neural Network Quantization using Generalized Gamma Distribution

arXiv.org Artificial Intelligence

As edge applications using convolutional neural networks (CNN) models grow, it is becoming necessary to introduce dedicated hardware accelerators in which network parameters and feature-map data are represented with limited precision. In this paper we propose a novel quantization algorithm for energy-efficient deployment of the hardware accelerators. For weights and biases, the optimal bit length of the fractional part is determined so that the quantization error is minimized over their distribution. For feature-map data, meanwhile, their sample distribution is well approximated with the generalized gamma distribution (GGD), and accordingly the optimal quantization step size can be obtained through the asymptotical closed form solution of GGD. The proposed quantization algorithm has a higher signal-to-quantization-noise ratio (SQNR) than other quantization schemes previously proposed for CNNs, and even can be more improved by tuning the quantization parameters, resulting in efficient implementation of the hardware accelerators for CNNs in terms of power consumption and memory bandwidth.


Taking Human out of Learning Applications: A Survey on Automated Machine Learning

arXiv.org Artificial Intelligence

Machine learning techniques have deeply rooted in our everyday life. However, since it is knowledge- and labor-intensive to pursuit good learning performance, human experts are heavily engaged in every aspect of machine learning. In order to make machine learning techniques easier to apply and reduce the demand for experienced human experts, automatic machine learning~(AutoML) has emerged as a hot topic of both in industry and academy. In this paper, we provide a survey on existing AutoML works. First, we introduce and define the AutoML problem, with inspiration from both realms of automation and machine learning. Then, we propose a general AutoML framework that not only covers almost all existing approaches but also guides the design for new methods. Afterward, we categorize and review the existing works from two aspects, i.e., the problem setup and the employed techniques. Finally, we provide a detailed analysis of AutoML approaches and explain the reasons underneath their successful applications. We hope this survey can serve as not only an insightful guideline for AutoML beginners but also an inspiration for future researches.


Consistency-based anomaly detection with adaptive multiple-hypotheses predictions

arXiv.org Artificial Intelligence

In out-of-distribution classification tasks, only some classes - the normal cases - can be modeled with data, whereas the variation of all possible anomalies is too large to be described sufficiently by samples. Thus, the widespread discriminative approaches cannot cover such learning tasks and rather generative models, which attempt to learn the input density of the ordinary cases, are used. However, generative models suffer under a large input dimensionality (as in images) and are typically inefficient learners. Motivated by the Local-Outlier-Factor (LOF) method, in this work, we propose to allow the network to directly estimate the local density functions since, for the detection of outliers, the local neighborhood is more important than the global one. At the same time, we retain consistency in the sense that the model must not support areas of the input space that are not covered by samples. Our method allows the model to identify out-of-distribution samples reliably. For the anomaly detection task on CIFAR-10, our ConAD model results in up to 5% points improvement over previously reported results. Anomaly detection tasks belong to the category of one-class-learning and are crucial in many applications, where a fixed set of classes cannot be defined, for instance, because a subset of classes is extremely rare or some classes are unknown at training time. For example, there might be a bear crossing the street as part of validation scenarios for automatic cars, unknown production anomalies due to critical change of the production environment, or unknown deviations from the healthy state in medical data.


The Many Moods of Emotion

arXiv.org Artificial Intelligence

Abstract-- This paper presents a novel approach to the facial expression generation problem. Building upon the assumption of the psychological community that emotion is intrinsically continuous, we first design our own continuous emotion representation with a 3-dimensional latent space issued from a neural network trained on discrete emotion classification. The so-obtained representation can be used to annotate large in the wild datasets and later used to trained a Generative Adversarial Network. We first show that our model is able to map back to discrete emotion classes with a objectively and subjectively better quality of the images than usual discrete approaches. But also that we are able to pave the larger space of possible facial expressions, generating the many moods of emotion. Moreover, two axis in this space may be found to generate similar expression changes as in traditional continuous representations such as arousalvalence. Finally we show from visual interpretation, that the third remaining dimension is highly related to the well-known dominance dimension from psychology. Affective computing is a topic of broad interest, finding applications in many fields such as healthcare, marketing or human-machine interfaces.


Gated Hierarchical Attention for Image Captioning

arXiv.org Artificial Intelligence

Attention modules connecting encoder and decoders have been widely applied in the field of object recognition, image captioning, visual question answering and neural machine translation, and significantly improves the performance. In this paper, we propose a bottom-up gated hierarchical attention (GHA) mechanism for image captioning. Our proposed model employs a CNN as the decoder which is able to learn different concepts at different layers, and apparently, different concepts correspond to different areas of an image. Therefore, we develop the GHA in which low-level concepts are merged into high-level concepts and simultaneously low-level attended features pass to the top to make predictions. Our GHA significantly improves the performance of the model that only applies one level attention, for example, the CIDEr score increases from 0.923 to 0.999, which is comparable to the state-of-the-art models that employ attributes boosting and reinforcement learning (RL). We also conduct extensive experiments to analyze the CNN decoder and our proposed GHA, and we find that deeper decoders cannot obtain better performance, and when the convolutional decoder becomes deeper the model is likely to collapse during training.


TensorFlow Agents: Efficient Batched Reinforcement Learning in TensorFlow

arXiv.org Artificial Intelligence

We introduce TensorFlow Agents, an efficient infrastructure paradigm for building parallel reinforcement learning algorithms in TensorFlow. We simulate multiple environments in parallel, and group them to perform the neural network computation on a batch rather than individual observations. This allows the TensorFlow execution engine to parallelize computation, without the need for manual synchronization. Environments are stepped in separate Python processes to progress them in parallel without interference of the global interpreter lock. As part of this project, we introduce BatchPPO, an efficient implementation of the proximal policy optimization algorithm. By open sourcing TensorFlow Agents, we hope to provide a flexible starting point for future projects that accelerates future research in the field.


Google Touts Speed, Accuracy From Machine Learning in DeepVariant

#artificialintelligence

CHICAGO (GenomeWeb) – Evidence published in the journal Nature Biotechnology in September demonstrated the efficacy of DeepVariant, Google's deep-learning-based variant caller, compared to older previous methods of calling genomic variants. DeepVariant "replaces the assortment of statistical modeling components with a single deep-learning model," according to the paper, whose authors represented Google and sister company Verily.


fast.ai · Making neural nets uncool again

#artificialintelligence

In machine learning and deep learning we can't do anything without data. So the people that create datasets for us to train our models are the (often under-appreciated) heroes. Some of the most useful and important datasets are those that become important "academic baselines"; that is, datasets that are widely studied by researchers and used to compare algorithmic changes. Some of these become household names (at least, among households that train models!), such as MNIST, CIFAR 10, and Imagenet. We all owe a debt of gratitude to those kind folks who have made datasets available for the research community.


Machine learning: more than a buzzword - CTOvision.com

#artificialintelligence

While artificial intelligence has long been heralded as the nextgen technology, many businesses are still afraid to dabble into it considering its very complicated nature. The IT industry has a severe case of Buzzword compliance. It is now compulsory for businesses to drop in words such as "machine learning," "artificial intelligence," or "deep learning" to invigorate any conversation around analytics. Unfortunately, this can make it harder to understand the real benefits that the latest evolution in analytics can bring.


Deep Learning for MR Angiography: Automated Detection of Cerebral Aneurysms

#artificialintelligence

To develop and evaluate a supportive algorithm using deep learning for detecting cerebral aneurysms at time-of-flight MR angiography to provide a second assessment of images already interpreted by radiologists. MR images reported by radiologists to contain aneurysms were extracted from four institutions for the period from November 2006 through October 2017. The images were divided into three data sets: training data set, internal test data set, and external test data set. The algorithm was constructed by deep learning with the training data set, and its sensitivity to detect aneurysms in the test data sets was evaluated. To find aneurysms that had been overlooked in the initial reports, two radiologists independently performed a blinded interpretation of aneurysm candidates detected by the algorithm. When there was disagreement, the final diagnosis was made in consensus. The number of newly detected aneurysms was also evaluated. The training data set, which provided training and validation data, included 748 aneurysms (mean size, 3.1 mm 2.0 [standard deviation]) from 683 examinations; 318 of these examinations were on male patients (mean age, 63 years 13) and 365 were on female patients (mean age, 64 years 13).