Goto

Collaborating Authors

 Asia


Spectral-Pruning: Compressing deep neural network via spectral analysis

arXiv.org Machine Learning

The model size of deep neural network is getting larger and larger to realize superior performance in complicated tasks. This makes it difficult to implement deep neural network in small edge-computing devices. To overcome this problem, model compression methods have been gathering much attention. However, there have been only few theoretical back-grounds that explain what kind of quantity determines the compression ability. To resolve this issue, we develop a new theoretical frame-work for model compression, and propose a new method called {\it Spectral-Pruning} based on the theory. Our theoretical analysis is based on the observation such that the eigenvalues of the covariance matrix of the output from nodes in the internal layers often shows rapid decay. We define "degree of freedom" to quantify an intrinsic dimensionality of the model by using the eigenvalue distribution and show that the compression ability is essentially controlled by this quantity. Along with this, we give a generalization error bound of the compressed model. Our proposed method is applicable to wide range of models, unlike the existing methods, e.g., ones possess complicated branches as implemented in SegNet and ResNet. Our method makes use of both "input" and "output" in each layer and is easy to implement. We apply our method to several datasets to justify our theoretical analyses and show that the proposed method achieves the state-of-the-art performance.


An Incremental Construction of Deep Neuro Fuzzy System for Continual Learning of Non-stationary Data Streams

arXiv.org Artificial Intelligence

Existing fuzzy neural networks (FNNs) are mostly developed under a shallow network configuration having lower generalization power than those of deep structures. This paper proposes a novel self-organizing deep fuzzy neural network, namely deep evolving fuzzy neural networks (DEVFNN). Fuzzy rules can be automatically extracted from data streams or removed if they play little role during their lifespan. The structure of the network can be deepened on demand by stacking additional layers using a drift detection method which not only detects the covariate drift, variations of input space, but also accurately identifies the real drift, dynamic changes of both feature space and target space. DEVFNN is developed under the stacked generalization principle via the feature augmentation concept where a recently developed algorithm, namely Generic Classifier (gClass), drives the hidden layer. It is equipped by an automatic feature selection method which controls activation and deactivation of input attributes to induce varying subsets of input features. A deep network simplification procedure is put forward using the concept of hidden layer merging to prevent uncontrollable growth of input space dimension due to the nature of feature augmentation approach in building a deep network structure. DEVFNN works in the sample-wise fashion and is compatible for data stream applications. The efficacy of DEVFNN has been thoroughly evaluated using six datasets with non-stationary properties under the prequential test-then-train protocol. It has been compared with four state-of the art data stream methods and its shallow counterpart where DEVFNN demonstrates improvement of classification accuracy.


A study on speech enhancement using exponent-only floating point quantized neural network (EOFP-QNN)

arXiv.org Artificial Intelligence

Numerous studies have investigated the effectiveness of neural network quantization on pattern classification tasks. The present study, for the first time, investigated the performance of speech enhancement (a regression task in speech processing) using a novel exponent-only floating-point quantized neural network (EOFP-QNN). The proposed EOFP-QNN consists of two stages: mantissa-quantization and exponent-quantization. In the mantissa-quantization stage, EOFP-QNN learns how to quantize the mantissa bits of the model parameters while preserving the regression accuracy using the least mantissa precision. In the exponent-quantization stage, the exponent part of the parameters is further quantized without causing any additional performance degradation. We evaluated the proposed EOFP quantization technique on two types of neural networks, namely, bidirectional long short-term memory (BLSTM) and fully convolutional neural network (FCN), on a speech enhancement task. Experimental results showed that the model sizes can be significantly reduced (the model sizes of the quantized BLSTM and FCN models were only 18.75% and 21.89%, respectively, compared to those of the original models) while maintaining satisfactory speech-enhancement performance.


Learning When to Concentrate or Divert Attention: Self-Adaptive Attention Temperature for Neural Machine Translation

arXiv.org Artificial Intelligence

Most of the Neural Machine Translation (NMT) models are based on the sequence-to-sequence (Seq2Seq) model with an encoder-decoder framework equipped with the attention mechanism. However, the conventional attention mechanism treats the decoding at each time step equally with the same matrix, which is problematic since the softness of the attention for different types of words (e.g. content words and function words) should differ. Therefore, we propose a new model with a mechanism called Self-Adaptive Control of Temperature (SACT) to control the softness of attention by means of an attention temperature. Experimental results on the Chinese-English translation and English-Vietnamese translation demonstrate that our model outperforms the baseline models, and the analysis and the case study show that our model can attend to the most relevant elements in the source-side contexts and generate the translation of high quality.


Learn more about the future of robotics at Disrupt SF

#artificialintelligence

At at Disrupt SF, we'll be joined by four experts to discuss how new technologies are changing the field. Those experts include Peter Barrett, founder and CTO and Playground, a venture fund and design studio focused on hardware startups. Barrett is a 30-year veteran of the tech industry, whose accomplishments include developing Cinepak (video compression software that was included as part of Apple QuickTime) and working at WebTV -- which was acquired by Microsoft, where he led Internet TV efforts for more than a decade. We'll also be joined by Helen Boniske, a partner at early stage hardware investor Lemnos. Before joining Lemnos, Boniske was a front office executive for the Arizona Diamondbacks.


Emotion and Sentiment Analysis: A Practitioner's Guide to NLP

#artificialintelligence

Sentiment analysis is perhaps one of the most popular applications of NLP, with a vast number of tutorials, courses, and applications that focus on analyzing sentiments of diverse datasets ranging from corporate surveys to movie reviews. The key aspect of sentiment analysis is to analyze a body of text for understanding the opinion expressed by it. Typically, we quantify this sentiment with a positive or negative value, called polarity. The overall sentiment is often inferred as positive, neutral or negative from the sign of the polarity score. Usually, sentiment analysis works best on text that has a subjective context than on text with only an objective context.


Odd Numbers -- Real Life

#artificialintelligence

Algorithms increasingly govern our social world, transforming data into scores or rankings that decide who gets credit, jobs, dates, policing, and much more. The field of "algorithmic accountability" has arisen to highlight the problems with such methods of classifying people, and it has great promise: Cutting-edge work in critical algorithm studies applies social theory to current events; law and policy experts seem to publish new articles daily on how artificial intelligence shapes our lives, and a growing community of researchers has developed a field known as "Fairness, Accuracy, and Transparency in Machine Learning." The social scientists, attorneys, and computer scientists promoting algorithmic accountability aspire to advance knowledge and promote justice. But what should such "accountability" more specifically consist of? At a two-day, interdisciplinary roundtable on AI ethics I recently attended, such questions featured prominently, and humanists, policy experts, and lawyers engaged in a free-wheeling discussion about topics ranging from robot arms races to computationally planned economies.


Busting Moves with DanceNet AI

#artificialintelligence

Inspired by STEM-focused YouTuber carykh, Indian developer Jaison Saji has produced a deep network system, DanceNet, that can automatically generate dance moves. Synced used DanceNet to produce a short clip (below) with code published on Github. Interested readers can use the system to improve this work or create their own. DanceNet uses a variational autoencoder (VAE) to automatically generate thousands of single dance pose pictures, then sequentially connect them to produce vigorous dance movements through joint training on Long Short-Term Memory (LSTM) and Mixture Density Networks (MDN). VAE is a commonly used generative model with two parts: an encoder transfers the image into a dense representation that has few dimensions and occupies less space than the original source and stores latent information about the input; while a decoder transfers dense-represented code back to its corresponding image.


AI FOR GOOD PRIVACY

#artificialintelligence

Get YouTube without the ads. Want to watch this again later? Sign in to add this video to a playlist. Report Need to report the video? Sign in to report inappropriate content.


Deflating Towards Utopia

#artificialintelligence

Central banks have inflation targets. We want inflation, because with inflation, money is more expensive in the future, thus better spent today. Inflation means growth, deflation means debt spiral. That is the lesson from 100 years of modern economics. I have long argued that we will never return to inflation again.