Asia
Applied AI Digest 112 – BootstrapLabs
Artificial-intelligence programs could develop some much-needed common sense by competing in scavenger hunts inside virtual homes filled with simulated coffee tables, couches, lamps, and other everyday things. Researchers at Facebook and Georgia Tech developed the scavenger-hunt challenge. Can you imagine being able to solve complex problems almost instantaneously? What would normally take months would take minutes. Seemingly impossible questions would become straightforward to answer.
Integral representation of the global minimizer
Sonoda, Sho, Ishikawa, Isao, Ikeda, Masahiro, Hagihara, Kei, Sawano, Yoshihiro, Matsubara, Takuo, Murata, Noboru
We have obtained an integral representation of the shallow neural network that attains the global minimum of its backpropagation (BP) training problem. According to our unpublished numerical simulations conducted several years prior to this study, we had noticed that such an integral representation may exist, but it was not proven until today. First, we introduced a Hilbert space of coefficient functions, and a reproducing kernel Hilbert space (RKHS) of hypotheses, associated with the integral representation. The RKHS reflects the approximation ability of neural networks. Second, we established the ridgelet analysis on RKHS. The analytic property of the integral representation is remarkably clear. Third, we reformulated the BP training as the optimization problem in the space of coefficient functions, and obtained a formal expression of the unique global minimizer, according to the Tikhonov regularization theory. Finally, we demonstrated that the global minimizer is the shrink ridgelet transform. Since the relation between an integral representation and an ordinary finite network is not clear, and BP is convex in the integral representation, we cannot immediately answer the question such as "Is a local minimum a global minimum?" However, the obtained integral representation provides an explicit expression of the global minimizer, without linearity-like assumptions, such as partial linearity and monotonicity. Furthermore, it indicates that the ordinary ridgelet transform provides the minimum norm solution to the original training equation.
Estimation of Non-Normalized Mixture Models and Clustering Using Deep Representation
Matsuda, Takeru, Hyvarinen, Aapo
We develop a general method for estimating a finite mixture of non-normalized models. Here, a non-normalized model is defined to be a parametric distribution with an intractable normalization constant. Existing methods for estimating non-normalized models without computing the normalization constant are not applicable to mixture models because they contain more than one intractable normalization constant. The proposed method is derived by extending noise contrastive estimation (NCE), which estimates non-normalized models by discriminating between the observed data and some artificially generated noise. We also propose an extension of NCE with multiple noise distributions. Then, based on the observation that conventional classification learning with neural networks is implicitly assuming an exponential family as a generative model, we introduce a method for clustering unlabeled data by estimating a finite mixture of distributions in an exponential family. Estimation of this mixture model is attained by the proposed extensions of NCE where the training data of neural networks are used as noise. Thus, the proposed method provides a probabilistically principled clustering method that is able to utilize a deep representation. Application to image clustering using a deep neural network gives promising results.
Comments on "Momentum fractional LMS for power signal parameter estimation"
Khan, Shujaat, Naseem, Imran, Sadiq, Alishba, Ahmad, Jawwad, Moinuddin, Muhammad
The purpose of this paper is to indicate that the recently proposed Momentum fractional least mean squares (mFLMS) algorithm has some serious flaws in its design and analysis. Our apprehensions are based on the evidence we found in the derivation and analysis in the paper titled: \textquotedblleft \textit{Momentum fractional LMS for power signal parameter estimation}\textquotedblright. In addition to the theoretical bases our claims are also verified through extensive simulation results. The experiments clearly show that the new method does not have any advantage over the classical least mean square (LMS) method.
Learning to Detect
Samuel, Neev, Diskin, Tzvi, Wiesel, Ami
We introduce two different deep architectures: a standard fully connected multi-layer network, and a Detection Network (DetNet) which is specifically designed for the task. The structure of DetNet is obtained by unfolding the iterations of a projected gradient descent algorithm into a network. We compare the accuracy and runtime complexity of the purposed approaches and achieve state-of-the-art performance while maintaining low computational requirements. Furthermore, we manage to train a single network to detect over an entire distribution of channels. Finally, we consider detection with soft outputs and show that the networks can easily be modified to produce soft decisions.
Nostalgic Adam: Weighing more of the past gradients when designing the adaptive learning rate
Huang, Haiwen, Wang, Chang, Dong, Bin
First-order optimization methods have been playing a prominent role in deep learning. Algorithms such as RMSProp and Adam are rather popular in training deep neural networks on large datasets. Recently, Reddi et al. discovered a flaw in the proof of convergence of Adam, and the authors proposed an alternative algorithm, AMSGrad, which has guaranteed convergence under certain conditions. In this paper, we propose a new algorithm, called Nostalgic Adam (NosAdam), which places bigger weights on the past gradients than the recent gradients when designing the adaptive learning rate. This is a new observation made through mathematical analysis of the algorithm. We also show that the estimate of the second moment of the gradient in NosAdam vanishes slower than Adam, which may account for faster convergence of NosAdam. We analyze the convergence of NosAdam and discover a convergence rate that achieves the best known convergence rate $O(1/\sqrt{T})$ for general convex online learning problems. Empirically, we show that NosAdam outperforms AMSGrad and Adam in some common machine learning problems.
Nvidia, AMD Get Buy Ratings On Artificial Intelligence Prospects
Graphics-chip makers Advanced Micro Devices (AMD) and Nvidia (NVDA) received fresh buy ratings on Friday, as investment bank Cowen initiated coverage of a host of semiconductor stocks. Cowen analyst Matthew Ramsay started coverage of 11 chip stocks. He rated eight as outperform, or buy, and three as market perform, or neutral. Ramsay said his top picks are AMD, Ambarella (AMBA), Broadcom (AVGO) and Monolithic Power Systems (MPWR). He also likes Nvidia, even though it has a high valuation.
VR in the sky is better than VR in your home
I'm watching someone on the edge of a helicopter as he counts down with his hands. Three, two, one, and we leap. Instead, I fall forward into an iFly indoor skydiving wind tunnel and I start to float. This is all happening at indoor skydiving facility, iFly in Union City, California. This week the company introduced its virtual reality flying experience to 28 of the company's 37 locations. In addition to the usual float-above-a-giant-fan, patrons can now put on a VR helmet for about $20 more than the usual cost of about $70 (prices are determined by region) and experience what it's like to leap out of an aircraft above Hawaii, Dubai, Southern California and the Alps.
The Camera, Transformed by Machine Learning - Core77
Together, they suggest a shared cultural understanding of a camera: a classic point-and-shoot. But the cameras we encounter every day bear little resemblance--in form or function--to this vestigial object. New capabilities in software, new hardware formats and imaging technologies, and emerging user behaviors around image creation are radically reshaping the object we know of as the "camera" into new categories. Perhaps the most impactful influence on the camera is being brought about by computer vision: empowering cameras to not only capture various kinds of images but to also parse visual information--effectively, to understand the world. Software trained on vast datasets of labeled images can recognize things like vehicles, dogs, cats, and people, along with facial features, emotions, and second-order information like movement vectors and gaze direction from raw images and videos.
Baidu Offloads Ticketing Platform to iQiyi to Focus on AI
Baidu Inc. is transferring its movie theater ticketing platform Nuomi Film to video streaming unit iQiyi Inc. so that it can focus efforts on the fast-growing artificial intelligence sector. Baidu had been long-rumored to be looking to offload parts of its Nuomi business, which had been hemorrhaging money after its parent company focused more intensely on the AI sector. The firm, most well-known for its dominant search engine, has adopted an AI strategy similar to Google-parent Alphabet Inc. covering consumer applications such as autonomous cars. Supported by the government's ambitions to become a world-leader by 2030, China is leading the race to develop AI products globally. The country dominated AI funding last year with Chinese startups accounting for a little under half of all financing pumped into the sector globally in 2017, leveraging China's vast market and well-established data infrastructure.