Genre
Approximate Inference with Amortised MCMC
Li, Yingzhen, Turner, Richard E., Liu, Qiang
We propose a novel approximate inference algorithm that approximates a target distribution by amortising the dynamics of a user-selected MCMC sampler. The idea is to initialise MCMC using samples from an approximation network, apply the MCMC operator to improve these samples, and finally use the samples to update the approximation network thereby improving its quality. This provides a new generic framework for approximate inference, allowing us to deploy highly complex, or implicitly defined approximation families with intractable densities, including approximations produced by warping a source of randomness through a deep neural network. Experiments consider image modelling with deep generative models as a challenging test for the method. Deep models trained using amortised MCMC are shown to generate realistic looking samples as well as producing diverse imputations for images with regions of missing pixels.
A unified view of entropy-regularized Markov decision processes
Neu, Gergely, Jonsson, Anders, Gรณmez, Vicenรง
We propose a general framework for entropy-regularized average-reward reinforcement learning in Markov decision processes (MDPs). Our approach is based on extending the linear-programming formulation of policy optimization in MDPs to accommodate convex regularization functions. Our key result is showing that using the conditional entropy of the joint state-action distributions as regularization yields a dual optimization problem closely resembling the Bellman optimality equations. This result enables us to formalize a number of state-of-the-art entropy-regularized reinforcement learning algorithms as approximate variants of Mirror Descent or Dual Averaging, and thus to argue about the convergence properties of these methods. In particular, we show that the exact version of the TRPO algorithm of Schulman et al. (2015) actually converges to the optimal policy, while the entropy-regularized policy gradient methods of Mnih et al. (2016) may fail to converge to a fixed point. Finally, we illustrate empirically the effects of using various regularization techniques on learning performance in a simple reinforcement learning setup.
Training Deep Convolutional Neural Networks with Resistive Cross-Point Devices
Gokmen, Tayfun, Onen, O. Murat, Haensch, Wilfried
In a previous work we have detailed the requirements to obtain a maximal performance benefit by implementing fully connected deep neural networks (DNN) in form of arrays of resistive devices for deep learning. This concept of Resistive Processing Unit (RPU) devices we extend here towards convolutional neural networks (CNNs). We show how to map the convolutional layers to RPU arrays such that the parallelism of the hardware can be fully utilized in all three cycles of the backpropagation algorithm. We find that the noise and bound limitations imposed due to analog nature of the computations performed on the arrays effect the training accuracy of the CNNs. Noise and bound management techniques are presented that mitigate these problems without introducing any additional complexity in the analog circuits and can be addressed by the digital circuits. In addition, we discuss digitally programmable update management and device variability reduction techniques that can be used selectively for some of the layers in a CNN. We show that combination of all those techniques enables a successful application of the RPU concept for training CNNs. The techniques discussed here are more general and can be applied beyond CNN architectures and therefore enables applicability of RPU approach for large class of neural network architectures.
On-the-fly Operation Batching in Dynamic Computation Graphs
Neubig, Graham, Goldberg, Yoav, Dyer, Chris
Dynamic neural network toolkits such as PyTorch, DyNet, and Chainer offer more flexibility for implementing models that cope with data of varying dimensions and structure, relative to toolkits that operate on statically declared computations (e.g., TensorFlow, CNTK, and Theano). However, existing toolkits - both static and dynamic - require that the developer organize the computations into the batches necessary for exploiting high-performance algorithms and hardware. This batching task is generally difficult, but it becomes a major hurdle as architectures become complex. In this paper, we present an algorithm, and its implementation in the DyNet toolkit, for automatically batching operations. Developers simply write minibatch computations as aggregations of single instance computations, and the batching algorithm seamlessly executes them, on the fly, using computationally efficient batched operations. On a variety of tasks, we obtain throughput similar to that obtained with manual batches, as well as comparable speedups over single-instance learning on architectures that are impractical to batch manually.
Google's AI Chief On Teaching Computers To LearnโAnd The Challenges Ahead
After the keynote, I caught up with Google senior VP of engineering John Giannandreaโwho, though he didn't appear onstage, is deeply involved in all of the above efforts and others as the company's lead for AI. "Last year, we talked about becoming an AI-first company and people weren't entirely sure what we meant," he told me. At I/O, Google announced Google.aiโwhich is maybe less of an actual thing than a statement (and accompanying website) designed to remind the world of the company's ambitious and far-flung efforts in AI. Giannandrea calls it "an umbrella brand" that shows off Google's work in hopes of inspiring others to build upon it. "We're saying, 'Come use this amazing stuff, see what you can do," he explains.
Some natural solutions to the p-value communication problem--and why they won't work.
Instead, do the work to present statistical conclusions with uncertainty rather than as dichotomies. Also, remember that most effects can't be zero (at least in social science and public health), and that an "effect" is usually a mean in a population (or something similar such as a regression coefficient)--a fact that seems to be lost from consciousness when researchers slip into binary statements about there being "an effect" or "no effect" as if they are writing about constants of nature. Again, it will be difficult to resolve the many problems with p-values and "statistical significance" without addressing the mistaken goal of certainty which such methods have been used to pursue.
AllAnalytics - Ariella Brown - Why Machine Learning Can Improve Customer Service
AI is changing our everyday interactions. What once required a human rep can now be handled by a virtual assistant whose programming allows customer problems to be solved more quickly. A recent Venturebeat article declared, "AI chatbots are the next big shift in customer service." Those of us of a certain generation expect to wait on a line or on the phone for a person to take care of our customer service issues. But the generation that favors texts to calls has come to have different expectations.
Artificial Intelligence Use Cases: An Overview - DATAVERSITY
The Artificial Intelligence Market Forecasts 2016 -2025 across 27 Industry Sectors has provided an overview of numerous Artificial Intelligence use cases, which includes Machine Learning, machine reasoning, Deep Learning, NLP, computer vision, and many other allied technologies. According to this study, food services, consumer products, advertising, and defense (along with others mentioned above) will significantly benefit from the growth of AI in the coming years.
The 10 Algorithms Machine Learning Engineers Need to Know
It is no doubt that the sub-field of machine learning / artificial intelligence has increasingly gained more popularity in the past couple of years. As Big Data is the hottest trend in the tech industry at the moment, machine learning is incredibly powerful to make predictions or calculated suggestions based on large amounts of data. Some of the most common examples of machine learning are Netflix's algorithms to make movie suggestions based on movies you have watched in the past or Amazon's algorithms that recommend books based on books you have bought before. So if you want to learn more about machine learning, how do you start? For me, my first introduction is when I took an Artificial Intelligence class when I was studying abroad in Copenhagen. My lecturer is a full-time Applied Math and CS professor at the Technical University of Denmark, in which his research areas are logic and artificial, focusing primarily on the use of logic to model human-like planning, reasoning and problem solving.
Spotify acquires Niland, a machine learning and AI startup - SlashGear
Spotify has announced the acquisition of Niland, a machine learning startup based out of Paris. The music company made the announcement itself this week, explaining that Niland shares its'passion for surfacing the right content to the right user at the right time.' Spotify plans to use the company's technology and know-how to improves its own recommendation abilities, doing so with the power of artificial intelligence behind it. Spotify announced the acquisition on Wednesday, saying that the Niland team will be joining the music company's own team in its New York office. The terms of the deal weren't revealed, such as how much Spotify paid for the company or when the deal was finalized. We do know, however, that personalized recommendations on Spotify are about to get much better than to the startup's work in machine learning and artificial intelligence.