Goto

Collaborating Authors

 Asia


Parameter Space Noise for Exploration

arXiv.org Machine Learning

Deep reinforcement learning (RL) methods generally engage in exploratory behavior through noise injection in the action space. An alternative is to add noise directly to the agent's parameters, which can lead to more consistent exploration and a richer set of behaviors. Methods such as evolutionary strategies use parameter perturbations, but discard all temporal structure in the process and require significantly more samples. Combining parameter noise with traditional RL methods allows to combine the best of both worlds. We demonstrate that both off- and on-policy methods benefit from this approach through experimental comparison of DQN, DDPG, and TRPO on high-dimensional discrete action environments as well as continuous control tasks. Our results show that RL with parameter noise learns more efficiently than traditional RL with action space noise and evolutionary strategies individually.


Alternating Multi-bit Quantization for Recurrent Neural Networks

arXiv.org Machine Learning

Recurrent neural networks have achieved excellent performance in many applications. However, on portable devices with limited resources, the models are often too large to deploy. For applications on the server with large scale concurrent requests, the latency during inference can also be very critical for costly computing resources. In this work, we address these problems by quantizing the network, both weights and activations, into multiple binary codes {-1,+1}. We formulate the quantization as an optimization problem. Under the key observation that once the quantization coefficients are fixed the binary codes can be derived efficiently by binary search tree, alternating minimization is then applied. We test the quantization for two well-known RNNs, i.e., long short term memory (LSTM) and gated recurrent unit (GRU), on the language models. Compared with the full-precision counter part, by 2-bit quantization we can achieve ~16x memory saving and ~6x real inference acceleration on CPUs, with only a reasonable loss in the accuracy. By 3-bit quantization, we can achieve almost no loss in the accuracy or even surpass the original model, with ~10.5x memory saving and ~3x real inference acceleration. Both results beat the exiting quantization works with large margins. We extend our alternating quantization to image classification tasks. In both RNNs and feedforward neural networks, the method also achieves excellent performance.


Distributed Newton Methods for Deep Neural Networks

arXiv.org Machine Learning

Deep learning involves a difficult non-convex optimization problem with a large number of weights between any two adjacent layers of a deep structure. To handle large data sets or complicated networks, distributed training is needed, but the calculation of function, gradient, and Hessian is expensive. In particular, the communication and the synchronization cost may become a bottleneck. In this paper, we focus on situations where the model is distributedly stored, and propose a novel distributed Newton method for training deep neural networks. By variable and feature-wise data partitions, and some careful designs, we are able to explicitly use the Jacobian matrix for matrix-vector products in the Newton method. Some techniques are incorporated to reduce the running time as well as the memory consumption. First, to reduce the communication cost, we propose a diagonalization method such that an approximate Newton direction can be obtained without communication between machines. Second, we consider subsampled Gauss-Newton matrices for reducing the running time as well as the communication cost. Third, to reduce the synchronization cost, we terminate the process of finding an approximate Newton direction even though some nodes have not finished their tasks. Details of some implementation issues in distributed environments are thoroughly investigated. Experiments demonstrate that the proposed method is effective for the distributed training of deep neural networks. In compared with stochastic gradient methods, it is more robust and may give better test accuracy.


A Modified Sigma-Pi-Sigma Neural Network with Adaptive Choice of Multinomials

arXiv.org Machine Learning

Sigma-Pi-Sigma neural networks (SPSNNs) as a kind of high-order neural networks can provide more powerful mapping capability than the traditional feedforward neural networks (Sigma-Sigma neural networks). In the existing literature, in order to reduce the number of the Pi nodes in the Pi layer, a special multinomial P_s is used in SPSNNs. Each monomial in P_s is linear with respect to each particular variable sigma_i when the other variables are taken as constants. Therefore, the monomials like sigma_i^n or sigma_i^n sigma_j with n>1 are not included. This choice may be somehow intuitive, but is not necessarily the best. We propose in this paper a modified Sigma-Pi-Sigma neural network (MSPSNN) with an adaptive approach to find a better multinomial for a given problem. To elaborate, we start from a complete multinomial with a given order. Then we employ a regularization technique in the learning process for the given problem to reduce the number of monomials used in the multinomial, and end up with a new SPSNN involving the same number of monomials (= the number of nodes in the Pi-layer) as in P_s. Numerical experiments on some benchmark problems show that our MSPSNN behaves better than the traditional SPSNN with P_s.


Model compression for faster structural separation of macromolecules captured by Cellular Electron Cryo-Tomography

arXiv.org Machine Learning

Electron Cryo-Tomography (ECT) enables 3D visualization of macromolecule structure inside single cells. Macromolecule classification approaches based on convolutional neural networks (CNN) were developed to separate millions of macromolecules captured from ECT systematically. However, given the fast accumulation of ECT data, it will soon become necessary to use CNN models to efficiently and accurately separate substantially more macromolecules at the prediction stage, which requires additional computational costs. To speed up the prediction, we compress classification models into compact neural networks with little in accuracy for deployment. Specifically, we propose to perform model compression through knowledge distillation. Firstly, a complex teacher network is trained to generate soft labels with better classification feasibility followed by training of customized student networks with simple architectures using the soft label to compress model complexity. Our tests demonstrate that our compressed models significantly reduce the number of parameters and time cost while maintaining similar classification accuracy.


China wants to create the chips to power the future of AI

#artificialintelligence

When you think about the major players in AI, a bunch of names leap to mind: Google, Microsoft, IBM, Facebook. One that you might overlook is Pinterest. But the company, whose whole business rests on wrangling vast quantities of imagery, has long done ambitious work in super-smart visual search.


Artificial intelligence is the weapon of the next Cold War

#artificialintelligence

With artificial intelligence weapons on both sides, are we in a new cold war? It is easy to confuse the current geopolitical situation with that of the 1980s. The United States and Russia each accuse the other of interfering in domestic affairs.…


Artificial Intelligence will open up new avenues: Experts - ET Telecom

#artificialintelligence

NEW DELHI: Citing the incident where Facebook had to abandon an experiment undertaken last year where two artificially intelligent programs or chat bots appeared to be chatting to each other in a strange language which they developed on their own and only they understood, Dr. Jitendra K. Das set the tone of the conclave on "The Confluence of Artificial Intelligence and Data Analytics" held recently at the FORE School of Management, New Delhi, in association with BRICS Chamber of Commerce and Industry. Dr. Das further explained how with the help of complex virtual learning techniques, a wide range of physical and cognitive tasks are being managed today with a high level of efficiency and accuracy. And as artificial intelligence or AI systems advance through machine learning these will continue to impact not just business but our lives as well. But, if indeed machines continue to improve their performance beyond human levels, a natural question to ask is whether machines will put humans' jobs at risk and reduce employment. According to Mr. Vijay Sethi, CIO and Head CSR at Hero MotoCorp Ltd., "Such a concern is not new and in fact dates back to the 1940s when AI and automation started developing."


People-carrying robot to the rescue

USATODAY - Tech Top Stories

Researchers in South Korea show off their people-carrying robot designed for rescue missions or helping people with disabilities. A link has been posted to your Facebook feed. Researchers in South Korea show off their people-carrying robot designed for rescue missions or helping people with disabilities.


Robot barista to serve coffee at travel agency's Tokyo cafe

The Japan Times

Travel agency H.I.S. Co. said Tuesday that a robot will serve coffee to customers at a cafe it plans to open next month at its flagship branch in central Tokyo. The agency will open its Henn Na Cafe (strange cafe) featuring a robotic arm and an automated coffee maker at its store in Tokyo's Shibuya Ward. The move will add to the company's series of services using robots, including at the Henn Na Hotel (strange hotel) in Sasebo, Nagasaki Prefecture, where a dinosaur robot welcomes guests at the front desk. At the Shibuya cafe, customers will be greeted by the U.S.-made robot, with an attached screen showing facial expressions. Would you like delicious coffee?" the machine asks in Japanese.