Asia
Efficient Reinforcement Learning with a Mind-Game for Full-Length StarCraft II
Liu, Ruo-Ze, Guo, Haifeng, Ji, Xiaozhong, Yu, Yang, Xiao, Zitai, Wu, Yuzhou, Pang, Zhen-Jia, Lu, Tong
StarCraft II provides an extremely challenging platform for reinforcement learning due to its huge state-space and game length. The previous fastest method requires days to train a full-length game policy in a single commercial machine. In this paper, we introduce the mind-game to facilitate the reinforcement learning, which is an abstract task model. With the mind-game, the policy is firstly trained in the mind-game fastly and is then mapped to the real game for the second phase training. In our experiments, the trained agent can achieve a 100% win-rate on the map Simple64 against the most difficult non-cheating built-in bot (level-7), and the training is 100 times faster than the previous ones under the same computational resource. To test the generalization performance of the agent, a Golden level of StarCraft II Ladder human player has competed with the agent. With restricted strategy, the agent wins the human player by 4 out of 5 games. The mind-game approach might shed some light for further studies of efficient reinforcement learning. The codes are publicly available (https://github.com/mindgameSC2/mind-SC2).
A Cooperative Multi-Agent Reinforcement Learning Framework for Resource Balancing in Complex Logistics Network
Li, Xihan, Zhang, Jia, Bian, Jiang, Tong, Yunhai, Liu, Tie-Yan
Resource balancing within complex transportation networks is one of the most important problems in real logistics domain. Traditional solutions on these problems leverage combinatorial optimization with demand and supply forecasting. However, the high complexity of transportation routes, severe uncertainty of future demand and supply, together with non-convex business constraints make it extremely challenging in the traditional resource management field. In this paper, we propose a novel sophisticated multi-agent reinforcement learning approach to address these challenges. In particular, inspired by the externalities especially the interactions among resource agents, we introduce an innovative cooperative mechanism for state and reward design resulting in more effective and efficient transportation. Extensive experiments on a simulated ocean transportation service demonstrate that our new approach can stimulate cooperation among agents and lead to much better performance. Compared with traditional solutions based on combinatorial optimization, our approach can give rise to a significant improvement in terms of both performance and stability.
Asynchronous Episodic Deep Deterministic Policy Gradient: Towards Continuous Control in Computationally Complex Environments
Zhang, Zhizheng, Chen, Jiale, Chen, Zhibo, Li, Weiping
Deep Deterministic Policy Gradient (DDPG) has been proved to be a successful reinforcement learning (RL) algorithm for continuous control tasks. However, DDPG still suffers from data insufficiency and training inefficiency, especially in computationally complex environments. In this paper, we propose Asynchronous Episodic DDPG (AE-DDPG), as an expansion of DDPG, which can achieve more effective learning with less training time required. First, we design a modified scheme for data collection in an asynchronous fashion. Generally, for asynchronous RL algorithms, sample efficiency or/and training stability diminish as the degree of parallelism increases. We consider this problem from the perspectives of both data generation and data utilization. In detail, we re-design experience replay by introducing the idea of episodic control so that the agent can latch on good trajectories rapidly. In addition, we also inject a new type of noise in action space to enrich the exploration behaviors. Experiments demonstrate that our AE-DDPG achieves higher rewards and requires less time consuming than most popular RL algorithms in Learning to Run task which has a computationally complex environment. Not limited to the control tasks in computationally complex environments, AE-DDPG also achieves higher rewards and 2- to 4-fold improvement in sample efficiency on average compared to other variants of DDPG in MuJoCo environments. Furthermore, we verify the effectiveness of each proposed technique component through abundant ablation study.
Matrix Completion via Nonconvex Regularization: Convergence of the Proximal Gradient Algorithm
Wen, Fei, Ying, Rendong, Liu, Peilin, Truong, Trieu-Kien
Matrix completion has attracted much interest in the past decade in machine learning and computer vision. For low-rank promotion in matrix completion, the nuclear norm penalty is convenient due to its convexity but has a bias problem. Recently, various algorithms using nonconvex penalties have been proposed, among which the proximal gradient descent (PGD) algorithm is one of the most efficient and effective. For the nonconvex PGD algorithm, whether it converges to a local minimizer and its convergence rate are still unclear. This work provides a nontrivial analysis on the PGD algorithm in the nonconvex case. Besides the convergence to a stationary point for a generalized nonconvex penalty, we provide more deep analysis on a popular and important class of nonconvex penalties which have discontinuous thresholding functions. For such penalties, we establish the finite rank convergence, convergence to restricted strictly local minimizer and eventually linear convergence rate of the PGD algorithm. Meanwhile, convergence to a local minimizer has been proved for the hard-thresholding penalty. Our result is the first shows that, nonconvex regularized matrix completion only has restricted strictly local minimizers, and the PGD algorithm can converge to such minimizers with eventually linear rate under certain conditions. Illustration of the PGD algorithm via experiments has also been provided. Code is available at https://github.com/FWen/nmc.
Video Friday: MIT's Mini Cheetah Robot, and More
Video Friday is your weekly selection of awesome robotics videos, collected by your Automaton bloggers. We'll also be posting a weekly calendar of upcoming robotics events for the next few months; here's what we have so far (send us your events!): Let us know if you have suggestions for next week, and enjoy today's videos. Impressive new video of MIT's Mini Cheetah doing backflips, and failing to do backflips, which is even cuter. MIT'S new mini cheetah robot is the first four-legged robot to do a backflip.
The Rise of Computer Vision Technology Analytics Insight
Computer vision Technology is rising, and increasingly gathering followers who want to adapt this technology to achieve new business heights. The rise of this technology can be attributed to the recent projections that have catapulted this technology to new zeniths. According to a market research, the computer vision market is valued at US$11.94 Billion and is likely to reach to US$17.38 Billion by 2023 growing at a CAGR of 7.80% from 2018 and 2023. The growth of the computer vision market is driven by the increasing adoption of computer vision into semi-autonomous and autonomous vehicles, and consumer drones which is dominated by the rising adoption of Industry 4.0. Recent advancements into computer vision technology with deep learning software, advanced cameras and image sensors, have expanded the scope for computer vision systems which can be deployed in a wide range of applications in different industries.
Neuromorphic computing and the brain that wouldn't die ZDNet
Inspired by a theory into the organisms of memory and recall in the brain, neural networking is a digital simulation of how synapses may retain information, after being trained to recognize patterns. For instance, neural nets enable a computer, or perhaps a cloud-based service, to recognize the characters of printed text without the need for programming explicitly specifying what text is, or how it can spot a certain face in a crowd after having seen several photographs of the same face. As a neural networking problem becomes linearly broader -- for example, distinguishing one form of written text from another -- the data required to train it grows exponentially larger. There's a valid argument that some of the tasks being envisioned for neural nets, such as spotting when anyone is getting depressed or agitated, may be impossible, even with today's storage and memory technologies. So the revelations by researchers that chemical structures comprised of completely random assemblies of nanometer-scale wires may exhibit the electrical characteristics of memory in a brain perhaps shouldn't continue to be dismissed for much longer.
Xbox Adaptive Controller: How Microsoft's latest Xbox accessory was born
One ad is a statement. Two ads – each during two of the most important moments for marketing – seems like commitment. And so when Xbox showed off its Adaptive Controller during both its Christmas and Super Bowl ad spots, it suggested they represented something important for the company. The special controller – much larger than the one traditionally sold with the console, and featuring a whole range of buttons built for people who might not usually be able to use them – was not just another accessory but one that represented a deeper aim for the company. That plan was to allow anyone to play on the Xbox, no matter the disabilities or other challenges that might have kept them from consoles in the past.
AirPods 2 release date: Rumours begin to swirl about Apple's updated wireless earphones for iPhones
Rumours are beginning to swirl that Apple will imminently release a new version of its AirPods. The much-loved, small and wireless earphones were first introduced in 2016. Almost since then, discussion has turned to any potential update from Apple, but the company has said very little. When it introduced its AirPower charging mat, it said it would offer a new version of the box the earphones come in so that they could be charged wirelessly, using the mat. But neither the mat or the new box have arrived, leading many to wonder when an update to the AirPods would actually be revealed.
Devil May Cry 5 is shaping up to be the best in the series so far
What officially began as the earliest incarnation of Resident Evil 4 in 1998, Devil May Cry quickly became a different beast entirely. Boasting numerous sequels, one ill-received spin off and an anime series, Devil May Cry is comfortably one of the flagship franchises for Capcom. Devil May Cry 5 is the long-awaited fifth entry in the mainline series and, after a little bit of retconning with the previous entries, it's chronologically the latest in the game world. Main character Dante, the wise-cracking monster slayer with his large sword and trademark red coat has returned and this time he's brought a lot of friends with him. In Devil May Cry 4, Nero had an arm called the Devil Bringer, which was his own arm, just demon-ified.