Deep Learning
Is Deeper Better only when Shallow is Good?
Malach, Eran, Shalev-Shwartz, Shai
Understanding the power of depth in feed-forward neural networks is an ongoing challenge in the field of deep learning theory. While current works account for the importance of depth for the expressive power of neural-networks, it remains an open question whether these benefits are exploited during a gradient-based optimization process. In this work we explore the relation between expressivity properties of deep networks and the ability to train them efficiently using gradient-based algorithms. We give a depth separation argument for distributions with fractal structure, showing that they can be expressed efficiently by deep networks, but not with shallow ones. These distributions have a natural coarse-to-fine structure, and we show that the balance between the coarse and fine details has a crucial effect on whether the optimization process is likely to succeed. We prove that when the distribution is concentrated on the fine details, gradient-based algorithms are likely to fail. Using this result we prove that, at least in some distributions, the success of learning deep networks depends on whether the distribution can be well approximated by shallower networks, and we conjecture that this property holds in general.
Approximating Optimisation Solutions for Travelling Officer Problem with Customised Deep Learning Network
Shao, Wei, Salim, Flora D., Chan, Jeffrey, Morrison, Sean, Zambetta, Fabio
Deep learning has been extended to a number of new domains with critical success, though some traditional orienteering problems such as the Travelling Salesman Problem (TSP) and its variants are not commonly solved using such techniques. Deep neural networks (DNNs) are a potentially promising and under-explored solution to solve these problems due to their powerful function approximation abilities, and their fast feed-forward computation. In this paper, we outline a method for converting an orienteering problem into a classification problem, and design a customised multi-layer deep learning network to approximate traditional optimisation solutions to this problem. We test the performance of the network on a real-world parking violation dataset, and conduct a generic study that empirically shows the critical architectural components that affect network performance for this problem.
Towards Understanding Chinese Checkers with Heuristics, Monte Carlo Tree Search, and Deep Reinforcement Learning
Liu, Ziyu, Zhou, Meng, Cao, Weiqing, Qu, Qiang, Yeung, Henry Wing Fung, Chung, Vera Yuk Ying
The game of Chinese Checkers is a challenging traditional board game of perfect information that differs from other traditional games in two main aspects: first, unlike Chess, all checkers remain indefinitely in the game and hence the branching factor of the search tree does not decrease as the game progresses; second, unlike Go, there are also no upper bounds on the depth of the search tree since repetitions and backward movements are allowed. Therefore, even in a restricted game instance, the state-space of the game can still be unbounded, making it challenging for a computer program to excel. In this work, we present an approach that effectively combines the use of heuristics, Monte Carlo tree search, and deep reinforcement learning for building a Chinese Checkers agent without the use of any human game-play data. Experiment results show that our agent is competent under different scenarios and reaches the level of experienced human players.
Improving Skin Condition Classification with a Visual Symptom Checker trained using Reinforcement Learning
Akrout, Mohamed, Farahmand, Amir-massoud, Jarmain, Tory, Abid, Latif
We present a visual symptom checker that combines a pre-trained Convolutional Neural Network (CNN) with a Reinforcement Learning (RL) agent as a Question Answering (QA) model. This method enables us to not only increase the classification confidence and accuracy of the visual symptom checker, but also decreases the average number of relevant questions asked to narrow down the differential diagnosis. By combining the CNN output in the form of classification probabilities as a part of the state structure of the simulated patient's environment, a DQN-based RL agent learns to ask the best symptom that maximizes its expected return over symptoms. We demonstrate that our RL approach increases the accuracy more than 20% as compared to the CNN alone, and up to 10% as compared to the decision tree model. We finally show that the RL approach not only outperforms the performance of the decision tree approach but also narrows down the diagonosis faster in terms of the average number of asked questions.
Towards Time-Aware Distant Supervision for Relation Extraction
Jiang, Tianwen, Zhao, Sendong, Liu, Jing, Yao, Jin-Ge, Liu, Ming, Qin, Bing, Liu, Ting, Lin, Chin-Yew
Distant supervision for relation extraction heavily suffers from the wrong labeling problem. To alleviate this issue in news data with the timestamp, we take a new factor time into consideration and propose a novel time-aware distant supervision framework (Time-DS). Time-DS is composed of a time series instance-popularity and two strategies. Instance-popularity is to encode the strong relevance of time and true relation mention. Therefore, instance-popularity would be an effective clue to reduce the noises generated through distant supervision labeling. The two strategies, i.e., hard filter and curriculum learning are both ways to implement instance-popularity for better relation extraction in the manner of Time-DS. The curriculum learning is a more sophisticated and flexible way to exploit instance-popularity to eliminate the bad effects of noises, thus get better relation extraction performance. Experiments on our collected multi-source news corpus show that Time-DS achieves significant improvements for relation extraction.
TF-Replicator: Distributed Machine Learning for Researchers DeepMind
At DeepMind, the Research Platform Team builds infrastructure to empower and accelerate our AI research. Today, we are excited to share how we developed TF-Replicator, a software library that helps researchers deploy their TensorFlow models on GPUs and Cloud TPUs with minimal effort and no previous experience with distributed systems. This blog post gives an overview of the ideas and technical challenges underlying TF-Replicator. For a more comprehensive description, please read our arXiv paper. A recurring theme in recent AI breakthroughs -- from AlphaFold to BigGAN to AlphaStar -- is the need for effortless and reliable scalability.
OpenAI Launches Neural MMO, a Massive Reinforcement Learning Simulator
Artificial intelligence that's beastly at World of Warcraft might not lie too far into the distant future, if OpenAI has its way. The San Francisco research nonprofit today released Neural MMO, a "massively multiagent" virtual training ground that plops agents in the middle of an RPG-like world -- one complete with a resource collection mechanic and player versus player combat. "The game genre of Massively Multiplayer Online Games (MMOs) simulates a large ecosystem of a variable number of players competing in persistent and extensive environments," OpenAI wrote in a blog post. "The inclusion of many agents and species leads to better exploration, divergent niche formation, and greater overall competence."READ
Improving molecular imaging using a deep learning approach
Generating comprehensive molecular images of organs and tumors in living organisms can be performed at ultra-fast speed using a new deep learning approach to image reconstruction developed by researchers at Rensselaer Polytechnic Institute. The research team's new technique has the potential to vastly improve the quality and speed of imaging in live subjects and was the focus of an article recently published in Light: Science and Applications, a Nature journal. Compressed sensing-based imaging is a signal processing technique that can be used to create images based on a limited set of point measurements. Recently, a Rensselaer research team proposed a novel instrumental approach to leverage this methodology to acquire comprehensive molecular data sets, as reported in Nature Photonics. While that approach produced more complete images, processing the data and forming an image could take hours. This latest methodology developed at Rensselaer builds on the previous advancement and has the potential to produce real-time images, while also improving the quality and usefulness of the images produced.
The Achilles' Heel of AI
According to a recent report from AI research and advisory firm Cognilytica, over 80% of the time spent in AI projects are spent dealing with and wrangling data. Even more importantly, and perhaps surprisingly, is how human-intensive much of this data preparation work is. In order for supervised forms of machine learning to work, especially the multi-layered deep learning neural network approaches, they must be fed large volumes of examples of correct data that is appropriately annotated, or "labeled", with the desired output result. For example, if you're trying to get your machine learning algorithm to correctly identify cats inside of images, you need to feed that algorithm thousands of images of cats, appropriately labeled as cats, with the images not having any extraneous or incorrect data that will throw the algorithm off as you build the model.
The Future of Marketing: Predicting Consumer Behavior with AI
Understanding what consumers want and need -- ideally, before they even do -- is an ongoing imperative for marketers. Artificial intelligence (AI) could make that job much easier, especially with the emergence of deep learning. A subset of AI, deep learning has the potential to transform the future of marketing by helping businesses to predict consumer behavior. It's a machine learning method that uses layered or "deep" neural networks, similar to those found in biological brains, to learn skills and solve complex problems faster than people can. It helps computers (or robots) handle "human" tasks, such as perceiving objects, recognizing voices, and translating languages.