Deep Learning
An AI Designed to Smell and Detect Illness from Human Health
Artificial intelligence (AI) researchers, specifically, a data science team from the Loughborough University are developing a model that has the extraordinary ability to detect a variety of illnesses, just by smelling the human breath. We all know that the human breath can reveal many characteristics of our physical condition, but with this deep learning network, many new cognitive paths in the scientific fields of medicine and AI can be explored. It has been evident that artificial intelligence has made great progress in its abilities to see and listen. Now, even the sense of smell has been conquered! Many substances that we breathe in and out can reveal illnesses that were previously unidentifiable.
DeepMind partners with gaming company for AI research
Artificial intelligence researchers are reaching out to gamers. We have been watching Alphabet's AI division DeepMind--acquired by Google in 2014--for years, particularly since it unveiled AlphaGo Zero, an AI capable of achieving superhuman intelligence without human assistance. Its latest project partners with Unity, one of the world's leading video game development platforms, and aims to research artificial intelligence agents and machine learning, with hopes of using it to improve costly technologies like robotics and self-driving cars. "The partnership will enable DeepMind to develop virtual environments and tasks in support of their fundamental AI research program," said Danny Lange, the vice president of AI and machine learning at Unity Technologies, in a blog post on Sept. 26. "Games and simulations have been a core part of DeepMind's research programme from the very beginning and this approach has already led to significant breakthroughs in AI research," Demis Hassabis, co-founder and CEO of DeepMind, tells the Daily Dot via email.
Robotics Heavyweights Embrace NVIDIA's Jetson AGX Xavier For AI Edge Intelligence
NVIDIA Isaac platform with Jetson Xavier, a computer designed specifically for robotics.NVIDIA Robots are a well-established part of manufacturing but have the opportunity to unlock new efficiencies in industries such as retail, food service and healthcare. To date, robots have primarily been enclosed or segmented into specific areas to protect people from possible injuries. Today, companies want to integrate robotics into various types of workplaces, but this requires a new design paradigm for robotics. Allowing a robot to move freely in an unpredictable environment requires fast, reliable, intelligent computing within the robot. The ability to deliver this level of complex computing at within a small component, at a low price point has held the robotics industry back.
The Battle for Top AI Talent Only Gets Tougher From Here
Andrew Ng helped create two of Silicon Valley's leading artificial intelligence labs. First, he built Google Brain, now the hub of AI research inside the internet giant. Then he built a lab in the Valley for Baidu, the company known as the Google of China. Ng was one of the primary figures behind the enormous and rapid rise of AI over the last five years as everyone from Facebook to Microsoft rebuilt themselves around deep learning. And on Tuesday night, he announced his departure from Baidu.
What Is Deep Learning AI? A Simple Guide With 8 Practical Examples
There's a lot of conversation lately about all the possibilities of machines learning to do things humans currently do in our factories, warehouses, offices and homes. While the technology is evolving--quickly--along with fears and excitement, terms such as artificial intelligence, machine learning and deep learning may leave you perplexed. I hope that this simple guide will help sort out the confusion around deep learning and that the 8 practical examples will help to clarify the actual use of deep learning technology today. The field of artificial intelligence is essentially when machines can do tasks that typically require human intelligence. It encompasses machine learning, where machines can learn by experience and acquire skills without human involvement.
Google DeepMind founder Demis Hassabis: Three truths about AI
The 2016 victory by a Google-built AI at the notoriously complex game of Go was a bold demonstration of the power of modern machine learning. That triumphant AlphaGo system, created by AI research group Google DeepMind, confounded expectations that computers were years away from beating a human champion. But as significant as that achievement was, DeepMind's co-founder Demis Hassabis expects it will be dwarfed by how AI will transform society in the years to come. "I would actually be very pessimistic about the world if something like AI wasn't coming down the road," he said. "The reason I say that is that if you look at the challenges that confront society: climate change, sustainability, mass inequality -- which is getting worse -- diseases, and healthcare, we're not making progress anywhere near fast enough in any of these areas. "Either we need an exponential improvement in human behavior -- less selfishness, less short-termism, more collaboration, more generosity -- or we need an exponential improvement in technology.
Extended Bit-Plane Compression for Convolutional Neural Network Accelerators
Cavigelli, Lukas, Benini, Luca
After the tremendous success of convolutional neural networks in image classification, object detection, speech recognition, etc., there is now rising demand for deployment of these compute-intensive ML models on tightly power constrained embedded and mobile systems at low cost as well as for pushing the throughput in data centers. This has triggered a wave of research towards specialized hardware accelerators. Their performance is often constrained by I/O bandwidth and the energy consumption is dominated by I/O transfers to off-chip memory. We introduce and evaluate a novel, hardware-friendly compression scheme for the feature maps present within convolutional neural networks. We show that an average compression ratio of 4.4x relative to uncompressed data and a gain of 60% over existing method can be achieved for ResNet-34 with a compression block requiring <300 bit of sequential cells and minimal combinational logic.
Visual Curiosity: Learning to Ask Questions to Learn Visual Recognition
Yang, Jianwei, Lu, Jiasen, Lee, Stefan, Batra, Dhruv, Parikh, Devi
In an open-world setting, it is inevitable that an intelligent agent (e.g., a robot) will encounter visual objects, attributes or relationships it does not recognize. In this work, we develop an agent empowered with visual curiosity, i.e. the ability to ask questions to an Oracle (e.g., human) about the contents in images (e.g., What is the object on the left side of the red cube?) and build visual recognition model based on the answers received (e.g., Cylinder). In order to do this, the agent must (1) understand what it recognizes and what it does not, (2) formulate a valid, unambiguous and informative language query (a question) to ask the Oracle, (3) derive the parameters of visual classifiers from the Oracle response and (4) leverage the updated visual classifiers to ask more clarified questions. Specifically, we propose a novel framework and formulate the learning of visual curiosity as a reinforcement learning problem. In this framework, all components of our agent, visual recognition module (to see), question generation policy (to ask), answer digestion module (to understand) and graph memory module (to memorize), are learned entirely end-to-end to maximize the reward derived from the scene graph obtained by the agent as a consequence of the dialog with the Oracle. Importantly, the question generation policy is disentangled from the visual recognition system and specifics of the environment. Consequently, we demonstrate a sort of double generalization. Our question generation policy generalizes to new environments and a new pair of eyes, i.e., new visual system. Trained on a synthetic dataset, our results show that our agent learns new visual concepts significantly faster than several heuristic baselines, even when tested on synthetic environments with novel objects, as well as in a realistic environment.
Modeling Online Discourse with Coupled Distributed Topics
Srivatsan, Akshay, Wojtowicz, Zachary, Berg-Kirkpatrick, Taylor
In this paper, we propose a deep, globally normalized topic model that incorporates structural relationships connecting documents in socially generated corpora, such as online forums. Our model (1) captures discursive interactions along observed reply links in addition to traditional topic information, and (2) incorporates latent distributed representations arranged in a deep architecture, which enables a GPU-based mean-field inference procedure that scales efficiently to large data. We apply our model to a new social media dataset consisting of 13M comments mined from the popular internet forum Reddit, a domain that poses significant challenges to models that do not account for relationships connecting user comments. We evaluate against existing methods across multiple metrics including perplexity and metadata prediction, and qualitatively analyze the learned interaction patterns.
Where and When to Look? Spatio-temporal Attention for Action Recognition in Videos
Meng, Lili, Zhao, Bo, Chang, Bo, Huang, Gao, Tung, Frederick, Sigal, Leonid
Inspired by the observation that humans are able to process videos efficiently by only paying attention when and where it is needed, we propose a novel spatial-temporal attention mechanism for video-based action recognition. For spatial attention, we learn a saliency mask to allow the model to focus on the most salient parts of the feature maps. For temporal attention, we employ a soft temporal attention mechanism to identify the most relevant frames from an input video. Further, we propose a set of regularizers that ensure that our attention mechanism attends to coherent regions in space and time. Our model is efficient, as it proposes a separable spatio-temporal mechanism for video attention, while being able to identify important parts of the video both spatially and temporally. We demonstrate the efficacy of our approach on three public video action recognition datasets. The proposed approach leads to state-of-the-art performance on all of them, including the new large-scale Moments in Time dataset. Furthermore, we quantitatively and qualitatively evaluate our model's ability to accurately localize discriminative regions spatially and critical frames temporally. This is despite our model only being trained with per video classification labels.