Deep Learning
Global Semantic Description of Objects based on Prototype Theory
Pino, Omar Vidal, Nascimento, Erickson Rangel, Campos, Mario Fernando Montenegro
In this paper, we introduce a novel semantic description approach inspired on Prototype Theory foundations. We propose a Computational Prototype Model (CPM) that encodes and stores the central semantic meaning of objects category: the semantic prototype. Also, we introduce a Prototype-based Description Model that encodes the semantic meaning of an object while describing its features using our CPM model. Our description method uses semantic prototypes computed by CNN-classifications models to create discriminative signatures that describe an object highlighting its most distinctive features within the category. Our experiments show that: i) our CPM model (semantic prototype + distance metric) is able to describe the internal semantic structure of objects categories; ii) our semantic distance metric can be understood as the object visual typicality score within a category; iii) our descriptor encoding is semantically interpretable and significantly outperforms other image global encodings in clustering and classification tasks.
Multi-modal Active Learning From Human Data: A Deep Reinforcement Learning Approach
Rudovic, Ognjen, Zhang, Meiru, Schuller, Bjorn, Picard, Rosalind W.
Human behavior expression and experience are inherently multi-modal, and characterized by vast individual and contextual heterogeneity. To achieve meaningful human-computer and human-robot interactions, multi-modal models of the users states (e.g., engagement) are therefore needed. Most of the existing works that try to build classifiers for the users states assume that the data to train the models are fully labeled. Nevertheless, data labeling is costly and tedious, and also prone to subjective interpretations by the human coders. This is even more pronounced when the data are multi-modal (e.g., some users are more expressive with their facial expressions, some with their voice). Thus, building models that can accurately estimate the users states during an interaction is challenging. To tackle this, we propose a novel multi-modal active learning (AL) approach that uses the notion of deep reinforcement learning (RL) to find an optimal policy for active selection of the users data, needed to train the target (modality-specific) models. We investigate different strategies for multi-modal data fusion, and show that the proposed model-level fusion coupled with RL outperforms the feature-level and modality-specific models, and the naive AL strategies such as random sampling, and the standard heuristics such as uncertainty sampling. We show the benefits of this approach on the task of engagement estimation from real-world child-robot interactions during an autism therapy. Importantly, we show that the proposed multi-modal AL approach can be used to efficiently personalize the engagement classifiers to the target user using a small amount of actively selected users data.
DeepBundle: Fiber Bundle Parcellation with Graph Convolution Neural Networks
Liu, Feihong, Feng, Jun, Chen, Geng, Wu, Ye, Hong, Yoonmi, Yap, Pew-Thian, Shen, Dinggang
Parcellation of whole-brain tractography streamlines is an important step for tract-based analysis of brain white matter microstructure. Existing fiber parcellation approaches rely on accurate registration between an atlas and the tractograms of an individual, however, due to large individual differences, accurate registration is hard to guarantee in practice. To resolve this issue, we propose a novel deep learning method, called DeepBundle, for registration-free fiber parcellation. Our method utilizes graph convolution neural networks (GCNNs) to predict the parcellation label of each fiber tract. GCNNs are capable of extracting the geometric features of each fiber tract and harnessing the resulting features for accurate fiber parcellation and ultimately avoiding the use of atlases and any registration method. We evaluate DeepBundle using data from the Human Connectome Project. Experimental results demonstrate the advantages of DeepBundle and suggest that the geometric features extracted from each fiber tract can be used to effectively parcellate the fiber tracts.
Building a Computer Mahjong Player via Deep Convolutional Neural Networks
Gao, Shiqi, Okuya, Fuminori, Kawahara, Yoshihiro, Tsuruoka, Yoshimasa
The evaluation function for imperfect information games is always hard to define but owns a significant impact on the playing strength of a program. Deep learning has made great achievements these years, and already exceeded the top human players' level even in the game of Go. In this paper, we introduce a new data model to represent the available imperfect information on the game table, and construct a well-designed convolutional neural network for game record training. We choose the accuracy of tile discarding which is also called as the agreement rate as the benchmark for this study. Our accuracy on test data reaches 70.44%, while the state-of-art baseline is 62.1% reported by Mizukami and Tsuruoka (2015), and is significantly higher than previous trials using deep learning, which shows the promising potential of our new model. For the AI program building, besides the tile discarding strategy, we adopt similar predicting strategies for other actions such as stealing (pon, chi, and kan) and riichi. With the simple combination of these several predicting networks and without any knowledge about the concrete rules of the game, a strength evaluation is made for the resulting program on the largest Japanese Mahjong site `Tenhou'. The program has achieved a rating of around 1850, which is significantly higher than that of an average human player and of programs among past studies.
Hailo launches its newest deep learning chip – TechCrunch
Hailo, a Tel Aviv-based AI chipmaker, today announced that it is now sampling its Hailo -8 chips, the first of its deep learning processors. The new chip promises up to 26 tera operations per second (TOPS), and the company is now testing it with a number of select customers, mostly in the automotive industry. Hailo first appeared on the radar last year, when it raised a $12.5 million Series A round. At the time, the company was still waiting for the first samples of its chips. Now, the company says that the Hailo-8 will outperform all other edge processors and do so at a smaller size and with fewer memory requirements.
Amazon Unveils Novel Alexa Dialog Modeling for Natural, Cross-Skill Conversations : Alexa Blogs
Today, customer exchanges with Alexa are generally either one-shot requests, like "Alexa, what's the weather?", or interactions that require multiple requests to complete more complex tasks. An Alexa customer planning a family movie night out, for example, must interact independently with multiple skills to find a list of local theaters playing a particular movie, identify a restaurant near one of them, and then purchase movie tickets, book a table, and perhaps order a ride. The cognitive burden of carrying information across skills -- such as time, number of people, and location -- rests with the customer. "We envision a world where customers will converse more naturally with Alexa: seamlessly transitioning between skills, asking questions, making choices, and speaking the same way they would with a friend, family member, or co-worker," says Rohit Prasad, Alexa vice president and head scientist. "Our objective is to shift the cognitive burden from the customer to Alexa."
Training neural belief-propagation decoders for quantum error-correcting codes
Two researchers at Université de Sherbrooke, in Canada, have recently developed and trained neural belief-propagation (BP) decoders for quantum low-density parity-check (LDPC) codes. Their study, outlined in a paper published in Physical Review Letters, suggests that training can enhance the performance of BP decoders significantly, helping to solve issues that are commonly associated with their application in quantum research. "Ten years ago, I wrote an article with Yeojin Chung explaining how standard decoding algorithms for LDPC codes, which are broadly used in classical communication, would fail in the quantum setting," David Poulin, one of the researchers who carried out the study, told Phys.org. "This problem has been obsessing me ever since. Recently, people have started to investigate the use of neural networks to decode quantum codes, but they all focused on a problem (decoding topological codes) that already had a number of good human-designed solutions. This was the perfect occasion to revisit my favorite open problem and use neural networks to decode quantum codes that had no previously known decoder."
The MIT Artificial Intelligence Podcast
MIT research scientist Lex Fridman hosts a great podcast on artificial intelligence. The following article is a short summary about Fridman and his podcast, which I highly recommend. Lex Fridman Mr. Fridman is not only a great moderator, he also does active research at MIT and teach courses on deep learning. Most of his research publications deal with the topic of autonomous vehicles. Topics of the Podcast Fridman has a number of well-known guests in his podcast and talks about many interesting topics.
Creating an AI can be five times worse for the planet than a car
Training artificial intelligence is an energy intensive process. New estimates suggest that the carbon footprint of training a single AI is as much as 284 tonnes of carbon dioxide equivalent – five times the lifetime emissions of an average car. Emma Strubell at the University of Massachusetts Amherst in the US and colleagues have assessed the energy consumption required to train four large neural networks, a type of AI used for processing language. Language-processing AIs underpin the algorithms that power Google Translate as well as OpenAI's GPT-2 text generator, which can convincingly pen fake news articles when given a few lines of text. These AIs are trained via deep learning, which involves processing vasts amounts of data. "In order to learn something as complex as language, the models have to be large," says Strubell.