Deep Learning
Deep Variational Transfer: Transfer Learning through Semi-supervised Deep Generative Models
Belhaj, Marouan, Protopapas, Pavlos, Pan, Weiwei
In real-world applications, it is often expensive and time-consuming to obtain labeled examples. In such cases, knowledge transfer from related domains, where labels are abundant, could greatly reduce the need for extensive labeling efforts. In this scenario, transfer learning comes in hand. In this paper, we propose Deep Variational Transfer (DVT), a variational autoencoder that transfers knowledge across domains using a shared latent Gaussian mixture model. Thanks to the combination of a semi-supervised ELBO and parameters sharing across domains, we are able to simultaneously: (i) align all supervised examples of the same class into the same latent Gaussian Mixture component, independently from their domain; (ii) predict the class of unsupervised examples from different domains and use them to better model the occurring shifts. We perform tests on MNIST and USPS digits datasets, showing DVT's ability to perform transfer learning across heterogeneous datasets. Additionally, we present DVT's top classification performances on the MNIST semi-supervised learning challenge. We further validate DVT on a astronomical datasets. DVT achieves states-of-the-art classification performances, transferring knowledge across real stars surveys datasets, EROS, MACHO and HiTS, . In the worst performance, we double the achieved F1-score for rare classes. These experiments show DVT's ability to tackle all major challenges posed by transfer learning: different covariate distributions, different and highly imbalanced class distributions and different feature spaces.
Training Complex Models with Multi-Task Weak Supervision
Ratner, Alexander, Hancock, Braden, Dunnmon, Jared, Sala, Frederic, Pandey, Shreyash, Ré, Christopher
As machine learning models continue to increase in complexity, collecting large hand-labeled training sets has become one of the biggest roadblocks in practice. Instead, weaker forms of supervision that provide noisier but cheaper labels are often used. However, these weak supervision sources have diverse and unknown accuracies, may output correlated labels, and may label different tasks or apply at different levels of granularity. We propose a framework for integrating and modeling such weak supervision sources by viewing them as labeling different related sub-tasks of a problem, which we refer to as the multi-task weak supervision setting. We show that by solving a matrix completion-style problem, we can recover the accuracies of these multi-task sources given their dependency structure, but without any labeled data, leading to higher-quality supervision for training an end model. Theoretically, we show that the generalization error of models trained with this approach improves with the number of unlabeled data points, and characterize the scaling with respect to the task and dependency structures. On three fine-grained classification problems, we show that our approach leads to average gains of 20.2 points in accuracy over a traditional supervised approach, 6.8 points over a majority vote baseline, and 4.1 points over a previously proposed weak supervision method that models tasks separately.
DeepMind's AlphaZero now showing human-like intuition in historical 'turning point' for AI
DeepMind's artificial intelligence programme AlphaZero is now showing signs of human-like intuition and creativity, in what developers have hailed as'turning point' in history. The computer system amazed the world last year when it mastered the game of chess from scratch within just four hours, despite not being programmed how to win. But now, after a year of testing and analysis by chess grandmasters, the machine has developed a new style of play unlike anything ever seen before, suggesting the programme is now improvising like a human. Unlike the world's best chess machine - Stockfish - which calculates millions of possible outcomes as it plays, AlphaZero learns from its past successes and...
A different kind of (deep) learning: part 1 – Towards Data Science
Deep learning has truly reshuffled things in machine learning field, and specifically in image recognition tasks. In 2012, Alex-net has initiated a (still far from ending) race towards solving, or at least significantly improving, computer vision tasks. Each of these research paths improves training quality (speed, accuracy, sometimes generalization), but it seems that doing more of the same thing may result in some gradual improvements, but not a in significant breakthrough. On the other hand, growing body of work in deep learning shows that there are significant flaws in current methods, especially in terms of generalization, e.g this recent one: generalization failure when objects are rotated: So there seems to be a need of improvements that are a bit more aggressive. Or perhaps expanding the research spectrum to ideas that may be a bit riskier.
Baidu has created a 'no code' platform to make building AI models easier - SiliconANGLE
Hoping to make up ground in the hotly contested artificial intelligence battleground, Chinese Internet giant Baidu Inc. is releasing a tool that allows businesses to create and deploy AI models without coding skills. Announced Saturday, EZDL is a "no-code platform to build custom machine learning models," designed with ease of use and security in mind, the company said. "EZDL is a service platform that allows users to build custom machine learning models with a drag-and-drop interface," Yongkang Xie, tech lead of Baidu EZDL, said in a statement. "It takes only four steps to train a deep learning model, built specifically for your unique business needs." EZDL is focused on three important aspects of machine learning that have already been popularized in numerous software apps: image classification, sound classification and object detection.
DeepMind's Go playing software can now beat you at two more games
DeepMind's AI Go master has taught itself new tricks. The latest version of the machine learning software, dubbed AlphaZero, can now also beat the world's best at chess and shogi – a Japanese game that is similar to chess but played on a bigger board with more pieces. DeepMind, a sister company to Google, claims that it is the first machine-learning system that can learn to do more than one task with superhuman ability. AlphaGo made headlines in 2016 when it beat the world's best players at a game long thought too hard for computers to crack. Then came AlphaGo Zero, which not only out-played AlphaGo but taught itself to do so without ever having seen a human play the game.
Why Robot Brains Need Symbols - Issue 67: Reboot
Humans can generalize a wide range of universals to arbitrary novel instances. They appear to do so in many areas of language (including syntax, morphology, and discourse) and thought (including transitive inference, entailments, and class-inclusion relationships). Advocates of symbol manipulation assume that the mind instantiates symbol-manipulating mechanisms including symbols, categories, and variables, and mechanisms for assigning instances to categories and representing and extending relationships between variables. This account provides a straightforward framework for understanding how universals are extended to arbitrary novel instances. Current eliminative connectionist models map input vectors to output vectors using the back-propagation algorithm (or one of its variants). To generalize universals to arbitrary novel instances, these models would need to generalize outside the training space. These models cannot generalize outside the training space. Therefore, current eliminative connectionist models cannot account for those cognitive phenomena that involve universals that can be freely extended to arbitrary cases. Richard Evans and Edward Grefenstette's recent paper at DeepMind, building on Joel Grus's blog post on the game Fizz-Buzz, follows remarkably similar lines, concluding that a canonical multilayer network was unable to solve the simple game on its own "because it did not capture the general, universally quantified rules needed to understand this task"--exactly what I said in 1998.
Amazon AWS Leaps Forward In Cloudy AI At re:Invent
Amazon Web Services hosted its annual conference last week. Over 50,000 attendees converged on Las Vegas to partake in the massive 4-day agenda of all things cloudy. In the realm of Artificial Intelligence (AI), AWS announced over a dozen completely new services for building and running smart applications, adding significantly to what was already a rich portfolio. "We want to help all of our customers embrace machine learning, no matter their size, budget, experience, or skill level," said Swami Sivasubramanian, Vice President of Amazon Machine Learning. AWS hopes to extend its lead in hosting the market for enterprises and AI startups.
EEG Classification based on Image Configuration in Social Anxiety Disorder
Mokatren, Lubna Shibly, Ansari, Rashid, Cetin, Ahmet Enis, Leow, Alex D., Ajilore, Olusola, Klumpp, Heide, Vural, Fatos T. Yarman
The problem of detecting the presence of Social Anxiety Disorder (SAD) using Electroencephalography (EEG) for classification has seen limited study and is addressed with a new approach that seeks to exploit the knowledge of EEG sensor spatial configuration. Two classification models, one which ignores the configuration (model 1) and one that exploits it with different interpolation methods (model 2), are studied. Performance of these two models is examined for analyzing 34 EEG data channels each consisting of five frequency bands and further decomposed with a filter bank. The data are collected from 64 subjects consisting of healthy controls and patients with SAD. Validity of our hypothesis that model 2 will significantly outperform model 1 is borne out in the results, with accuracy $6$--$7\%$ higher for model 2 for each machine learning algorithm we investigated. Convolutional Neural Networks (CNN) were found to provide much better performance than SVM and kNNs.
From Word To Sense Embeddings: A Survey on Vector Representations of Meaning
Camacho-Collados, Jose, Pilehvar, Mohammad Taher
Over the past years, distributed semantic representations have proved to be effective and flexible keepers of prior knowledge to be integrated into downstream applications. This survey focuses on the representation of meaning. We start from the theoretical background behind word vector space models and highlight one of their major limitations: the meaning conflation deficiency, which arises from representing a word with all its possible meanings as a single vector. Then, we explain how this deficiency can be addressed through a transition from the word level to the more fine-grained level of word senses (in its broader acceptation) as a method for modelling unambiguous lexical meaning. We present a comprehensive overview of the wide range of techniques in the two main branches of sense representation, i.e., unsupervised and knowledge-based. Finally, this survey covers the main evaluation procedures and applications for this type of representation, and provides an analysis of four of its important aspects: interpretability, sense granularity, adaptability to different domains and compositionality.