Deep Learning
Exploring the Landscape of Spatial Robustness
Engstrom, Logan, Tran, Brandon, Tsipras, Dimitris, Schmidt, Ludwig, Madry, Aleksander
The study of adversarial robustness has so far largely focused on perturbations bound in p-norms. However, state-of-the-art models turn out to be also vulnerable to other, more natural classes of perturbations such as translations and rotations. In this work, we thoroughly investigate the vulnerability of neural network--based classifiers to rotations and translations. While data augmentation offers relatively small robustness, we use ideas from robust optimization and test-time input aggregation to significantly improve robustness. Finally we find that, in contrast to the p-norm case, first-order methods cannot reliably find worst-case perturbations. This highlights spatial robustness as a fundamentally different setting requiring additional study. Code available at https://github.com/MadryLab/adversarial_spatial and https://github.com/MadryLab/spatial-pytorch.
Emergent Tool Use From Multi-Agent Autocurricula
Baker, Bowen, Kanitscheider, Ingmar, Markov, Todor, Wu, Yi, Powell, Glenn, McGrew, Bob, Mordatch, Igor
Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordination. We find clear evidence of six emergent phases in agent strategy in our environment, each of which creates a new pressure for the opposing team to adapt; for instance, agents learn to build multi-object shelters using moveable boxes which in turn leads to agents discovering that they can overcome obstacles using ramps. We further provide evidence that multi-agent competition may scale better with increasing environment complexity and leads to behavior that centers around far more human-relevant skills than other self-supervised reinforcement learning methods such as intrinsic motivation. Finally, we propose transfer and fine-tuning as a way to quantitatively evaluate targeted capabilities, and we compare hide-and-seek agents to both intrinsic motivation and random initialization baselines in a suite of domain-specific intelligence tests.
They Might NOT Be Giants: Crafting Black-Box Adversarial Examples with Fewer Queries Using Particle Swarm Optimization
Mosli, Rayan, Wright, Matthew, Yuan, Bo, Pan, Yin
--Machine learning models have been found to be susceptible to adversarial examples that are often indistinguishable from the original inputs. These adversarial examples are created by applying adversarial perturbations to input samples, which would cause them to be misclassified by the target models. Attacks that search and apply the perturbations to create adversarial examples are performed in both white-box and black-box settings, depending on the information available to the attacker about the target. For black-box attacks, the only capability available to the attacker is the ability to query the target with specially crafted inputs and observing the labels returned by the model. Current black-box attacks either have low success rates, requires a high number of queries, or produce adversarial examples that are easily distinguishable from their sources. In this paper, we present AdversarialPSO, a black-box attack that uses fewer queries to create adversarial examples with high success rates. AdversarialPSO is based on the evolutionary search algorithm Particle Swarm Optimization, a population-based gradient-free optimization algorithm. It is flexible in balancing the number of queries submitted to the target vs the quality of imperceptible adversarial examples. The attack has been evaluated using the image classification benchmark datasets CIF AR-10, MNIST, and Imagenet, achieving success rates of 99.6%, 96.3%, and 82.0%, respectively, while submitting substantially fewer queries than the state-of-the-art. We also present a black-box method for isolating salient features used by models when making classifications. This method, called Swarms with Individual Search Spaces or SWISS, creates adversarial examples by finding and modifying the most important features in the input. The purpose of these two attacks is to help evaluate the robustness of machine learning models and to encourage the exploration of much-needed defenses. Deep learning (DL) is being used to solve a wide variety of problems in many different domains, such as image classification [1], malware detection [2], speech recognition [3], and medicine [4]. Despite state-of-the-art performances, DL models have been shown to suffer from a general flaw that makes them vulnerable to external attack. Adversaries can cause models to misclassify inputs by applying small perturbations to samples at test time [5].
Deep Neural Networks for Choice Analysis: Architectural Design with Alternative-Specific Utility Functions
Whereas deep neural network (DNN) is increasingly applied to choice analysis, it is challenging to reconcile domain-specific behavioral knowledge with generic-purpose DNN, to improve DNN's interpretability and predictive power, and to identify effective regularization methods for specific tasks. This study designs a particular DNN architecture with alternative-specific utility functions (ASU-DNN) by using prior behavioral knowledge. Unlike a fully connected DNN (F-DNN), which computes the utility value of an alternative k by using the attributes of all the alternatives, ASU-DNN computes it by using only k's own attributes. Theoretically, ASU-DNN can dramatically reduce the estimation error of F-DNN because of its lighter architecture and sparser connectivity. Empirically, ASU-DNN has 2-3% higher prediction accuracy than F-DNN over the whole hyperparameter space in a private dataset that we collected in Singapore and a public dataset in R mlogit package. The alternative-specific connectivity constraint, as a domain-knowledge-based regularization method, is more effective than the most popular generic-purpose explicit and implicit regularization methods and architectural hyperparameters. ASU-DNN is also more interpretable because it provides a more regular substitution pattern of travel mode choices than F-DNN does. The comparison between ASU-DNN and F-DNN can also aid in testing the behavioral knowledge. Our results reveal that individuals are more likely to compute utility by using an alternative's own attributes, supporting the long-standing practice in choice modeling. Overall, this study demonstrates that prior behavioral knowledge could be used to guide the architecture design of DNN, to function as an effective domain-knowledge-based regularization method, and to improve both the interpretability and predictive power of DNN in choice analysis.
Learning Invariants through Soft Unification
Cingillioglu, Nuri, Russo, Alessandra
Human reasoning involves recognising common underlying principles across many examples by utilising variables. The byproducts of such reasoning are invariants that capture patterns across examples such as "if someone went somewhere then they are there" without mentioning specific people or places. Humans learn what variables are and how to use them at a young age, and the question this paper addresses is whether machines can also learn and use variables solely from examples without requiring human pre-engineering. We propose Unification Networks that incorporate soft unification into neural networks to learn variables and by doing so lift examples into invariants that can then be used to solve a given task. We evaluate our approach on four datasets to demonstrate that learning invariants captures patterns in the data and can improve performance over baselines. Humans have the ability to process symbolic knowledge and maintain symbolic thought (Unger & Deacon, 1998). When reasoning, humans do not require combinatorial enumeration of examples but instead utilise invariant patterns with placeholders replacing specific entities. Symbolic cognitive models (Lewis, 1999) embrace this perspective with the human mind seen as an information processing system operating on formal symbols such as reading a stream of tokens in natural language. The language of thought hypothesis (Morton & Fodor, 1978) frames human thought as a structural construct with varying sub-components such as "X went to Y".
MFCC-based Recurrent Neural Network for Automatic Clinical Depression Recognition and Assessment from Speech
Rejaibi, Emna, Komaty, Ali, Meriaudeau, Fabrice, Agrebi, Said, Othmani, Alice
MFCC-based Recurrent Neural Network for Automatic Clinical Depression Recognition and Assessment from Speech Emna Rejaibi a,b,c, Ali Komaty d, Fabrice Meriaudeau e, Said Agrebi c, Alice Othmani a a Universit e Paris-Est, LISSI, UPEC, 94400 Vitry sur Seine, France b INSAT Institut National des Sciences Appliqu ees et de T echnologie, Centre Urbain Nord BP 676-1080, Tunis, Tunisie c Y obitrust, T echnopark El Gazala B11 Route de Raoued Km 3.5, 2088 Ariana, Tunisie d University of Sciences and Arts in Lebanon, Ghobeiry, Liban e Universit e de Bourgogne Franche Comt e, ImvIA EA7535/ IFTIM Abstract Major depression, also known as clinical depression, is a constant sense of despair and hopelessness. It is a major mental disorder that can a ff ect people of any age including children and that a ff ect negatively person's personal life, work life, social life and health conditions. Globally, over 300 million people of all ages are estimated to su ff er from clinical depression. A deep recurrent neural network-based framework is presented in this paper to detect depression and to predict its severity level from speech. Low-level and high-level audio features are extracted from audio recordings to predict the 24 scores of the Patient Health Questionnaire (a depression assessment test) and the binary class of depression diagnosis. To overcome the problem of the small size of Speech Depression Recognition (SDR) datasets, data augmentation techniques are used to expand the labeled training set and also transfer learning is performed where the proposed model is trained on a related task and reused as starting point for the proposed model on SDR task. The proposed framework is evaluated on the DAIC-WOZ corpus of the A VEC2017 challenge and promising results are obtained. An overall accuracy of 76.27% with a root mean square error of 0.4 is achieved in assessing depression, while a root mean square error of 0.168 is achieved in predicting the depression severity levels. Introduction Depression is a mental disorder caused by several factors: psychological, social or even physical factors. Psychological factors are related to permanent stress and the inability to successfully cope with di fficult situations. Social factors concern relationship struggles with family or friends and physical factors cover head injuries. Depression describes a loss of interest in every exciting and joyful aspect of everyday life. Mood disorders and mood swings are temporary mental states taking an essential part of daily events, whereas, depression is more permanent and can lead to suicide at its extreme severity levels.
A Self-Attentional Neural Architecture for Code Completion with Multi-Task Learning
Liu, Fang, Li, Ge, Wei, Bolin, Xia, Xin, Li, Ming, Fu, Zhiyi, Jin, Zhi
--Code completion, one of the most useful features in the integrated development environments, can accelerate software development by suggesting the libraries, APIs, method names in real-time. Recent studies have shown that statistical language models can improve the performance of code completion tools through learning from large-scale software repositories. However, these models suffer from three major drawbacks: a) The hierarchical structural information of the programs is not fully utilized in the program's representation; b) In programs, the semantic relationships can be very long, existing LSTM based language models are not sufficient to model the long-term dependency. In this paper, we present a novel method that introduces the hierarchical structural information into the representation of programs by considering the path from the predicting node to the root node. T o capture the long-term dependency in the input programs, we apply Transformer-XL network as the base language model. Besides, we creatively propose a Multi-T ask Learning (MTL) framework to learn two related tasks in code completion jointly, where knowledge acquired from one task could be beneficial to another task. Experiments on three real-world datasets demonstrate the effectiveness of our model when compared with state-of-the-art methods. As the complexity and scale of the software developing continue to grow, code completion has become an essential feature of Integrated Development Environments (IDEs). It can speed up the process of software development by suggesting the next probable token based on existing code. However, traditional code completion tools rely on compile-time type information or heuristics rules to make recommendations [1], [2], which are costly and could not well capture human's programming patterns. To alleviate this problem, code completion research started to focus on learning from large-scale codebases in recent years. Based on the observation of source code's repeatability and predictability [3], statistical language models are generally used for modeling source code. N-gram is one of the most widely used language models [3]-[5]. Most recently, as the success of deep learning, source code modeling techniques have turned to Recurrent Neural Network (RNN) based models [2], [6]. In these models, a piece of source code is represented as source code token sequence or Abstract Syntactic Tree (AST) node sequence. Given a partial code sequence, the model computes the probability of the next token or AST node and recommends the one with the highest probability.
Microsoft Vision AI Developer Kit Simplifies Building Vision-Based Deep Learning Projects
For the Vision AI Developer Kit, Microsoft and Qualcomm have partnered to simplify training and deploying computer vision-based AI models. Developers can use Microsoft's cloud-based AI and IoT services on Azure to train models while deploying them on the smart camera edge device powered by a Qualcomm's AI accelerator. Let's take a close look at Vision AI Developer Kit. The Vision AI Developer Kit not only looks stylish and sophisticated, but also boasts of an impressive configuration. The kit is powered by a Qualcomm Snapdragon 603 processor, 4GB of LDDR4X memory and 16GB of eMMC storage.
Huawei Wants To Tackle NVIDIA And Google With A Solid AI Strategy
It supports mainstream deep learning frameworks such as TensorFlow, PyTorch and PaddlePaddle. Tensor Engine and its operators are Huawei's equivalent of NVIDIA cuDNN, a library that makes CUDA accessible to AI developers. MindSpore is Huawei's own unified training/inference framework architected to be design-friendly, operations-friendly that's adaptable to multiple scenarios. It includes core subsystems, such as a model library, graph compute, and tuning toolkit; a unified, distributed architecture for machine learning, deep learning, and reinforcement learning; a flexible program interface along with support for multiple languages. MindSpore is highly optimized for Ascend chips. It takes advantage of the hardware innovations that went into the design of the AI chips.
How deep learning can maximize player performance in sports
We treat athletes as if they are real-life superheroes that overcome physical challenges to achieve greatness in their respective sports. Today's athletes are physically faster, stronger and more agile than the generation before, but something is wrong. We have not made the same progress in improving athletes' mental skills and health as we have physical skills and health. The focus of any individual or team sport is to maximize player performance. In our sports culture, we are obsessed with team and player statistics using traditional measures in each sport.