Deep Learning
Certified Defenses: Why Tighter Relaxations May Hurt Training?
Jovanović, Nikola, Balunović, Mislav, Baader, Maximilian, Vechev, Martin
Certified defenses based on convex relaxations are an established technique for training provably robust models. The key component is the choice of relaxation, varying from simple intervals to tight polyhedra. Paradoxically, however, it was empirically observed that training with tighter relaxations can worsen certified robustness. While several methods were designed to partially mitigate this issue, the underlying causes are poorly understood. In this work we investigate the above phenomenon and show that tightness may not be the determining factor for reduced certified robustness. Concretely, we identify two key features of relaxations that impact training dynamics: continuity and sensitivity. We then experimentally demonstrate that these two factors explain the drop in certified robustness when using popular relaxations. Further, we show, for the first time, that it is possible to successfully train with tighter relaxations (i.e., triangle), a result supported by our two properties. Overall, we believe the insights of this work can help drive the systematic discovery of new effective certified defenses.
Rethinking Eye-blink: Assessing Task Difficulty through Physiological Representation of Spontaneous Blinking
Continuous assessment of task difficulty and mental workload is essential in improving the usability and accessibility of interactive systems. Eye tracking data has often been investigated to achieve this ability, with reports on the limited role of standard blink metrics. Here, we propose a new approach to the analysis of eye-blink responses for automated estimation of task difficulty. The core module is a time-frequency representation of eye-blink, which aims to capture the richness of information reflected on blinking. In our first study, we show that this method significantly improves the sensitivity to task difficulty. We then demonstrate how to form a framework where the represented patterns are analyzed with multi-dimensional Long Short-Term Memory recurrent neural networks for their non-linear mapping onto difficulty-related parameters. This framework outperformed other methods that used hand-engineered features. This approach works with any built-in camera, without requiring specialized devices. We conclude by discussing how Rethinking Eye-blink can benefit real-world applications.
Exploiting Spline Models for the Training of Fully Connected Layers in Neural Network
Mo, Kanya, Zheng, Shen, Wang, Xiwei, Wang, Jinghua, Schewe, Klaus-Dieter
The fully connected (FC) layer, one of the most fundamental modules in artificial neural networks (ANN), is often considered difficult and inefficient to train due to issues including the risk of overfitting caused by its large amount of parameters. Based on previous work studying ANN from linear spline perspectives, we propose a spline-based approach that eases the difficulty of training FC layers. Given some dataset, we first obtain a continuous piece-wise linear (CPWL) fit through spline methods such as multivariate adaptive regression spline (MARS). Next, we construct an ANN model from the linear spline model and continue to train the ANN model on the dataset using gradient descent optimization algorithms. Our experimental results and theoretical analysis show that our approach reduces the computational cost, accelerates the convergence of FC layers, and significantly increases the interpretability of the resulting model (FC layers) compared with standard ANN training with random parameter initialization followed by gradient descent optimizations.
Improving Object Detection in Art Images Using Only Style Transfer
Kadish, David, Risi, Sebastian, Løvlie, Anders Sundnes
Despite recent advances in object detection using deep learning neural networks, these neural networks still struggle to identify objects in art images such as paintings and drawings. This challenge is known as the cross depiction problem and it stems in part from the tendency of neural networks to prioritize identification of an object's texture over its shape. In this paper we propose and evaluate a process for training neural networks to localize objects - specifically people - in art images. We generate a large dataset for training and validation by modifying the images in the COCO dataset using AdaIn style transfer. This dataset is used to fine-tune a Faster R-CNN object detection network, which is then tested on the existing People-Art testing dataset. The result is a significant improvement on the state of the art and a new way forward for creating datasets to train neural networks to process art images.
Transformer Language Models with LSTM-based Cross-utterance Information Representation
Sun, G., Zhang, C., Woodland, P. C.
The effective incorporation of cross-utterance information has the potential to improve language models (LMs) for automatic speech recognition (ASR). To extract more powerful and robust cross-utterance representations for the Transformer LM (TLM), this paper proposes the R-TLM which uses hidden states in a long short-term memory (LSTM) LM. To encode the cross-utterance information, the R-TLM incorporates an LSTM module together with a segment-wise recurrence in some of the Transformer blocks. In addition to the LSTM module output, a shortcut connection using a fusion layer that bypasses the LSTM module is also investigated. The proposed system was evaluated on the AMI meeting corpus, the Eval2000 and the RT03 telephone conversation evaluation sets. The best R-TLM achieved 0.9%, 0.6%, and 0.8% absolute WER reductions over the single-utterance TLM baseline, and 0.5%, 0.3%, 0.2% absolute WER reductions over a strong cross-utterance TLM baseline on the AMI evaluation set, Eval2000 and RT03 respectively. Improvements on Eval2000 and RT03 were further supported by significance tests. R-TLMs were found to have better LM scores on words where recognition errors are more likely to occur. The R-TLM WER can be further reduced by interpolation with an LSTM-LM.
SCOUT: Socially-COnsistent and UndersTandable Graph Attention Network for Trajectory Prediction of Vehicles and VRUs
Carrasco, Sandra, Llorca, David Fernández, Sotelo, Miguel Ángel
Autonomous vehicles navigate in dynamically changing environments under a wide variety of conditions, being continuously influenced by surrounding objects. Modelling interactions among agents is essential for accurately forecasting other agents' behaviour and achieving safe and comfortable motion planning. In this work, we propose SCOUT, a novel Attention-based Graph Neural Network that uses a flexible and generic representation of the scene as a graph for modelling interactions, and predicts socially-consistent trajectories of vehicles and Vulnerable Road Users (VRUs) under mixed traffic conditions. We explore three different attention mechanisms and test our scheme with both bird-eye-view and on-vehicle urban data, achieving superior performance than existing state-of-the-art approaches on InD and ApolloScape Trajectory benchmarks. Additionally, we evaluate our model's flexibility and transferability by testing it under completely new scenarios on RounD dataset. The importance and influence of each interaction in the final prediction is explored by means of Integrated Gradients technique and the visualization of the attention learned.
Min-Max-Plus Neural Networks
Conventional artificial neural networks typically have a fixed nonlinear activation function that applies to all neurons. Introducing a trainable nonlinear part of the network usually further enhances fitting capability of the network. For example, the seminal work of He et al. [1] showed for the first time that human-level performance on ImageNet Classification (experimentally with an error rate of 5.1%) could be surpassed by the performance of a large scale deep neural network (experimentally with an error rate of 4.94%). A key ingredient of their work is to make an extension of the classical Rectified Linear Unit (ReLU) as the nonlinear activation function to Parametric Rectified Linear Unit (PReLU) in which the slope of the negative part of the input is learnable. In this paper, we propose a new framework of neural networks called Min-Max-Plus Neural Networks (MMP-NNs) whose nonlinear part is systematically complexified. The mathematical foundation of this model is called tropical mathematics [2] which is a fast developing area in mathematics and whose connections to neural networks have been established only very recently. A special feature of tropical mathematics is that the operations of usual multiplications and additions degenerate to operations of additions (called "tropical multiplications") and min/max operations (called "tropical additions") respectively. Consequently, the nonlinear part of an MMP-NN only involves additions and min/max operations. Since the nonlinear part of an MMP-NN is trainable, the overall fitting capability of the model is determined by the fitting capabilities of both the linear and nonlinear parts of the network.
Good health and well-being: summarising AI and robotics in healthcare – diagnostics, personalised care, drug discovery, and more
In December 2020 we announced the launch of our focus series AI for Good: UN sustainable development goals (SDGs). Each month we pick a different sustainable development goal (SDG) and highlight work in that area. Following a terrific response to our first focus on "good health and well-being", we bring you the first of our monthly summary articles where we provide a brief overview of the topic and some highlights from the series. With the COVID-19 pandemic dominating our lives at the moment, research relating to the disease has rightly received considerable coverage in our focus series. In partnership with CLAIRE's COVID-19 taskforce initiative we have brought you articles covering the formation of the taskforce, and about some of the research from the participants.
Combining convolutional neural network with computational neuroscience to simulate cochlear mechanics
A trio of researchers at Ghent University has combined a convolutional neural network with computational neuroscience to create a model that simulates human cochlear mechanics. In their paper published in Nature Machine Intelligence, Deepak Baby, Arthur Van Den Broucke and Sarah Verhulst describe how they built their model and the ways they believe it can be used. Over the past several decades, great strides have been made in speech and voice recognition technology. Customers are routinely serviced by phone-based agents, for example. Also, voice recognition and response systems on smartphones have become ubiquitous.
High-Performance Large-Scale Image Recognition Without Normalization
Brock, Andrew, De, Soham, Smith, Samuel L., Simonyan, Karen
Batch normalization is a key component of most image classification models, but it has many undesirable properties stemming from its dependence on the batch size and interactions between examples. Although recent work has succeeded in training deep ResNets without normalization layers, these models do not match the test accuracies of the best batch-normalized networks, and are often unstable for large learning rates or strong data augmentations. In this work, we develop an adaptive gradient clipping technique which overcomes these instabilities, and design a significantly improved class of Normalizer-Free ResNets. Our smaller models match the test accuracy of an EfficientNet-B7 on ImageNet while being up to 8.7x faster to train, and our largest models attain a new state-of-the-art top-1 accuracy of 86.5%. In addition, Normalizer-Free models attain significantly better performance than their batch-normalized counterparts when finetuning on ImageNet after large-scale pre-training on a dataset of 300 million labeled images, with our best models obtaining an accuracy of 89.2%. Our code is available at https://github.com/deepmind/ deepmind-research/tree/master/nfnets