Deep Learning
Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
Michaelis, Claudio, Mitzkus, Benjamin, Geirhos, Robert, Rusak, Evgenia, Bringmann, Oliver, Ecker, Alexander S., Bethge, Matthias, Brendel, Wieland
The ability to detect objects regardless of image distortions or weather conditions is crucial for real-world applications of deep learning like autonomous driving. We here provide an easy-to-use benchmark to assess how object detection models perform when image quality degrades. The three resulting benchmark datasets, termed Pascal-C, Coco-C and Cityscapes-C, contain a large variety of image corruptions. We show that a range of standard object detection models suffer a severe performance loss on corrupted images (down to 30-60% of the original performance). However, a simple data augmentation trick - stylizing the training images - leads to a substantial increase in robustness across corruption type, severity and dataset. We envision our comprehensive benchmark to track future progress towards building robust object detection models. Benchmark, code and data are available at: http://github.com/bethgelab/robust-detection-benchmark
Deep Learning to Address Candidate Generation and Cold Start Challenges in Recommender Systems: A Research Survey
Rama, Kiran, Kumar, Pradeep, Bhasker, Bharat
Among the machine learning applications to business, recommender systems would take one of the top places when it comes to success and adoption. They help the user in accelerating the process of search while helping businesses maximize sales. Post phenomenal success in computer vision and speech recognition, deep learning methods are beginning to get applied to recommender systems. Current survey papers on deep learning in recommender systems provide a historical overview and taxonomy of recommender systems based on type. Our paper addresses the gaps of providing a taxonomy of deep learning approaches to address recommender systems problems in the areas of cold start and candidate generation in recommender systems. We outline different challenges in recommender systems into those related to the recommendations themselves (include relevance, speed, accuracy and scalability), those related to the nature of the data (cold start problem, imbalance and sparsity) and candidate generation. We then provide a taxonomy of deep learning techniques to address these challenges. Deep learning techniques are mapped to the different challenges in recommender systems providing an overview of how deep learning techniques can be used to address them. We contribute a taxonomy of deep learning techniques to address the cold start and candidate generation problems in recommender systems. Cold Start is addressed through additional features (for audio, images, text) and by learning hidden user and item representations. Candidate generation has been addressed by separate networks, RNNs, autoencoders and hybrid methods. We also summarize the advantages and limitations of these techniques while outlining areas for future research.
OmniNet: A unified architecture for multi-modal multi-task learning
Pramanik, Subhojeet, Agrawal, Priyanka, Hussain, Aman
Transformer is a popularly used neural network architecture, especially for language understanding. We introduce an extended and unified architecture which can be used for tasks involving a variety of modalities like image, text, videos, etc. We propose a spatio-temporal cache mechanism that enables learning spatial dimension of the input in addition to the hidden states corresponding to the temporal input sequence. The proposed architecture further enables a single model to support tasks with multiple input modalities as well as asynchronous multi-task learning, thus we refer to it as OmniNet. For example, a single instance of OmniNet can concurrently learn to perform the tasks of part-of-speech tagging, image captioning, visual question answering and video activity recognition. We demonstrate that training these four tasks together results in about three times compressed model while retaining the performance in comparison to training them individually. We also show that using this neural network pre-trained on some modalities assists in learning an unseen task. This illustrates the generalization capacity of the self-attention mechanism on the spatio-temporal cache present in OmniNet.
Multi-Purposing Domain Adaptation Discriminators for Pseudo Labeling Confidence
Wilson, Garrett, Cook, Diane J.
Often domain adaptation is performed using a discriminator (domain One such adversarial domain-invariant feature learning method classifier) to learn domain-invariant feature representations is the domain-adversarial neural network (DANN) [14, 15], which so that a classifier trained on labeled source data will generalize is a typical baseline for other variants. This method consists of a feature well to unlabeled target data. A line of research stemming from extractor network followed by two additional networks: a task semi-supervised learning uses pseudo labeling to directly generate classifier and a domain classifier (Figure 1). The network is updated "pseudo labels" for the unlabeled target data and trains a classifier by two competing objectives: (1) the feature extractor followed by on the now-labeled target data, where the samples are selected the task classifier learns to correctly classify the labeled source data or weighted based on some measure of confidence. In this paper, while the domain classifier learns to correctly predict whether the we propose multi-purposing the discriminator to not only aid in features originated from source or target data, and (2) the feature producing domain-invariant representations but also to provide extractor learns to make the domain classifier predict the domain pseudo labeling confidence.
Design and Evaluation of Product Aesthetics: A Human-Machine Hybrid Approach
Burnap, Alex, Hauser, John R., Timoshenko, Artem
Aesthetics are critically important to market acceptance in many product categories. In the automotive industry in particular, an improved aesthetic design can boost sales by 30% or more. Firms invest heavily in designing and testing new product aesthetics. A single automotive "theme clinic" costs between \$100,000 and \$1,000,000, and hundreds are conducted annually. We use machine learning to augment human judgment when designing and testing new product aesthetics. The model combines a probabilistic variational autoencoder (VAE) and adversarial components from generative adversarial networks (GAN), along with modeling assumptions that address managerial requirements for firm adoption. We train our model with data from an automotive partner-7,000 images evaluated by targeted consumers and 180,000 high-quality unrated images. Our model predicts well the appeal of new aesthetic designs-38% improvement relative to a baseline and substantial improvement over both conventional machine learning models and pretrained deep learning models. New automotive designs are generated in a controllable manner for the design team to consider, which we also empirically verify are appealing to consumers. These results, combining human and machine inputs for practical managerial usage, suggest that machine learning offers significant opportunity to augment aesthetic design.
Deep Invertible Networks for EEG-based brain-signal decoding
Schirrmeister, Robin Tibor, Ball, Tonio
Deep-learning-based brain-signal decoding has recently achieved competitive accuracies compared with traditional feature-based decoding approaches. For example, they were used to decode movement-related EEG signals with accuracies at least as good as well-established movement-decoding approaches (Schirrmeister et al., 2017a) and been applied to error or event-related-based decoding (Lawhern et al., 2018, Vรถlker et al., 2018) as well as automatic diagnosis of pathologies (Schirrmeister et.
Deep Multi-View Learning via Task-Optimal CCA
Couture, Heather D., Kwitt, Roland, Marron, J. S., Troester, Melissa, Perou, Charles M., Niethammer, Marc
Canonical Correlation Analysis (CCA) is widely used for multimodal data analysis and, more recently, for discriminative tasks such as multi-view learning; however, it makes no use of class labels. Recent CCA methods have started to address this weakness but are limited in that they do not simultaneously optimize the CCA projection for discrimination and the CCA projection itself, or they are linear only. We address these deficiencies by simultaneously optimizing a CCA-based and a task objective in an end-to-end manner. Together, these two objectives learn a non-linear CCA projection to a shared latent space that is highly correlated and discriminative. Our method shows a significant improvement over previous state-of-the-art (including deep supervised approaches) for cross-view classification, regularization with a second view, and semi-supervised learning on real data.
An AI-Augmented Lesion Detection Framework For Liver Metastases With Model Interpretability
Hunt, Xin J., Abbey, Ralph, Tharrington, Ricky, Huiskens, Joost, Wesdorp, Nina
Colorectal cancer (CRC) is the third most common cancer and the second leading cause of cancer-related deaths worldwide. Most CRC deaths are the result of progression of metastases. The assessment of metastases is done using the RECIST criterion, which is time consuming and subjective, as clinicians need to manually measure anatomical tumor sizes. AI has many successes in image object detection, but often suffers because the models used are not interpretable, leading to issues in trust and implementation in the clinical setting. We propose a framework for an AI-augmented system in which an interactive AI system assists clinicians in the metastasis assessment. We include model interpretability to give explanations of the reasoning of the underlying models.
Mitigating Uncertainty in Document Classification
Zhang, Xuchao, Chen, Fanglan, Lu, Chang-Tien, Ramakrishnan, Naren
The uncertainty measurement of classifiers' predictions is especially important in applications such as medical diagnoses that need to ensure limited human resources can focus on the most uncertain predictions returned by machine learning models. However, few existing uncertainty models attempt to improve overall prediction accuracy where human resources are involved in the text classification task. In this paper, we propose a novel neural-network-based model that applies a new dropout-entropy method for uncertainty measurement. We also design a metric learning method on feature representations, which can boost the performance of dropout-based uncertainty methods with smaller prediction variance in accurate prediction trials. Extensive experiments on real-world data sets demonstrate that our method can achieve a considerable improvement in overall prediction accuracy compared to existing approaches. In particular, our model improved the accuracy from 0.78 to 0.92 when 30\% of the most uncertain predictions were handed over to human experts in "20NewsGroup" data.
AquaSight: Automatic Water Impurity Detection Utilizing Convolutional Neural Networks
Gupta, Ankit, Ruebush, Elliott
According to the United Nations World Water Assessment Programme, every day, 2 million tons of sewage and industrial and agricultural waste are discharged into the worlds water. In order to address this pervasive issue of increasing water pollution, while ensuring that the global population has an efficient, accurate, and low cost method to assess whether the water they drink is contaminated, we propose AquaSight, a novel mobile application that utilizes deep learning methods, specifically Convolutional Neural Networks, for automated water impurity detection. After comprehensive training with a dataset of 105 images representing varying magnitudes of contamination, the deep learning algorithm achieved a 96 percent accuracy and loss of 0.108. Furthermore, the machine learning model uses efficient analysis of the turbidity and transparency levels of water to estimate a particular sample of waters level of contamination. When deployed, the AquaSight system will provide an efficient way for individuals to secure an estimation of water quality, alerting local and national government to take action and potentially saving millions of lives worldwide.