Deep Learning
End to End Vehicle Lateral Control Using a Single Fisheye Camera
Toromanoff, Marin, Wirbel, Emilie, Wilhelm, Frรฉdรฉric, Vejarano, Camilo, Perrotton, Xavier, Moutarde, Fabien
Abstract-- Convolutional neural networks are commonly used to control the steering angle for autonomous cars. Most of the time, multiple long range cameras are used to generate lateral failure cases. In this paper we present a novel model to generate this data and label augmentation using only one short range fisheye camera. We present our simulator and how it can be used as a consistent metric for lateral end-to-end control evaluation. Experiments are conducted on a custom dataset corresponding to more than 10000 km and 200 hours of open road driving. Finally we evaluate this model on real world driving scenarios, open road and a custom test track with challenging obstacle avoidance and sharp turns. In our simulator based on real-world videos, the final model was capable of more than 99% autonomy on urban road. The ultimate goal for autonomous vehicles is to drive in any environment without any human input. To achieve this, autonomous cars have to analyze their environment using data coming from different sensors and control the car accordingly. In the most common approach, this task is cut into different modules then fed into a rule-based control algorithm which actually drives the car.
Controlling Over-generalization and its Effect on Adversarial Examples Generation and Detection
Abbasi, Mahdieh, Rajabi, Arezoo, Mozafari, Azadeh Sadat, Bobba, Rakesh B., Gagne, Christian
Convolutional Neural Networks (CNNs) allowed improving the state-of-the-art for many vision applications. However, naive CNNs suffer from two serious issues: vulnerability to adversarial examples and making incorrect but confident predictions for out-distribution samples. In this paper, we draw a connection between these two issues of CNNs through over-generalization. We reveal an augmented CNN (an extra output class added) as a simple yet effective end-to-end approach has the capacity for controlling over-generalization. We demonstrate training an augmented CNN on only a properly selected natural out-distribution dataset and interpolated samples empowers it to classify a wide range of unseen out-distribution samples as dustbin. Meanwhile, its misclassification rates on a broad spectrum of well-known black-box adversaries drop drastically as it classifies a portion of adversaries as dustbin class (rejection option) while correctly classifies some of the remaining. However, such an augmented CNN is never trained with any types of adversaries. Finally, generation of white-box adversarial attacks using augmented CNNs can be harder as the attack algorithms have to avoid dustbin regions for generating actual adversaries.
Class2Str: End to End Latent Hierarchy Learning
Saha, Soham, Varma, Girish, Jawahar, C. V.
Abstract--Deep neural networks for image classification typically consists of a convolutional feature extractor followed by a fully connected classifier network. The predicted and the ground truth labels are represented as one hot vectors. Such a representation assumes that all classes are equally dissimilar . However, classes have visual similarities and often form a hierarchy. Learning this latent hierarchy explicitly in the architecture could provide invaluable insights. We propose an alternate architecture to the classifier network called the Latent Hierarchy (LH) Classifier and an end to end learned Class2Str mapping which discovers a latent hierarchy of the classes. We show that for some of the best performing architectures on CIF AR and Imagenet datasets, the proposed replacement and training by LH classifier recovers the accuracy, with a fraction of the number of parameters in the classifier part. Compared to the previous work of HDCNN, which also learns a 2 level hierarchy, we are able to learn a hierarchy at an arbitrary number of levels as well as obtain an accuracy improvement on the Imagenet classification task over them. We also verify that many visually similar classes are grouped together, under the learnt hierarchy.
Deep Multimodal Image-Repurposing Detection
Sabir, Ekraam, AbdAlmageed, Wael, Wu, Yue, Natarajan, Prem
Nefarious actors on social media and other platforms often spread rumors and falsehoods through images whose metadata (e.g., captions) have been modified to provide visual substantiation of the rumor/falsehood. This type of modification is referred to as image repurposing, in which often an unmanipulated image is published along with incorrect or manipulated metadata to serve the actor's ulterior motives. We present the Multimodal Entity Image Repurposing (MEIR) dataset, a substantially challenging dataset over that which has been previously available to support research into image repurposing detection. The new dataset includes location, person, and organization manipulations on real-world data sourced from Flickr. We also present a novel, end-to-end, deep multimodal learning model for assessing the integrity of an image by combining information extracted from the image with related information from a knowledge base. The proposed method is compared against state-of-the-art techniques on existing datasets as well as MEIR, where it outperforms existing methods across the board, with AUC improvement up to 0.23.
LSTM-Based Goal Recognition in Latent Space
Amado, Leonardo, Aires, Joรฃo Paulo, Pereira, Ramon Fraga, Magnaguagno, Maurรญcio C., Granada, Roger, Meneguzzi, Felipe
Approaches to goal recognition have progressively relaxed the requirements about the amount of domain knowledge and available observations, yielding accurate and efficient algorithms capable of recognizing goals. However, to recognize goals in raw data, recent approaches require either human engineered domain knowledge, or samples of behavior that account for almost all actions being observed to infer possible goals. This is clearly too strong a requirement for real-world applications of goal recognition, and we develop an approach that leverages advances in recurrent neural networks to perform goal recognition as a classification task, using encoded plan traces for training. We empirically evaluate our approach against the state-of-the-art in goal recognition with image-based domains, and discuss under which conditions our approach is superior to previous ones.
Life-Long Disentangled Representation Learning with Cross-Domain Latent Homologies
Achille, Alessandro, Eccles, Tom, Matthey, Loic, Burgess, Christopher P., Watters, Nick, Lerchner, Alexander, Higgins, Irina
Intelligent behaviour in the real-world requires the ability to acquire new knowledge from an ongoing sequence of experiences while preserving and reusing past knowledge. We propose a novel algorithm for unsupervised representation learning from piece-wise stationary visual data: Variational Autoencoder with Shared Embeddings (VASE). Based on the Minimum Description Length principle, VASE automatically detects shifts in the data distribution and allocates spare representational capacity to new knowledge, while simultaneously protecting previously learnt representations from catastrophic forgetting. Our approach encourages the learnt representations to be disentangled, which imparts a number of desirable properties: VASE can deal sensibly with ambiguous inputs, it can enhance its own representations through imagination-based exploration, and most importantly, it exhibits semantically meaningful sharing of latents between different datasets. Compared to baselines with entangled representations, our approach is able to reason beyond surface-level statistics and perform semantically meaningful cross-domain inference.
Deep Dive Into Computer Vision With Neural Networks: Part 1 - DZone AI
Machine vision, or computer vision, is a popular research topic in artificial intelligence (AI) that has been around for many years. However, machine vision still remains as one of the biggest challenges in AI. In this article, we will explore the use of deep neural networks to address some of the fundamental challenges of computer vision. In particular, we will be looking at applications such as network compression, fine-grained image classification, captioning, texture synthesis, image search, and object tracking. Even though deep neural networks feature incredible performance, their demands for computing power and storage pose a significant challenge to their deployment in actual application.
Hear and Speak Your Natural -- NLP keras โ Data Driven Investor โ Medium
The Human's are evolved about 2.3 to 2.4 million years ago. Since the 18th century, Scientists thought the great apes to be closely related to human beings. In the 19th century, They speculated that closest living relatives of humans were either chimpanzees or gorillas. Do you know what made us different from our closest living relatives? Humans have a persistent process of thinking.
The 5 Steps To Build A Business's Deep Learning Workflow
Deep learning, or using massive amounts of data to build intelligent models, is a hot topic. Many companies, now seeing the benefits of AI materialize, have decided they need to get started on deep learning or risk getting left behind. How do you connect the big idea (businesses need to get started on deep learning) to the specific topic at hand (set up a workflow)? Here are the five main steps on setting up a deep learning workflow. Computers are great at optimizing models, but not so great at setting strategic goals.