Deep Learning
HDMapNet: An Online HD Map Construction and Evaluation Framework
Li, Qi, Wang, Yue, Wang, Yilun, Zhao, Hang
High-definition map (HD map) construction is a crucial problem for autonomous driving. This problem typically involves collecting high-quality point clouds, fusing multiple point clouds of the same scene, annotating map elements, and updating maps constantly. This pipeline, however, requires a vast amount of human efforts and resources which limits its scalability. Additionally, traditional HD maps are coupled with centimeter-level accurate localization which is unreliable in many scenarios. In this paper, we argue that online map learning, which dynamically constructs the HD maps based on local sensor observations, is a more scalable way to provide semantic and geometry priors to self-driving vehicles than traditional pre-annotated HD maps. Meanwhile, we introduce an online map learning method, titled HDMapNet. It encodes image features from surrounding cameras and/or point clouds from LiDAR, and predicts vectorized map elements in the bird's-eye view. We benchmark HDMapNet on the nuScenes dataset and show that in all settings, it performs better than baseline methods. Of note, our fusion-based HDMapNet outperforms existing methods by more than 50% in all metrics. To accelerate future research, we develop customized metrics to evaluate map learning performance, including both semantic-level and instance-level ones. By introducing this method and metrics, we invite the community to study this novel map learning problem. We will release our code and evaluation kit to facilitate future development.
Real-Time Super-Resolution System of 4K-Video Based on Deep Learning
Cao, Yanpeng, Wang, Chengcheng, Song, Changjun, Tang, Yongming, Li, He
Video super-resolution (VSR) technology excels in reconstructing low-quality video, avoiding unpleasant blur effect caused by interpolation-based algorithms. However, vast computation complexity and memory occupation hampers the edge of deplorability and the runtime inference in real-life applications, especially for large-scale VSR task. This paper explores the possibility of real-time VSR system and designs an efficient and generic VSR network, termed EGVSR. The proposed EGVSR is based on spatio-temporal adversarial learning for temporal coherence. In order to pursue faster VSR processing ability up to 4K resolution, this paper tries to choose lightweight network structure and efficient upsampling method to reduce the computation required by EGVSR network under the guarantee of high visual quality. Besides, we implement the batch normalization computation fusion, convolutional acceleration algorithm and other neural network acceleration techniques on the actual hardware platform to optimize the inference process of EGVSR network. Finally, our EGVSR achieves the real-time processing capacity of 4K@29.61FPS. Compared with TecoGAN, the most advanced VSR network at present, we achieve 85.04% reduction of computation density and 7.92x performance speedups. In terms of visual quality, the proposed EGVSR tops the list of most metrics (such as LPIPS, tOF, tLP, etc.) on the public test dataset Vid4 and surpasses other state-of-the-art methods in overall performance score. The source code of this project can be found on https://github.com/Thmen/EGVSR.
Deep Metric Learning Model for Imbalanced Fault Diagnosis
Intelligent diagnosis method based on data-driven and deep learning is an attractive and meaningful field in recent years. However, in practical application scenarios, the imbalance of time-series fault is an urgent problem to be solved. This paper proposes a novel deep metric learning model, where imbalanced fault data and a quadruplet data pair design manner are considered. Based on such data pair, a quadruplet loss function which takes into account the inter-class distance and the intra-class data distribution are proposed. This quadruplet loss pays special attention to imbalanced sample pair. The reasonable combination of quadruplet loss and softmax loss function can reduce the impact of imbalance. Experiment results on two open-source datasets show that the proposed method can effectively and robustly improve the performance of imbalanced fault diagnosis.
A Survey on Data Augmentation for Text Classification
Bayer, Markus, Kaufhold, Marc-André, Reuter, Christian
Data augmentation, the artificial creation of training data for machine learning by transformations, is a widely studied research field across machine learning disciplines. While it is useful for increasing the generalization capabilities of a model, it can also address many other challenges and problems, from overcoming a limited amount of training data over regularizing the objective to limiting the amount data used to protect privacy. Based on a precise description of the goals and applications of data augmentation (C1) and a taxonomy for existing works (C2), this survey is concerned with data augmentation methods for textual classification and aims to achieve a concise and comprehensive overview for researchers and practitioners (C3). Derived from the taxonomy, we divided more than 100 methods into 12 different groupings and provide state-of-the-art references expounding which methods are highly promising (C4). Finally, research perspectives that may constitute a building block for future work are given (C5).
The Causal-Neural Connection: Expressiveness, Learnability, and Inference
Xia, Kevin, Lee, Kai-Zhan, Bengio, Yoshua, Bareinboim, Elias
One of the central elements of any causal inference is an object called structural causal model (SCM), which represents a collection of mechanisms and exogenous sources of random variation of the system under investigation (Pearl, 2000). An important property of many kinds of neural networks is universal approximability: the ability to approximate any function to arbitrary precision. Given this property, one may be tempted to surmise that a collection of neural nets is capable of learning any SCM by training on data generated by that SCM. In this paper, we show this is not the case by disentangling the notions of expressivity and learnability. Specifically, we show that the causal hierarchy theorem (Thm. 1, Bareinboim et al., 2020), which describes the limits of what can be learned from data, still holds for neural models. For instance, an arbitrarily complex and expressive neural net is unable to predict the effects of interventions given observational data alone. Given this result, we introduce a special type of SCM called a neural causal model (NCM), and formalize a new type of inductive bias to encode structural constraints necessary for performing causal inferences. Building on this new class of models, we focus on solving two canonical tasks found in the literature known as causal identification and estimation. Leveraging the neural toolbox, we develop an algorithm that is both sufficient and necessary to determine whether a causal effect can be learned from data (i.e., causal identifiability); it then estimates the effect whenever identifiability holds (causal estimation). Simulations corroborate the proposed approach.
Self-Contrastive Learning
Bae, Sangmin, Kim, Sungnyun, Ko, Jongwoo, Lee, Gihun, Noh, Seungjong, Yun, Se-Young
This paper proposes a novel contrastive learning framework, coined as Self-Contrastive (SelfCon) Learning, that self-contrasts within multiple outputs from the different levels of a network. We confirmed that SelfCon loss guarantees the lower bound of mutual information (MI) between the intermediate and last representations. Besides, we empirically showed, via various MI estimators, that SelfCon loss highly correlates to the increase of MI and better classification performance. In our experiments, SelfCon surpasses supervised contrastive (SupCon) learning without the need for a multi-viewed batch and with the cheaper computational cost. Especially on ResNet-18, we achieved top-1 classification accuracy of 76.45% for the CIFAR-100 dataset, which is 2.87% and 4.36% higher than SupCon and cross-entropy loss, respectively. We found that mitigating both vanishing gradient and overfitting issue makes our method outperform the counterparts.
This AI Asks Questions, Finds Answers And Suggests Actions, All At Scale
The Covid-19 pandemic has revealed to us a very inconvenient truth: Regardless of all the technological advances of the last century and a half, our lives can be stopped and destroyed suddenly and unexpectedly by an invisible plague. The excitement over the most recent technological breakthrough--artificial intelligence or deep learning--has met the grim reality of our inability to adequately prepare for, manage, and overcome society's most important and consequential challenges. While Google's DeepMind set itself up to "solve intelligence," i.e., to advance the state of AI and then use the hoped-for "superior intelligence" to solve humanity's challenges, Sagie Davidovich and Ron Karidi co-founded SparkBeyond 7 years ago "to harness the world's collective intelligence in order to solve the world's toughest challenges." They aimed to use existing AI, algorithms, and knowledge to advance humanity's problem-solving capabilities. Most important, SparkBeyond wanted to go beyond the typical use of AI which is basically an extension and an upgrade of what has been called "predictive analytics" before the 2010s.
5 AI startups out to change the world
Advances in deep learning and neural networks have delivered huge breakthroughs in natural language processing and computer vision, and they have the potential to solve big problems in manufacturing, retail, supply chain, agriculture, and countless other business domains. Naturally, technology startups are behind some of the most important innovations. In recent articles, we looked at startups revolutionizing natural language processing and startups leading the way in MLops. Here we'll take a look at "applied AI" startups. These are companies that are applying different techniques--whether it be processing images, text, audio, video, categorical or tabular data, or combinations of the above--to address various industry challenges, from fulfilling the promise of self-driving cars to pushing the boundaries of agricultural production.
Video AI in the Cloud: 6 Platforms and APIs
Artificial intelligence (AI) is increasingly being used to manage video content. Deep learning-based computer vision techniques can help recognize concepts and faces in video streams, categorize videos, automatically add captions, and enhance videos and images using techniques like super-resolution. Developing video AI from scratch is a huge investment. Today you can leverage video AI capabilities out of the box, using the powerful video APIs offered by a number of cloud platforms. In this article I'll describe a few of the world's most advanced video AI platforms: This service makes it easy to add image and video analytics to your applications using mature, highly scalable deep learning technology.
The Threat of Artificial Intelligence
The technologies referred to as "artificial intelligence" or "AI" are more momentous than most people realize. Their impact will be at least equal to, and may well exceed, that of electricity, the computer, and the internet. What's more, their impact will be massive and rapid, faster than what the internet has wrought in the past thirty years. Much of it will be wondrous, giving sight to the blind and enabling self-driving vehicles, for example, but AI-engendered technology may also devastate job rolls, enable an all- encompassing surveillance state, and provoke social upheavals yet unforeseen. The time we have to understand this fast-moving technology and establish principles for its governance is very short. The term "AI" was coined by a computer scientist in 1956.