Goto

Collaborating Authors

 Statistical Learning


HDMapNet: An Online HD Map Construction and Evaluation Framework

arXiv.org Artificial Intelligence

High-definition map (HD map) construction is a crucial problem for autonomous driving. This problem typically involves collecting high-quality point clouds, fusing multiple point clouds of the same scene, annotating map elements, and updating maps constantly. This pipeline, however, requires a vast amount of human efforts and resources which limits its scalability. Additionally, traditional HD maps are coupled with centimeter-level accurate localization which is unreliable in many scenarios. In this paper, we argue that online map learning, which dynamically constructs the HD maps based on local sensor observations, is a more scalable way to provide semantic and geometry priors to self-driving vehicles than traditional pre-annotated HD maps. Meanwhile, we introduce an online map learning method, titled HDMapNet. It encodes image features from surrounding cameras and/or point clouds from LiDAR, and predicts vectorized map elements in the bird's-eye view. We benchmark HDMapNet on the nuScenes dataset and show that in all settings, it performs better than baseline methods. Of note, our fusion-based HDMapNet outperforms existing methods by more than 50% in all metrics. To accelerate future research, we develop customized metrics to evaluate map learning performance, including both semantic-level and instance-level ones. By introducing this method and metrics, we invite the community to study this novel map learning problem. We will release our code and evaluation kit to facilitate future development.


Kernel Continual Learning

arXiv.org Artificial Intelligence

This paper introduces kernel continual learning, a simple but effective variant of continual learning that leverages the non-parametric nature of kernel methods to tackle catastrophic forgetting. We deploy an episodic memory unit that stores a subset of samples for each task to learn task-specific classifiers based on kernel ridge regression. This does not require memory replay and systematically avoids task interference in the classifiers. We further introduce variational random features to learn a data-driven kernel for each task. To do so, we formulate kernel continual learning as a variational inference problem, where a random Fourier basis is incorporated as the latent variable. The variational posterior distribution over the random Fourier basis is inferred from the coreset of each task. In this way, we are able to generate more informative kernels specific to each task, and, more importantly, the coreset size can be reduced to achieve more compact memory, resulting in more efficient continual learning based on episodic memory. Extensive evaluation on four benchmarks demonstrates the effectiveness and promise of kernels for continual learning.


Deep Metric Learning Model for Imbalanced Fault Diagnosis

arXiv.org Artificial Intelligence

Intelligent diagnosis method based on data-driven and deep learning is an attractive and meaningful field in recent years. However, in practical application scenarios, the imbalance of time-series fault is an urgent problem to be solved. This paper proposes a novel deep metric learning model, where imbalanced fault data and a quadruplet data pair design manner are considered. Based on such data pair, a quadruplet loss function which takes into account the inter-class distance and the intra-class data distribution are proposed. This quadruplet loss pays special attention to imbalanced sample pair. The reasonable combination of quadruplet loss and softmax loss function can reduce the impact of imbalance. Experiment results on two open-source datasets show that the proposed method can effectively and robustly improve the performance of imbalanced fault diagnosis.


A Survey on Data Augmentation for Text Classification

arXiv.org Artificial Intelligence

Data augmentation, the artificial creation of training data for machine learning by transformations, is a widely studied research field across machine learning disciplines. While it is useful for increasing the generalization capabilities of a model, it can also address many other challenges and problems, from overcoming a limited amount of training data over regularizing the objective to limiting the amount data used to protect privacy. Based on a precise description of the goals and applications of data augmentation (C1) and a taxonomy for existing works (C2), this survey is concerned with data augmentation methods for textual classification and aims to achieve a concise and comprehensive overview for researchers and practitioners (C3). Derived from the taxonomy, we divided more than 100 methods into 12 different groupings and provide state-of-the-art references expounding which methods are highly promising (C4). Finally, research perspectives that may constitute a building block for future work are given (C5).


The Causal-Neural Connection: Expressiveness, Learnability, and Inference

arXiv.org Artificial Intelligence

One of the central elements of any causal inference is an object called structural causal model (SCM), which represents a collection of mechanisms and exogenous sources of random variation of the system under investigation (Pearl, 2000). An important property of many kinds of neural networks is universal approximability: the ability to approximate any function to arbitrary precision. Given this property, one may be tempted to surmise that a collection of neural nets is capable of learning any SCM by training on data generated by that SCM. In this paper, we show this is not the case by disentangling the notions of expressivity and learnability. Specifically, we show that the causal hierarchy theorem (Thm. 1, Bareinboim et al., 2020), which describes the limits of what can be learned from data, still holds for neural models. For instance, an arbitrarily complex and expressive neural net is unable to predict the effects of interventions given observational data alone. Given this result, we introduce a special type of SCM called a neural causal model (NCM), and formalize a new type of inductive bias to encode structural constraints necessary for performing causal inferences. Building on this new class of models, we focus on solving two canonical tasks found in the literature known as causal identification and estimation. Leveraging the neural toolbox, we develop an algorithm that is both sufficient and necessary to determine whether a causal effect can be learned from data (i.e., causal identifiability); it then estimates the effect whenever identifiability holds (causal estimation). Simulations corroborate the proposed approach.


Self-Contrastive Learning

arXiv.org Artificial Intelligence

This paper proposes a novel contrastive learning framework, coined as Self-Contrastive (SelfCon) Learning, that self-contrasts within multiple outputs from the different levels of a network. We confirmed that SelfCon loss guarantees the lower bound of mutual information (MI) between the intermediate and last representations. Besides, we empirically showed, via various MI estimators, that SelfCon loss highly correlates to the increase of MI and better classification performance. In our experiments, SelfCon surpasses supervised contrastive (SupCon) learning without the need for a multi-viewed batch and with the cheaper computational cost. Especially on ResNet-18, we achieved top-1 classification accuracy of 76.45% for the CIFAR-100 dataset, which is 2.87% and 4.36% higher than SupCon and cross-entropy loss, respectively. We found that mitigating both vanishing gradient and overfitting issue makes our method outperform the counterparts.


How AI Can Solve Trader's Never-ending Search For Edge In Markets

#artificialintelligence

Artificial Intelligence can help you make decisions more objectively as it is based on numerous data points. AI is evolving at a very fast pace and is being used in every field. AI mimics the human brain, we have been using it ourselves from ages, when we track the prices of stock and try to remember them we are creating memory which creates patterns. AI can do the same task efficiently and can include more data point. AI systems work by crunching through data, trying to find patterns and correlations, and teaching themselves how to approximate future outcomes by formulating algorithms that use the patterns and correlations found in the data.


Colour Quantization Using K-Means Clustering and OpenCV

#artificialintelligence

Have you ever wondered how we can implement a machine learning algorithm on the pixel intensity value with a common K-means clustering algorithm? In this method, we would generate a compressed variant of our picture with more scattered colours. The image will be processed in a lower intensity resolution, whereas the fraction of pixels will prevail. This procedure is very interesting, so I expect that you will like it. This article can appear as a particularly impressive and unexpected one, so here is the link to the article, please have a read and hope you like it.


Gaussian process interpolation: the choice of the family of models is more important than that of the selection criterion

arXiv.org Machine Learning

Regression and interpolation with Gaussian processes, or kriging, is a popular statistical tool for non-parametric function estimation, originating from geostatistics and time series analysis, and later adopted in many other areas such as machine learning and the design and analysis of computer experiments (see, e.g., Stein, 1999; Santner et al., 2003; Rasmussen and Williams, 2006, and references therein). It is widely used for constructing fast approximations of time-consuming computer models, with applications to calibration and validation (Kennedy and O'Hagan, 2001; Bayarri et al., 2007), engineering design (Jones et al., 1998; Forrester et al., 2008), Bayesian inference (Calderhead et al., 2009; Wilkinson, 2014), and the optimization of machine learning algorithms (Bergstra et al., 2011)--to name but a few. A Gaussian process (GP) prior is characterized by its mean and covariance functions. They are usually chosen within parametric families (for instance, constant or linear mean functions, and Matérn covariance functions), which transfers the problem of choosing the mean and covariance functions to that of selecting parameters. The selection is most often carried out by optimization of a criterion that measures the goodness of fit of the predictive distributions, and a variety of such criteria--the likelihood function, the leave-one-out (LOO) squared-predictionerror criterion (hereafter denoted by LOO-SPE), and others--is available from the literature.


A Graph Data Augmentation Strategy with Entropy Preserving

arXiv.org Artificial Intelligence

All of these domains and many more can be readily modeled as graphs, which contain information about connection between individual units. For instance, citation graph describes interactions among science research papers which are represented as nodes with labels to indicate category, and the citation links between papers are mapped into edges. Information from single node or local dense nodes propagates along edges, and this makes graphs be useful structured knowledge repositories for machine learning tasks like link prediction and node classification. Graph Convolutional Networks (GCNs) [4, 5, 6, 7, 8] draw support from convolutional operation on graph to aggregate neighbor nodes information from low-to high-order hierarchical structures to get central node representation. In the course of time, GCNs and subsequent variants have emerged as powerful approaches for a variety of tasks like semi-supervised node classification [4, 9], which is also the main focus of this paper. In order to enable GCNs with more expressivity to wider neighbors, one may stack more layers to the network. But unfortunately, deeper layer network model fails to achieve the expectation partly due to the phenomenon of over-smoothing [10], which is an inherent issue of graph convolutional calculation mechanism. It has been proven that graph convolution operation is a type of Laplacian smoothing, thus representations of nodes in same region converge to same values and tend to be indistinguishable across different classes in embedding space as model goes deeper [11]. An easy but effective way to tackle with over-smoothing is to generate perturbed graph data for training.