Goto

Collaborating Authors

 Statistical Learning


Cross Spline Net and a Unified World

arXiv.org Machine Learning

In today's machine learning world for tabular data, XGBoost and fully connected neural network (FCNN) are two most popular methods due to their good model performance and convenience to use. However, they are highly complicated, hard to interpret, and can be overfitted. In this paper, we propose a new modeling framework called cross spline net (CSN) that is based on a combination of spline transformation and cross-network (Wang et al. 2017, 2021). We will show CSN is as performant and convenient to use, and is less complicated, more interpretable and robust. Moreover, the CSN framework is flexible, as the spline layer can be configured differently to yield different models. With different choices of the spline layer, we can reproduce or approximate a set of non-neural network models, including linear and spline-based statistical models, tree, rule-fit, tree-ensembles (gradient boosting trees, random forest), oblique tree/forests, multi-variate adaptive regression spline (MARS), SVM with polynomial kernel, etc. Therefore, CSN provides a unified modeling framework that puts the above set of non-neural network models under the same neural network framework. By using scalable and powerful gradient descent algorithms available in neural network libraries, CSN avoids some pitfalls (such as being ad-hoc, greedy or non-scalable) in the case-specific optimization methods used in the above non-neural network models. We will use a special type of CSN, TreeNet, to illustrate our point. We will compare TreeNet with XGBoost and FCNN to show the benefits of TreeNet. We believe CSN will provide a flexible and convenient framework for practitioners to build performant, robust and more interpretable models.


Provable Tempered Overfitting of Minimal Nets and Typical Nets

arXiv.org Machine Learning

We study the overfitting behavior of fully connected deep Neural Networks (NNs) with binary weights fitted to perfectly classify a noisy training set. We consider interpolation using both the smallest NN (having the minimal number of weights) and a random interpolating NN. For both learning rules, we prove overfitting is tempered. Our analysis rests on a new bound on the size of a threshold circuit consistent with a partial function. To the best of our knowledge, ours are the first theoretical results on benign or tempered overfitting that: (1) apply to deep NNs, and (2) do not require a very high or very low input dimension.


Learn 2 Rage: Experiencing The Emotional Roller Coaster That Is Reinforcement Learning

arXiv.org Artificial Intelligence

This work presents the experiments and solution outline for our teams winning submission in the Learn To Race Autonomous Racing Virtual Challenge 2022 hosted by AIcrowd. The objective of the Learn-to-Race competition is to push the boundary of autonomous technology, with a focus on achieving the safety benefits of autonomous driving. In the description the competition is framed as a reinforcement learning (RL) challenge. We focused our initial efforts on implementation of Soft Actor Critic (SAC) variants. Our goal was to learn non-trivial control of the race car exclusively from visual and geometric features, directly mapping pixels to control actions. We made suitable modifications to the default reward policy aiming to promote smooth steering and acceleration control. The framework for the competition provided real time simulation, meaning a single episode (learning experience) is measured in minutes. Instead of pursuing parallelisation of episodes we opted to explore a more traditional approach in which the visual perception was processed (via learned operators) and fed into rule-based controllers. Such a system, while not as academically "attractive" as a pixels-to-actions approach, results in a system that requires less training, is more explainable, generalises better and is easily tuned and ultimately out-performed all other agents in the competition by a large margin.


Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens

arXiv.org Artificial Intelligence

Current progress in artificial intelligence is centered around so-called large language models that consist of neural networks processing long sequences of high-dimensional vectors called tokens. Statistical physics provides powerful tools to study the functioning of learning with neural networks and has played a recognized role in the development of modern machine learning. The statistical physics approach relies on simplified and analytically tractable models of data. However, simple tractable models for long sequences of high-dimensional tokens are largely underexplored. Inspired by the crucial role models such as the single-layer teacher-student perceptron (aka generalized linear regression) played in the theory of fully connected neural networks, in this paper, we introduce and study the bilinear sequence regression (BSR) as one of the most basic models for sequences of tokens. We note that modern architectures naturally subsume the BSR model due to the skip connections. Building on recent methodological progress, we compute the Bayes-optimal generalization error for the model in the limit of long sequences of high-dimensional tokens, and provide a message-passing algorithm that matches this performance. We quantify the improvement that optimal learning brings with respect to vectorizing the sequence of tokens and learning via simple linear regression. We also unveil surprising properties of the gradient descent algorithms in the BSR model.


Maximum a Posteriori Inference for Factor Graphs via Benders' Decomposition

arXiv.org Machine Learning

Many Bayesian statistical inference problems come down to computing a maximum a-posteriori (MAP) assignment of latent variables. Yet, standard methods for estimating the MAP assignment do not have a finite time guarantee that the algorithm has converged to a fixed point. Previous research has found that MAP inference can be represented in dual form as a linear programming problem with a non-polynomial number of constraints. A Lagrangian relaxation of the dual yields a statistical inference algorithm as a linear programming problem. However, the decision as to which constraints to remove in the relaxation is often heuristic. We present a method for maximum a-posteriori inference in general Bayesian factor models that sequentially adds constraints to the fully relaxed dual problem using Benders' decomposition. Our method enables the incorporation of expressive integer and logical constraints in clustering problems such as must-link, cannot-link, and a minimum number of whole samples allocated to each cluster. Using this approach, we derive MAP estimation algorithms for the Bayesian Gaussian mixture model and latent Dirichlet allocation. Empirical results show that our method produces a higher optimal posterior value compared to Gibbs sampling and variational Bayes methods for standard data sets and provides certificate of convergence.


Retrieving snow depth distribution by downscaling ERA5 Reanalysis with ICESat-2 laser altimetry

arXiv.org Artificial Intelligence

Estimating the variability of seasonal snow cover, in particular snow depth in remote areas, poses significant challenges due to limited spatial and temporal data availability. This study uses snow depth measurements from the ICESat-2 satellite laser altimeter, which are sparse in both space and time, and incorporates them with climate reanalysis data into a downscaling-calibration scheme to produce monthly gridded snow depth maps at microscale (10 m). Snow surface elevation measurements from ICESat-2 along profiles are compared to a digital elevation model to determine snow depth at each point. To efficiently turn sparse measurements into snow depth maps, a regression model is fitted to establish a relationship between the retrieved snow depth and the corresponding ERA5 Land snow depth. This relationship, referred to as subgrid variability, is then applied to downscale the monthly ERA5 Land snow depth data. The method can provide timeseries of monthly snow depth maps for the entire ERA5 time range (since 1950). The validation of downscaled snow depth data was performed at an intermediate scale (100 m x 500 m) using datasets from airborne laser scanning (ALS) in the Hardangervidda region of southern Norway. Results show that snow depth prediction achieved R2 values ranging from 0.74 to 0.88 (post-calibration). The method relies on globally available data and is applicable to other snow regions above the treeline. Though requiring area-specific calibration, our approach has the potential to provide snow depth maps in areas where no such data exist and can be used to extrapolate existing snow surveys in time and over larger areas. With this, it can offer valuable input data for hydrological, ecological or permafrost modeling tasks.


Population stratification for prediction of mortality in post-AKI patients

arXiv.org Artificial Intelligence

AKI is associated with increases in (1) post-discharge mortality risk, (2) length of hospital stay and (3) healthcare expenditures [19], as well as short term unplanned re-admissions and mid term progressive chronic conditions. Around 33% of AKI patients require unplanned re-admissions within 90 days after discharge and around 15% develop progressive chronic kidney disease over the first year after discharge [14, 16]. AKI is multi-factorial, and accurate follow up planning is challenging. Machine learning has been viewed as promising to build tools to support decision making in clinical follow-up planning. Broadly speaking, recent initiatives can be structured along two alternatives: 1. Tools grounded on prior medical expert knowledge, which is used to stratify patients according to meaningful attributes, in such way that specialised plans can be devised for each group of patients [5, 13, 15, 19, 23]. 2. Tools grounded on machine learning techniques, which take control of the planning process and build accurate decision procedures which, however, demand extreme care in selection of new patients, to ensure compliance with population definitions that are used during preparation of decision procedures [1, 2, 3, 7, 21, 22]. Compliance with ethical standards demands that such tools are fair, transparent, and optimised for the benefit of patients. Technical requirements to ensure ethical compliance must include algorithmic transparency to support fairness and transparency in decision making and optimised, goal-oriented patient stratification to ensure human-centred optimised performance. The research initiative presented in this article focused on the development of a tool to support clinical follow up planning for post-AKI patients after hospital discharge, with particular attention to ethical compliance based on technical requirements.


Robust Variable Selection for High-dimensional Regression with Missing Data and Measurement Errors

arXiv.org Machine Learning

The linear relationship between response variables and covariates has been the topic of interest.In the classical squared loss function,it is usually assumed that the data obey a normal distribution.However,the data discussed in this paper contain a large number of missing data and measurement errors,such that the datausually do not conform to any of the common forms of data distribution.We propose a method based on an exponential squared loss function with tuning parameter.For data with different distributions,a better result of linear regression can be achieved by changing the value of the tuning parameter h.Therefore,forany kind of data distribution,going with an exponential squared loss function with moderating variables will be highly robust.For any data distribution,the loss function is strongly robust for h (0,+x).In previous studies,when using the traditional squared loss function,the data distribution requirements are very high,resulting in the traditional exponential squared loss function being very sensitive to anomalies.This reduces the estimation efficiency of the model,and this drawback becomes more obvious in data containing missing data with measurement errors.In contrast,the use of exponential squared loss functions can improve the estimation efficiency of the model by varying thetuning parameter h in a way that adapts to more distributed forms of data sets and produces more reliable estimates. In the traditional squared loss function,the values of the covariates are always defaulted to be free ofmissingdata and measurement errors.Even if missing data and measurement errors exist,they are assumed to be absent or these data are removed.However,this assumption is often broken in studies in disciplines such as health and epidemiology.As an illustration,Zhang and Zhou(1)looked at a collection of breast cancer patients to identify the gene expression that was associated with long-term disease-free survival.The datacollection consists of 24481 gene probes collected from 78 breast cancer patients.In particular,using the log-value of the ratio (log1o(Ratio)),which could be denoted as Y,it is possible to forecast the disease-free survival.In truth,gene sensors will inevitably lead to measurement errors.In this breast cancer data set,the(log1o(Ratio))numbers have missing data. When there are a large numberof missing data and measurement errors in a dataset,if we ignore the missing data and measurement errors and use the traditional square loss function for estimation,the estimation accuracy of the model will be greatly affected due to the chaotic data distribution,resulting in significant estimation bias.In the above dataset, We discover that employing the traditional squared loss function,which handles data with measurement errors and Robust Variable Selection for High-dimensional Regression with Missing Data and Measurement Errors


Rethinking Positive Pairs in Contrastive Learning

arXiv.org Artificial Intelligence

Contrastive learning, a prominent approach to representation learning, traditionally assumes positive pairs are closely related samples (the same image or class) and negative pairs are distinct samples. We challenge this assumption by proposing to learn from arbitrary pairs, allowing any pair of samples to be positive within our framework.The primary challenge of the proposed approach lies in applying contrastive learning to disparate pairs which are semantically distant. Motivated by the discovery that SimCLR can separate given arbitrary pairs (e.g., garter snake and table lamp) in a subspace, we propose a feature filter in the condition of class pairs that creates the requisite subspaces by gate vectors selectively activating or deactivating dimensions. This filter can be optimized through gradient descent within a conventional contrastive learning mechanism. We present Hydra, a universal contrastive learning framework for visual representations that extends conventional contrastive learning to accommodate arbitrary pairs. Our approach is validated using IN1K, where 1K diverse classes compose 500,500 pairs, most of them being distinct. Surprisingly, Hydra achieves superior performance in this challenging setting. Additional benefits include the prevention of dimensional collapse and the discovery of class relationships. Our work highlights the value of learning common features of arbitrary pairs and potentially broadens the applicability of contrastive learning techniques on the sample pairs with weak relationships.


Neuropsychology and Explainability of AI: A Distributional Approach to the Relationship Between Activation Similarity of Neural Categories in Synthetic Cognition

arXiv.org Artificial Intelligence

Within an explainability framework, the neuropsychology of artificial intelligence focuses on studying synthetic neural cognitive mechanisms, considering them as new subjects of cognitive psychology research. The goal is to make artificial neural networks used in language models understandable by adapting concepts from human cognitive psychology to the interpretation of artificial neural cognition. In this context, the notion of categorization is particularly relevant because it plays a key role as a process of segmentation and reconstruction of informational data by the neural vectors of synthetic cognition. Thus, in this study, the aim is to use the concept of categorization, as understood in human cognitive psychology (particularly in its relation to the notion of similarity), to apply it to the analysis of neural behavior and to infer certain synthetic cognitive processes underlying the observed behaviors.