Statistical Learning
Wasserstein-Aitchison GAN for angular measures of multivariate extremes
Lhaut, Stéphane, Rootzén, Holger, Segers, Johan
Economically responsible mitigation of multivariate extreme risks -- extreme rainfall in a large area, huge variations of many stock prices, widespread breakdowns in transportation systems -- requires estimates of the probabilities that such risks will materialize in the future. This paper develops a new method, Wasserstein--Aitchison Generative Adversarial Networks (WA-GAN), which provides simulated values of future $d$-dimensional multivariate extreme events and which hence can be used to give estimates of such probabilities. The main hypothesis is that, after transforming the observations to the unit-Pareto scale, their distribution is regularly varying in the sense that the distributions of their radial and angular components (with respect to the $L_1$-norm) converge and become asymptotically independent as the radius gets large. The method is a combination of standard extreme value analysis modeling of the tails of the marginal distributions with nonparametric GAN modeling of the angular distribution. For the latter, the angular values are transformed to Aitchison coordinates in a full $(d-1)$-dimensional linear space, and a Wasserstein GAN is trained on these coordinates and used to generate new values. A reverse transformation is then applied to these values and gives simulated values on the original data scale. The method shows good performance compared to other existing methods in the literature, both in terms of capturing the dependence structure of the extremes in the data, as well as in generating accurate new extremes of the data distribution. The comparison is performed on simulated multivariate extremes from a logistic model in dimensions up to 50 and on a 30-dimensional financial data set.
Conditional independence testing with a single realization of a multivariate nonstationary nonlinear time series
Wieck-Sosa, Michael, Haddad, Michel F. C., Ramdas, Aaditya
That is, testing whether two random vectors X and Y are independent given a third random vector Z . For example, there are conditional independence tests based on conditional densities [SW08], characteristic functions [SW07], empirical likelihood ratios [SW14], discretization [Mar05; Hua10], permutation [Dor+14; Sen+17], kernels [Fuk+07; Zha+11; SP11], copulas [BRT12], and conditional mutual information [Run18b]. Also, there are many conditional independence tests based on regressing X on Z and Y on Z followed by testing for independence between the residuals [Pat+09; Pet+14; Ram14; FFX20; ZZG17; Zha+19]. Unfortunately, conditional independence tests oftentimes struggle to control the Type-I error in finite samples, as shown by Shah and Peters [SP20]. In fact, Shah and Peters [SP20] prove that conditional independence testing is fundamentally impossible without making further assumptions. This issue has sparked significant interest in conditional independence testing over the last several years. We begin by providing an overview of recent advances in conditional independence testing. Afterwards, we discuss how our work addresses limitations in the existing literature. Finally, we motivate our work by reviewing key applications of conditional independence tests for time series in areas such as variable selection and causal discovery.
Assessing Racial Disparities in Healthcare Expenditures Using Causal Path-Specific Effects
Ou, Xiaxian, He, Xinwei, Benkeser, David, Nabi, Razieh
Racial disparities in healthcare expenditures are well-documented, yet the underlying drivers remain complex and require further investigation. This study employs causal and counterfactual path-specific effects to quantify how various factors, including socioeconomic status, insurance access, health behaviors, and health status, mediate these disparities. Using data from the Medical Expenditures Panel Survey, we estimate how expenditures would differ under counterfactual scenarios in which the values of specific mediators were aligned across racial groups along selected causal pathways. A key challenge in this analysis is ensuring robustness against model misspecification while addressing the zero-inflation and right-skewness of healthcare expenditures. For reliable inference, we derive asymptotically linear estimators by integrating influence function-based techniques with flexible machine learning methods, including super learners and a two-part model tailored to the zero-inflated, right-skewed nature of healthcare expenditures.
Cert-SSB: Toward Certified Sample-Specific Backdoor Defense
Qiao, Ting, Wang, Yingjia, Liu, Xing, Wu, Sixing, Li, Jianbing, Li, Yiming
Deep neural networks (DNNs) are vulnerable to backdoor attacks, where an attacker manipulates a small portion of the training data to implant hidden backdoors into the model. The compromised model behaves normally on clean samples but misclassifies backdoored samples into the attacker-specified target class, posing a significant threat to real-world DNN applications. Currently, several empirical defense methods have been proposed to mitigate backdoor attacks, but they are often bypassed by more advanced backdoor techniques. In contrast, certified defenses based on randomized smoothing have shown promise by adding random noise to training and testing samples to counteract backdoor attacks. In this paper, we reveal that existing randomized smoothing defenses implicitly assume that all samples are equidistant from the decision boundary. However, it may not hold in practice, leading to suboptimal certification performance. To address this issue, we propose a sample-specific certified backdoor defense method, termed Cert-SSB. Cert-SSB first employs stochastic gradient ascent to optimize the noise magnitude for each sample, ensuring a sample-specific noise level that is then applied to multiple poisoned training sets to retrain several smoothed models. After that, Cert-SSB aggregates the predictions of multiple smoothed models to generate the final robust prediction. In particular, in this case, existing certification methods become inapplicable since the optimized noise varies across different samples. To conquer this challenge, we introduce a storage-update-based certification method, which dynamically adjusts each sample's certification region to improve certification performance. We conduct extensive experiments on multiple benchmark datasets, demonstrating the effectiveness of our proposed method. Our code is available at https://github.com/NcepuQiaoTing/Cert-SSB.
Towards proactive self-adaptive AI for non-stationary environments with dataset shifts
Narro, David Fernández, Ferri, Pablo, García-Gómez, Juan M., Sáez, Carlos
Artificial Intelligence (AI) models deployed in production frequently face challenges in maintaining their performance in non-stationary environments. This issue is particularly noticeable in medical settings, where temporal dataset shifts often occur. These shifts arise when the distributions of training data differ from those of the data encountered during deployment over time. Further, new labeled data to continuously retrain AI is not typically available in a timely manner due to data access limitations. To address these challenges, we propose a proactive self-adaptive AI approach, or pro-adaptive, where we model the temporal trajectory of AI parameters, allowing us to short-term forecast parameter values. To this end, we use polynomial spline bases, within an extensible Functional Data Analysis framework. We validate our methodology with a logistic regression model addressing prior probability shift, covariate shift, and concept shift. This validation is conducted on both a controlled simulated dataset and a publicly available real-world COVID-19 dataset from Mexico, with various shifts occurring between 2020 and 2024. Our results indicate that this approach enhances the performance of AI against shifts compared to baseline stable models trained at different time distances from the present, without requiring updated training data. This work lays the foundation for pro-adaptive AI research against dynamic, non-stationary environments, being compatible with data protection, in resilient AI production environments for health.
MPEC: Manifold-Preserved EEG Classification via an Ensemble of Clustering-Based Classifiers
Shahbazi, Shermin, Nasiri, Mohammad-Reza, Ramezani, Majid
ORCID: 0000 - 0003 - 0886 - 7023 Abstract -- Accurate classification of EEG signals is crucial for brain - computer interfaces (BCIs) and neuroprosthetic applications, yet many existing methods fail to account for the non - Euclidean, manifold structure of EEG data, resulting in suboptimal performance. Preserving this manifold information is essential to capture the true geometry of EEG signals, but tradition al classification techniques largely overlook this need. To this end, w e propose MPEC (Manifold - Preserved EEG Classification via an Ensemble of Clus tering - Based Classifiers), that introduces two key innovations: (1) a feature engineering phase that combines covariance matrices and Radial Basis Function (RBF) kernels to capture both linear and non - linear relationships among EEG channels, and (2) a clustering phase that employs a modified K - means al gorithm tailored for the Riemannian manifold space, ensuring local geometric sensitivity. Ensembling multiple clustering - based classifiers, MPEC achieves superior results, validated by significant improvements on the BCI Competition IV dataset 2a. Keywords -- brain - computer interfaces (BCIs), EEG signal classification, ensemble modeling, clustering - based classification. EEG signal classification is essential in brain - computer interfaces (BCIs) and neuroprosthetics, where precise interpretation supports real - time control and cognitive applications. However, traditional techniques often overlook the non - Euclidean, manifold structure of EEG data, leading to suboptimal results [1] . We propose Manifold - Preserved EEG Classification via an Ensemble of Clustering - Based Classifiers (MPEC), a novel method that enhances classification accuracy by preserving the intrinsic manifold structure of EEG signals.
A comparative study of deep learning and ensemble learning to extend the horizon of traffic forecasting
Zheng, Xiao, Bagloee, Saeed Asadi, Sarvi, Majid
Traffic forecasting is vital for Intelligent Transportation Systems, for which Machine Learning (ML) methods have been extensively explored to develop data-driven Artificial Intelligence (AI) solutions. Recent research focuses on modelling spatial-temporal correlations for short-term traffic prediction, leaving the favourable long-term forecasting a challenging and open issue. This paper presents a comparative study on large-scale real-world signalized arterials and freeway traffic flow datasets, aiming to evaluate promising ML methods in the context of large forecasting horizons up to 30 days. Focusing on modelling capacity for temporal dynamics, we develop one ensemble ML method, eXtreme Gradient Boosting (XGBoost), and a range of Deep Learning (DL) methods, including Recurrent Neural Network (RNN)-based methods and the state-of-the-art Transformer-based method. Time embedding is leveraged to enhance their understanding of seasonality and event factors. Experimental results highlight that while the attention mechanism/Transformer framework is effective for capturing long-range dependencies in sequential data, as the forecasting horizon extends, the key to effective traffic forecasting gradually shifts from temporal dependency capturing to periodicity modelling. Time embedding is particularly effective in this context, helping naive RNN outperform Informer by 31.1% for 30-day-ahead forecasting. Meanwhile, as an efficient and robust model, XGBoost, while learning solely from time features, performs competitively with DL methods. Moreover, we investigate the impacts of various factors like input sequence length, holiday traffic, data granularity, and training data size. The findings offer valuable insights and serve as a reference for future long-term traffic forecasting research and the improvement of AI's corresponding learning capabilities.
Orthogonal Factor-Based Biclustering Algorithm (BCBOF) for High-Dimensional Data and Its Application in Stock Trend Prediction
Biclustering is an effective technique in data mining and pattern recognition. Biclustering algorithms based on traditional clustering face two fundamental limitations when processing high-dimensional data: (1) The distance concentration phenomenon in high-dimensional spaces leads to data sparsity, rendering similarity measures ineffective; (2) Mainstream linear dimensionality reduction methods disrupt critical local structural patterns. To apply biclustering to high-dimensional datasets, we propose an orthogonal factor-based bicluster-ing algorithm (BCBOF). First, we constructed orthogonal factors in the vector space of the high-dimensional dataset. Then, we performed clustering using the coordinates of the original data in the orthogonal subspace as clustering targets. Finally, we obtained biclustering results of the original dataset. Since dimensionality reduction was applied before clustering, the proposed algorithm effectively mitigated the data sparsity problem caused by high dimensionality. Additionally, we applied this biclustering algorithm to stock technical indicator combinations and stock price trend prediction. Biclustering results were transformed into fuzzy rules, and we incorporated profit-preserving and stop-loss rules into the rule set, ultimately forming a fuzzy inference system for stock price trend predictions and trading signals. The results showed that our algorithm outperformed other biclustering techniques. To validate the effectiveness of the fuzzy inference system, we conducted virtual trading experiments using historical data from 10 A-share stocks. The experimental results showed that the generated trading strategies yielded higher returns for investors. Introduction Since its initial proposal by Cheng and Church[1], biclustering has evolved into a sophisticated analytical approach.
Passive Measurement of Autonomic Arousal in Real-World Settings
Abdel-Ghaffar, Samy, Galatzer-Levy, Isaac, Heneghan, Conor, Liu, Xin, Kernasovskiy, Sarah, Garrett, Brennan, Barakat, Andrew, McDuff, Daniel
The autonomic nervous system (ANS) is activated during stress, which can have negative effects on cardiovascular health, sleep, the immune system, and mental health. While there are ways to quantify ANS activity in laboratories, there is a paucity of methods that have been validated in real-world contexts. We present the Fitbit Body Response Algorithm, an approach to continuous remote measurement of ANS activation through widely available remote wrist-based sensors. The design was validated via two experiments, a Trier Social Stress Test (n = 45) and ecological momentary assessments (EMA) of perceived stress (n=87), providing both controlled and ecologically valid test data. Model performance predicting perceived stress when using all available sensor modalities was consistent with expectations (accuracy=0.85) and outperformed models with access to only a subset of the signals. We discuss and address challenges to sensing that arise in real world settings that do not present in conventional lab environments.
Generalised Label-free Artefact Cleaning for Real-time Medical Pulsatile Time Series
Chen, Xuhang, Olakorede, Ihsane, Bögli, Stefan Yu, Xu, Wenhao, Beqiri, Erta, Li, Xuemeng, Tang, Chenyu, Gao, Zeyu, Gao, Shuo, Ercole, Ari, Smielewski, Peter
Artefacts compromise clinical decision-making in the use of medical time series. Pulsatile waveforms offer probabilities for accurate artefact detection, yet most approaches rely on supervised manners and overlook patient-level distribution shifts. To address these issues, we introduce a generalised label-free framework, GenClean, for real-time artefact cleaning and leverage an in-house dataset of 180,000 ten-second arterial blood pressure (ABP) samples for training. We first investigate patient-level generalisation, demonstrating robust performances under both intra- and inter-patient distribution shifts. We further validate its effectiveness through challenging cross-disease cohort experiments on the MIMIC-III database. Additionally, we extend our method to photoplethysmography (PPG), highlighting its applicability to diverse medical pulsatile signals. Finally, its integration into ICM+, a clinical research monitoring software, confirms the real-time feasibility of our framework, emphasising its practical utility in continuous physiological monitoring. This work provides a foundational step toward precision medicine in improving the reliability of high-resolution medical time series analysis