Goto

Collaborating Authors

 Statistical Learning


Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator

arXiv.org Artificial Intelligence

In addition to a policy update, AC methods employ a parallel critic update to bootstrap the Q-value for policy gradient estimation, which often enjoys reduced variance and fast convergence in training. Despite the empirical success, theoretical analysis of AC in the most practical form remains challenging. Existing works mostly focus on either the double-loop or the two-timescale variants. In double-loop AC, the actor is updated in the outer loop only after the critic takes sufficiently many steps to have an accurate estimation of the Q-value in the inner loop [ Y anget al., 2019; Kumar et al., 2019; Wang et al., 2019 ] . Hence, the convergence of the critic is decoupled from that of the actor. The analysis is separated into a policy evaluation sub-problem in the inner loop and a perturbed gradient descent in the outer loop. In two-timescale AC, the actor and the critic are updated simultaneously in each iteration using stepsizes of different timescales. The actor stepsize (denoted by ฮฑ t in the sequel) is typically smaller than that of the critic (denoted by ฮฒ t in the sequel), with their ratio going to zero as the iteration number goes to infinity (i.e., lim t ฮฑ t/ฮฒ t = 0). The two-timescale allows the critic to approximate the correct Q-value asymptotically.


Adaptive and Robust DBSCAN with Multi-agent Reinforcement Learning

arXiv.org Artificial Intelligence

DBSCAN, a well-known density-based clustering algorithm, has gained widespread popularity and usage due to its effectiveness in identifying clusters of arbitrary shapes and handling noisy data. However, it encounters challenges in producing satisfactory cluster results when confronted with datasets of varying density scales, a common scenario in real-world applications. In this paper, we propose a novel Adaptive and Robust DBSCAN with Multi-agent Reinforcement Learning cluster framework, namely AR-DBSCAN. First, we model the initial dataset as a two-level encoding tree and categorize the data vertices into distinct density partitions according to the information uncertainty determined in the encoding tree. Each partition is then assigned to an agent to find the best clustering parameters without manual assistance. The allocation is density-adaptive, enabling AR-DBSCAN to effectively handle diverse density distributions within the dataset by utilizing distinct agents for different partitions. Second, a multi-agent deep reinforcement learning guided automatic parameter searching process is designed. The process of adjusting the parameter search direction by perceiving the clustering environment is modeled as a Markov decision process. Using a weakly-supervised reward training policy network, each agent adaptively learns the optimal clustering parameters by interacting with the clusters. Third, a recursive search mechanism adaptable to the data's scale is presented, enabling efficient and controlled exploration of large parameter spaces. Extensive experiments are conducted on nine artificial datasets and a real-world dataset. The results of offline and online tasks show that AR-DBSCAN not only improves clustering accuracy by up to 144.1% and 175.3% in the NMI and ARI metrics, respectively, but also is capable of robustly finding dominant parameters.


Boosting Statistic Learning with Synthetic Data from Pretrained Large Models

arXiv.org Machine Learning

The rapid advancement of generative models, such as Stable Diffusion, raises a key question: how can synthetic data from these models enhance predictive modeling? While they can generate vast amounts of datasets, only a subset meaningfully improves performance. We propose a novel end-to-end framework that generates and systematically filters synthetic data through domain-specific statistical methods, selectively integrating high-quality samples for effective augmentation. Our experiments demonstrate consistent improvements in predictive performance across various settings, highlighting the potential of our framework while underscoring the inherent limitations of generative models for data augmentation. Despite the ability to produce large volumes of synthetic data, the proportion that effectively improves model performance is limited.


Local linear Frรฉchet curve regression in manifolds

arXiv.org Machine Learning

Global Frรฉchet functional regression has been recently addressed from time correlated bivariate curve data evaluated in a manifold (see Torres et al. 2025). For this type of curve data sets, the present paper solves the problem of local linear approximation of the Frรฉchet conditional mean in an extrinsic and intrinsic way. The extrinsic local linear Frรฉchet functional regression predictor is obtained in the time varying tangent space by projection into an orthornormal basis of the ambient Hilbert space. The conditions assumed ensure the existence and uniqueness of this predictor, and its computation via exponential and logarithmic maps. A weighted Frรฉchet mean approach is adopted in the computation of an intrinsic local linear Frรฉchet functional regression predictor. The asymptotic optimality of this intrinsic local approximation is also proved. The performance of the empirical version of both, extrinsic and intrinsic functional predictors, and of a Nadaraya-Watson type Frรฉchet curve predictor is illustrated in the simulation study undertaken. The finite-sample size properties are also tested in a real-data application via cross-validation. Specifically, functional prediction of the magnetic vector field from the time-varying geocentric latitude and longitude of the satellite NASA's MAGSAT spacecraft is addressed.


Conformal Prediction with Cellwise Outliers: A Detect-then-Impute Approach

arXiv.org Machine Learning

Conformal prediction is a powerful tool for constructing prediction intervals for black-box models, providing a finite sample coverage guarantee for exchangeable data. However, this exchangeability is compromised when some entries of the test feature are contaminated, such as in the case of cellwise outliers. To address this issue, this paper introduces a novel framework called detect-then-impute conformal prediction. This framework first employs an outlier detection procedure on the test feature and then utilizes an imputation method to fill in those cells identified as outliers. To quantify the uncertainty in the processed test feature, we adaptively apply the detection and imputation procedures to the calibration set, thereby constructing exchangeable features for the conformal prediction interval of the test label. We develop two practical algorithms, PDI-CP and JDI-CP, and provide a distribution-free coverage analysis under some commonly used detection and imputation procedures. Notably, JDI-CP achieves a finite sample $1-2ฮฑ$ coverage guarantee. Numerical experiments on both synthetic and real datasets demonstrate that our proposed algorithms exhibit robust coverage properties and comparable efficiency to the oracle baseline.


Precise gradient descent training dynamics for finite-width multi-layer neural networks

arXiv.org Machine Learning

In this paper, we provide the first precise distributional characterization of gradient descent iterates for general multi-layer neural networks under the canonical single-index regression model, in the `finite-width proportional regime' where the sample size and feature dimension grow proportionally while the network width and depth remain bounded. Our non-asymptotic state evolution theory captures Gaussian fluctuations in first-layer weights and concentration in deeper-layer weights, and remains valid for non-Gaussian features. Our theory differs from existing neural tangent kernel (NTK), mean-field (MF) theories and tensor program (TP) in several key aspects. First, our theory operates in the finite-width regime whereas these existing theories are fundamentally infinite-width. Second, our theory allows weights to evolve from individual initializations beyond the lazy training regime, whereas NTK and MF are either frozen at or only weakly sensitive to initialization, and TP relies on special initialization schemes. Third, our theory characterizes both training and generalization errors for general multi-layer neural networks beyond the uniform convergence regime, whereas existing theories study generalization almost exclusively in two-layer settings. As a statistical application, we show that vanilla gradient descent can be augmented to yield consistent estimates of the generalization error at each iteration, which can be used to guide early stopping and hyperparameter tuning. As a further theoretical implication, we show that despite model misspecification, the model learned by gradient descent retains the structure of a single-index function with an effective signal determined by a linear combination of the true signal and the initialization.


Clustering with Communication: A Variational Framework for Single Cell Representation Learning

arXiv.org Machine Learning

Single-cell RNA sequencing (scRNA-seq) has revealed complex cellular heterogeneity, but recent studies emphasize that understanding biological function also requires modeling cell-cell communication (CCC), the signaling interactions mediated by ligand-receptor pairs that coordinate cellular behavior. Tools like CellChat have demonstrated that CCC plays a critical role in processes such as cell differentiation, tissue regeneration, and immune response, and that transcriptomic data inherently encodes rich information about intercellular signaling. We propose CCCVAE, a novel variational autoencoder framework that incorporates CCC signals into single-cell representation learning. By leveraging a communication-aware kernel derived from ligand-receptor interactions and a sparse Gaussian process, CCCVAE encodes biologically informed priors into the latent space. Unlike conventional VAEs that treat each cell independently, CCCVAE encourages latent embeddings to reflect both transcriptional similarity and intercellular signaling context. Empirical results across four scRNA-seq datasets show that CCCVAE improves clustering performance, achieving higher evaluation scores than standard VAE baselines. This work demonstrates the value of embedding biological priors into deep generative models for unsupervised single-cell analysis.


Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering Perspective

arXiv.org Artificial Intelligence

Although numerous complex algorithms for treatment effect estimation have been developed in recent years, their effectiveness remains limited when handling insufficiently labeled training sets due to the high cost of labeling the effect after treatment, e.g., expensive tumor imaging or biopsy procedures needed to evaluate treatment effects. Therefore, it becomes essential to actively incorporate more high-quality labeled data, all while adhering to a constrained labeling budget. To enable data-efficient treatment effect estimation, we formalize the problem through rigorous theoretical analysis within the active learning context, where the derived key measures -- \textit{factual} and \textit{counterfactual covering radius} determine the risk upper bound. To reduce the bound, we propose a greedy radius reduction algorithm, which excels under an idealized, balanced data distribution. To generalize to more realistic data distributions, we further propose FCCM, which transforms the optimization objective into the \textit{Factual} and \textit{Counterfactual Coverage Maximization} to ensure effective radius reduction during data acquisition. Furthermore, benchmarking FCCM against other baselines demonstrates its superiority across both fully synthetic and semi-synthetic datasets.


Incentive-Aware Machine Learning; Robustness, Fairness, Improvement & Causality

arXiv.org Artificial Intelligence

Machine Learning (ML) algorithms are deeply embedded in var ious aspects of modern life, influencing everything from enhancing daily conveniences and sh aping online purchasing behavior to making critical decisions in areas such as hiring, loan appr ovals, college admissions, and probation rulings. Given the high stakes of these decisions, individu als often have strong incentives to strategically modify the data they provide to these algorithms to s ecure more favorable outcomes. For instance, individuals might open additional credit accoun ts or take other steps to improve their credit scores before applying for a loan. In the context of co llege admissions, applicants may retake standardized tests like the GRE, enroll in test preparation courses, or even switch schools to boost their class rankings, all in efforts to present themselves as m ore competitive candidates. Such instances of "strategic adaptation" have been extensi vely documented across disciplines including Economics, CS, and Public Policy Bj orkegren et al. [ 2020 ], Dee et al. [ 2019 ], Dranove et al. [ 2003 ], Greenstone et al. [ 2022 ], Gonzalez-Lira and Mobarak [ 2019 ], Chang et al. [ 2024 ]. The challenge arises when decision-makers deploying ML algorithms fail to account for these adaptations, potentially undermining the original goals of the policies the algorithms are intended to support. For example, in college admissions, a student's decision to change schools solely to improve their class ranking may not necessarily reflect a substantive impr ovement in their qualifications. This literature review was recently published in SIGEcom Ex changes.


Balancing Client Participation in Federated Learning Using AoI

arXiv.org Artificial Intelligence

--Federated Learning (FL) offers a decentralized framework that preserves data privacy while enabling collaborative model training across distributed clients. However, FL faces significant challenges due to limited communication resources, statistical heterogeneity, and the need for balanced client participation. This paper proposes an Age of Information (AoI)- based client selection policy that addresses these challenges by minimizing load imbalance through controlled selection intervals. Our method employs a decentralized Markov scheduling policy, allowing clients to independently manage participation based on age-dependent selection probabilities, which balances client updates across training rounds with minimal central oversight. We provide a convergence proof for our method, demonstrating that it ensures stable and efficient model convergence. Specifically, we derive optimal parameters for the Markov selection model to achieve balanced and consistent client participation, highlighting the benefits of AoI in enhancing convergence stability. Through extensive simulations, we demonstrate that our AoI-based method, particularly the optimal Markov variant, improves convergence over the FedA vg selection approach across both IID and non-IID data settings by 7. 5% and up to 20%. Our findings underscore the effectiveness of AoI-based scheduling for scalable, fair, and efficient FL systems across diverse learning environments. Federated learning (FL), introduced by McMahan et al. [1], emerged as a solution to the limitations of traditional machine learning models that require centralized data collection and processing. FL enables client devices to collaboratively train a global model while keeping all the training data localized, thus addressing privacy concerns. Traditional machine-learning approaches require centralized data training in data centers, which often becomes impractical for edge devices due to privacy constraints in wireless networks and limited wireless communication resources. Federated learning overcomes these challenges by enabling devices to train machine learning models without data sharing and transmission, fulfilling the needs of data privacy and security. Compared to traditional distributed machine learning, Federated Learning introduces several challenges [2], including system heterogeneity from diverse device capabilities causing aggregation delays due to stragglers, statistical heterogeneity arising from non-IID and imbalanced client data affecting model convergence, and privacy concerns as exchanged model updates may inadvertently expose sensitive information. This paper is presented in part at the 2024 IEEE Global Communications Conference (Globecom). The authors are with the Center for Pervasive Communications and Computing, University of California, Irvine (e-mail: ajavani@uci.edu,