Statistical Learning
On the Convergence of Loss and Uncertainty-based Active Learning Algorithms
We investigate the convergence rates and data sample sizes required for training a machine learning model using a stochastic gradient descent (SGD) algorithm, where data points are sampled based on either their loss value or uncertainty value. These training methods are particularly relevant for active learning and data subset selection problems. For SGD with a constant step size update, we present convergence results for linear classifiers and linearly separable datasets using squared hinge loss and similar training loss functions.
A Non-parametric Direct Learning Approach to Heterogeneous Treatment Effect Estimation under Unmeasured Confounding
In various domains, different subjects may exhibit different responses to the same set of treatments. The exploration of this heterogeneity in the effects resulting from exposure has gained substantial interest in recent years. For instance, inferring the heterogeneous effect of a medical treatment on clinical outcome can contribute to the development of personalized treatment (Cai et al., 2011). A similar concept has found application in personalized marketing as well (Chandra et al., 2022).