Goto

Collaborating Authors

 Oceania


412758d043dd247bddea07c7ec558c31-Paper.pdf

Neural Information Processing Systems

Numerical results demonstrate significant regret reductions by our method, in comparison with several baselines in a range of multi-armed and linear bandit problems.




Appendices

Neural Information Processing Systems

Each dataset contains miscellaneous series, categorized into six domains (micro, industry, macro, finance, demographic, other). Thus, atime series regression dataset consists ofT input-target pairs: {(X1,y1),...,(XT,yT). For each synthesized training set withT samples, we synthesize100T samples as the testing set. D.3 Models The NN we use has six fully-connected layers with ReLU activation function and three residual connections. D.4 Results There are three methods tobe compared.


Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood

arXiv.org Machine Learning

Policy inference plays an essential role in the contextual bandit problem. In this paper, we use empirical likelihood to develop a Bayesian inference method for the joint analysis of multiple contextual bandit policies in finite sample regimes. The proposed inference method is robust to small sample sizes and is able to provide accurate uncertainty measurements for policy value evaluation. In addition, it allows for flexible inferences on policy comparison with full uncertainty quantification. We demonstrate the effectiveness of the proposed inference method using Monte Carlo simulations and its application to an adolescent body mass index data set.


Do More Predictions Improve Statistical Inference? Filtered Prediction-Powered Inference

arXiv.org Machine Learning

Recent advances in artificial intelligence have enabled the generation of large-scale, low-cost predictions with increasingly high fidelity. As a result, the primary challenge in statistical inference has shifted from data scarcity to data reliability. Prediction-powered inference methods seek to exploit such predictions to improve efficiency when labeled data are limited. However, existing approaches implicitly adopt a use-all philosophy, under which incorporating more predictions is presumed to improve inference. When prediction quality is heterogeneous, this assumption can fail, and indiscriminate use of unlabeled data may dilute informative signals and degrade inferential accuracy. In this paper, we propose Filtered Prediction-Powered Inference (FPPI), a framework that selectively incorporates predictions by identifying a data-adaptive filtered region in which predictions are informative for inference. We show that this region can be consistently estimated under a margin condition, achieving fast rates of convergence. By restricting the prediction-powered correction to the estimated filtered region, FPPI adaptively mitigates the impact of biased or noisy predictions. We establish that FPPI attains strictly improved asymptotic efficiency compared with existing prediction-powered inference methods. Numerical studies and a real-data application to large language model evaluation demonstrate that FPPI substantially reduces reliance on expensive labels by selectively leveraging reliable predictions, yielding accurate inference even in the presence of heterogeneous prediction quality.