Triply Robust Off-Policy Evaluation

Liu, Anqi, Liu, Hao, Anandkumar, Anima, Yue, Yisong

Nov-15-2019–arXiv.org Machine Learning

We frame OPE as a covariate-shift problem and leverage modern robust regression tools. Ours is a general approach that can be used to augment any existing OPE method that utilizes the direct method. When augmenting doubly robust methods, we call the resulting method triply robust, since we add robustness to the direct method used in doubly robust. We prove upper bounds on the resulting bias and variance, as well as derive novel minimax bounds based on robust minimax analysis for covariate shift. Our robust regression method is compatible with deep learning, and is thus applicable to complex OPE settings that require powerful function approximators. Finally, we demonstrate superior empirical performance across the standard OPE benchmarks, especially in the case where the logging policy is unknown and must be estimated from data. 1 Introduction Contextual bandits is the online learning setting where a policy repeatedly observes a context, takes an action, and then observes a reward only for the chosen action [Langford and Zhang, 2007].

evaluation, robust regression, variance, (13 more...)

arXiv.org Machine Learning

Nov-15-2019

arXiv.org PDF

Add feedback

Country:
- North America > United States > Pennsylvania > Allegheny County > Pittsburgh (0.04)

Genre:
- Research Report (1.00)

Technology:
- Information Technology
  - Data Science > Data Mining (0.93)
  - Artificial Intelligence
    - Representation & Reasoning > Search (0.70)
    - Machine Learning
      - Statistical Learning > Regression (0.48)
      - Neural Networks > Deep Learning (0.48)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found