Orthogonalized Estimation of Difference of $Q$-functions
Learning optimal dynamic treatment rules, or sequential policies for taking actions, is important, although often only observational data is available. Many recent works in offline reinforcement learning develop methodology to evaluate and optimize sequential decision rules, without the ability to conduct online exploration. An extensive literature on causal inference and machine learning establishes methodologies for learning causal contrasts, such as the conditional average treatment effect (CATE) (Wager and Athey, 2018; Foster and Syrgkanis, 2019; Künzel et al., 2019; Kennedy, 2020), which is sufficient for making optimal decisions. Methods that specifically estimate causal contrasts (such as the CATE), can better adapt to potentially smoother or more structured contrast functions, while methods that instead contrast estimates (by taking the difference of outcome regressions or Q functions) can not. Additionally, estimation of causal contrasts can be improved via orthogonalization or double machine learning (Kennedy, 2022; Chernozhukov et al., 2018). Estimating the causal contrast is both sufficient for optimal decisions and statistically favorable.
Jun-12-2024
- Country:
- North America > United States
- California (0.14)
- Europe > United Kingdom
- England > Cambridgeshire > Cambridge (0.04)
- North America > United States
- Genre:
- Research Report (0.40)
- Technology: