Double machine learning for sample selection models
Bia, Michela, Huber, Martin, Lafférs, Lukáš
This paper considers treatment evaluation when outcomes are only observed for a subpopulation due to sample selection or outcome attrition/non-response. For identification, we combine a selection-on-observables assumption for treatment assignment with either selection-on-observables or instrumental variable assumptions concerning the outcome attrition/sample selection process. To control in a data-driven way for potentially high dimensional pre-treatment covariates that motivate the selectionon-observables assumptions, we adapt the double machine learning framework to sample selection problems. That is, we make use of (a) Neyman-orthogonal and doubly robust score functions, which imply the robustness of treatment effect estimation to moderate regularization biases in the machine learningbased estimation of the outcome, treatment, or sample selection models and (b) sample splitting (or cross-fitting) to prevent overfitting bias. We demonstrate that the proposed estimators are asymptotically normal and root-n consistent under specific regularity conditions concerning the machine learners and investigate their finite sample properties in a simulation study. The estimator is available in the causalweight package for the statistical software R. Keywords: sample selection, double machine learning, doubly robust estimation, efficient score.
Dec-9-2020
- Country:
- Europe
- Slovakia > Banska Bystrica
- Banská Bystrica (0.04)
- Switzerland > Fribourg
- Fribourg (0.04)
- Slovakia > Banska Bystrica
- North America > United States
- Illinois > Cook County
- Chicago (0.04)
- Michigan (0.04)
- New York (0.04)
- Illinois > Cook County
- South America > Colombia (0.04)
- Europe
- Genre:
- Research Report (1.00)
- Industry:
- Education (0.67)
- Technology: