A spectral algorithm for robust regression with subgaussian rates
Much work concerning the prototypical problem of regression focuses on the study of rates of error of a given statistical procedure while making strong assumptions on the underlying distributions of samples, assuming for instance that they are i.i.d. and subgaussian or bounded (see for instance, [28, 41, 31]). It is however of fundamental importance to understand what happens when the data violates such strong assumptions, for instance, when the underlying distribution of samples is heavy-tailed and/or when the dataset is corrupted by outliers. In such cases - which are everyday cases for real-world datasets - classical estimators such as OLS or MLE exhibit, at best, far-from-optimal statistical behaviours and at worst completely non-sens outputs. In this work, we study the statistical properties (non-asymptotic estimations and predictions results) of algorithms coming with actual working code constructed on this type of real-word datasets. We want to put forward that it is an algorithm and not only a purely theoretical estimator and that this algorithm can be coded efficiently (we provide a simulation study in the following) since its most time consuming fundamental building block is to find a top singular vector of a reasonable size matrix. However, our theoretical results show that even though the dataset is far from the ideal i.i.d.
Jul-12-2020
- Country:
- North America > United States
- Ohio (0.04)
- Pennsylvania > Philadelphia County
- Philadelphia (0.04)
- New York > New York County
- New York City (0.04)
- New Jersey > Hudson County
- Hoboken (0.04)
- Europe
- North America > United States
- Genre:
- Research Report > New Finding (0.34)
- Technology: