A spectral algorithm for robust regression with subgaussian rates

Depersin, Jules

arXiv.org Machine Learning 

Much work concerning the prototypical problem of regression focuses on the study of rates of error of a given statistical procedure while making strong assumptions on the underlying distributions of samples, assuming for instance that they are i.i.d. and subgaussian or bounded (see for instance, [28, 41, 31]). It is however of fundamental importance to understand what happens when the data violates such strong assumptions, for instance, when the underlying distribution of samples is heavy-tailed and/or when the dataset is corrupted by outliers. In such cases - which are everyday cases for real-world datasets - classical estimators such as OLS or MLE exhibit, at best, far-from-optimal statistical behaviours and at worst completely non-sens outputs. In this work, we study the statistical properties (non-asymptotic estimations and predictions results) of algorithms coming with actual working code constructed on this type of real-word datasets. We want to put forward that it is an algorithm and not only a purely theoretical estimator and that this algorithm can be coded efficiently (we provide a simulation study in the following) since its most time consuming fundamental building block is to find a top singular vector of a reasonable size matrix. However, our theoretical results show that even though the dataset is far from the ideal i.i.d.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found