On Least Squares Estimation under Heteroscedastic and Heavy-Tailed Errors
Kuchibhotla, Arun K., Patra, Rohit K.
The study of the LSE has received considerable attention in statistics as well as machine learning; see [9, 36, 41, 54, 55] for important contributions. LSEs are particularly useful when F is known to satisfy some shape constraints such as monotonicity, convexity, or unimodality. In such cases, the LSEs are tuning parameter-free, can be computed as the solution to convex optimization problems, and are adaptive, i.e., the rate of convergence of the LSE changes depending on the "structure" of f 0 [4, 23, 28, 43, 47]. For example, if F is the class of monotone functions and f 0 is a strictly increasing function then ˆ f converges at an n 1/ 3 rate (i.e., null ˆ f f 0null O p(n 1/ 3)); however, if f 0 0, then ˆ f converges at an n 1/ 2 rate [10, 25, 56]. Tsirelson [50] and van de Geer and Wegkamp [53] have established necessary and sufficient conditions on F and null for consistency of ˆ f . Our goal in this work is to provide some general sufficient conditions on F and null under which the LSE is "rate-optimal"; see discussions after Corollary 2.1 for more details, also see Example 3.1. Our results significantly expand the scenarios under which LSE can be proven to be rate-optimal or "be a safe choice" for the model at hand. Theorem 3.2.5 of [55] implies that the LSE ˆ f defined on F satisfies null ˆ f f 0null O p(r n) for any r n such that E null sup f F: null f f 0null r nnull null G n[2null ( f f 0)( X) (f f 0) 2 (X)]null null null C nr 2 n, (3) where C denotes a constant 1 . In the rest of this paper, we make the convention that the constant C is not necessarily the same on each occurrence.
Sep-14-2019