Data splitting improves statistical performance in overparametrized regimes
Mücke, Nicole, Reiss, Enrico, Rungenhagen, Jonas, Klein, Markus
Modern machine learning applications often involve learning statistical models of great complexity and datasets of massive size become increasingly available. However, while increasing the size of the training datasets generally offers improvement in model performance, the training process is very computation-intensive and thus time-consuming. Indeed, hardware architectures have physical limits in terms of storage, memory, processing speed and communication. A central challenge is thus to design efficient large-scale algorithms. Distributed learning and parallel computing is a common and simple approach to deal with large datasets. The n observations are evenly split to M machines (or local nodes, workers), each having access to only a subset of n/M training samples. Each machine performs local computations to fit a model and transmits it to a central node for merging. This simple divide and conquer approach having been proposed in e.g.
Oct-21-2021
- Country:
- Europe
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- Germany > Brandenburg
- Potsdam (0.04)
- United Kingdom > England
- Asia > Middle East
- Jordan (0.04)
- Europe
- Genre:
- Research Report (0.82)
- Technology: