Data Twinning
Vakayil, Akhil, Joseph, V. Roshan
Often in statistics and machine learning we are required to partition a dataset, e.g., when (i) splitting a dataset for training and testing, (ii) subsampling from Big Data for conducting tractable statistical analysis or to save storage space, (iii) generating multiple splits of a dataset for divide-and-conquer procedures to act upon, and (iv) creating k-fold cross validation sets for model tuning and validation. For this purpose, we propose a novel method named Twinning that can be used for partitioning a dataset into statistically similar sets. Twinning is motivated from the recent work on optimal data splitting for model validation, by Joseph and Vakayil (2021). For model validation, the common practice is to randomly split the dataset into training and testing sets, e.g., for an 80-20 split, 20% of the dataset is selected randomly for testing, while the remaining 80% is used for training the model. It is easy to see that such random splitting can plausibly give rise to pathological splits, wherein the training and testing sets cover roughly disjoint regions of the feature space, thereby resulting in poor testing performance of the model.
Oct-6-2021
- Country:
- North America > United States
- New York (0.04)
- Georgia > Fulton County
- Atlanta (0.04)
- North America > United States
- Genre:
- Research Report (0.84)
- Industry:
- Technology: