Statistical Learning
Beyond Effi ciency: Molecular Data Pruning for Enhanced Generalization
With the emergence of various molecular tasks and massive datasets, how to perform e ffi cient training has become an urgent yet under-explored issue in the area. Data pruning (DP), as an oft-stated approach to saving training burdens, filters out less influential samples to form a coreset for training.