Active Learning for Regression based on Wasserstein distance and GroupSort Neural Networks
Bobbia, Benjamin, Picard, Matthias
Collecting data is a significant challenge in machine learning, and more generally in statistics. The amount of data necessary to get a sharp estimation of a function can get unreasonably high, especially when the task involves highdimensional objects. In various applications, extracting and gathering enough unlabeled data is not a big challenge. However, labeling them can be a very costly and time-consuming process. In fields such as statistical physics, we would often need to run complex simulations or call an expert to label data manually. Hence we want to be able to get a satisfying estimation with a more compact set of labeled data. Among the proposed solutions, we can mention fewshot learning (Li et al., 2006) and transfer learning (Ben-David et al., 2010). Those solutions leverage from another model previously trained on a similar task to simplify our model's training. To dodge the issue, we can also resort to generative adversarial models (Goodfellow et al., 2014) or diffusion models (Trabucco et al., 2023): the idea is to augment a dataset using generated data,
Mar-22-2024
- Country:
- North America > United States
- California > San Francisco County > San Francisco (0.14)
- Europe > Netherlands
- South Holland > Delft (0.04)
- North America > United States
- Genre:
- Research Report (0.40)
- Technology: