Imbalanced Continual Learning with Partitioning Reservoir Sampling
Kim, Chris Dongjoo, Jeong, Jinseo, Kim, Gunhee
–arXiv.org Artificial Intelligence
Continual learning from a sequential stream of data is a crucial challenge for machine learning research. Most studies have been conducted on this topic under the single-label classification setting along with an assumption of balanced label distribution. This work expands this research horizon towards multi-label classification. In doing so, we identify unanticipated adversity innately existent in many multi-label datasets, the long-tailed distribution. We jointly address the two independently solved problems, Catastropic Forgetting and the long-tailed label distribution by first empirically showing a new challenge of destructive forgetting of the minority concepts on the tail. Then, we curate two benchmark datasets, COCOseq and NUS-WIDEseq, that allow the study of both intra- and inter-task imbalances. Lastly, we propose a new sampling strategy for replay-based approach named Partitioning Reservoir Sampling (PRS), which allows the model to maintain a balanced knowledge of both head and tail classes. We publicly release the dataset and the code in our project page.
arXiv.org Artificial Intelligence
Sep-8-2020
- Country:
- North America > United States
- California > Alameda County > Berkeley (0.04)
- Europe > Belgium
- Flanders > Flemish Brabant > Leuven (0.04)
- Asia
- Singapore (0.04)
- South Korea > Seoul
- Seoul (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Education > Educational Setting > Online (0.46)
- Technology: