Imbalanced Classification via a Tabular Translation GAN
Gradstein, Jonathan, Salhov, Moshe, Tulpan, Yoav, Lindenbaum, Ofir, Averbuch, Amir
Data that exhibits class imbalance appears frequently in real-world scenarios [1], in varying domains and applications: detecting pathologies or diseases in medical records [2], preventing network attacks in cybersecurity [3], detecting fraudulent financial transactions [4], distinguishing between earthquakes and explosions [5] and detecting spam communications [6]. In addition to these applications, where the class distribution is naturally skewed due to the frequency of events, some applications may exhibit class imbalance caused by extrinsic factors such as collection and storage limitations [7]. Most standard classification models are designed around and implicitly assume a relatively balanced class distribution; when applied without proper adjustments they may fail to accurately model the minority class and converge on a solution that over-classifies the majority class due to its increased prior probability [8]. These models thus neglect recall on the minority class and lead to unsatisfactory results when we desire high performance on a more balanced testing criterion. This issue is exacerbated by the fact that commonly used metrics such as accuracy may be misleading in evaluating the performance of the model. Even models that naively classify all samples as majority may have high accuracy under severe class imbalance. Most approaches to dealing with these shortcomings fall broadly into two categories: re-weighting the loss objective to more heavily account for the minority class, and resampling the input dataset such that the minority class is more prominent.
Apr-19-2022
- Country:
- North America > United States (0.04)
- Asia > Middle East
- Israel > Tel Aviv District > Tel Aviv (0.04)
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Information Technology > Security & Privacy (1.00)
- Health & Medicine (0.86)
- Technology: