Fair and Diverse DPP-based Data Summarization
Celis, L. Elisa, Keswani, Vijay, Straszak, Damian, Deshpande, Amit, Kathuria, Tarun, Vishnoi, Nisheeth K.
A problem facing many services - from search engines and news feeds to machine learning - is data summarization: how can one select a small but representative, i.e., diverse, subset from a large dataset. For instance, Google Images outputs a small subset of images from its enormous dataset given a user query. Similarly, in training a learning algorithm one may be required to choose a subset of data points to train on as training on the entire dataset may be costly. However, data summarization algorithms prevalent in the online world have been recently shown to be biased with respect to sensitive attributes such as gender, race and ethnicity. For instance, a recent study found evidence of systematic under-representation of women in search results [14]. Concretely, the above work studied the output of Google Images for various search terms involving occupations and found, e.g., that for the search term "CEO", the percentage of women in top 100 results was 11%, significantly lower than the ground truth of 27%. Through studies on human subjects, they also found that such misrepresentations have the power to influence people's perception about reality. Beyond humans, since data summaries are used to train algorithms, there is a danger that these biases in the data might be passed on to the algorithms that use them; a phenomena that is being revealed more and more in automated data-driven processes in education, recruitment, banking, and judiciary systems, see [22]. A robust and widely deployed method for data summarization is to associate a diversity score to each subset and select a subset with probability proportional to this score; see [13].
Feb-12-2018
- Country:
- Europe (1.00)
- North America > United States
- New York (0.28)
- Genre:
- Research Report > New Finding (1.00)
- Technology: