Standardizing and Benchmarking Crisis-related Social Media Datasets for Humanitarian Information Processing
Alam, Firoj, Sajjad, Hassan, Imran, Muhammad, Ofli, Ferda
–arXiv.org Artificial Intelligence
Time-critical analysis of social media streams is important for humanitarian organizations to plan rapid response during disasters. The crisis informatics research community has developed several techniques and systems to process and classify big crisis related data posted on social media. However, due to the dispersed nature of the datasets used in the literature, it is not possible to compare the results and measure the progress made towards better models for crisis informatics. In this work, we attempt to bridge this gap by standardizing various existing crisis-related datasets. We consolidate labels of eight annotated data sources and provide 166.1k and 141.5k tweets for informativeness and humanitarian classification tasks, respectively. The consolidation results in a larger dataset that affords the ability to train more sophisticated models. To that end, we provide baseline results using CNN and BERT models.
arXiv.org Artificial Intelligence
Apr-29-2020
- Country:
- Asia
- Bangladesh (0.04)
- India (0.04)
- Malaysia (0.04)
- Middle East > Qatar
- Nepal (0.04)
- Pakistan (0.04)
- Philippines > Luzon
- National Capital Region > City of Manila (0.04)
- Singapore (0.04)
- Europe
- France > Île-de-France
- Iceland (0.04)
- Italy
- Emilia-Romagna > Metropolitan City of Bologna
- Bologna (0.04)
- Sardinia (0.04)
- Emilia-Romagna > Metropolitan City of Bologna
- United Kingdom
- England > Cambridgeshire
- Cambridge (0.04)
- Wales (0.04)
- England > Cambridgeshire
- North America
- Canada
- Alberta (0.04)
- Quebec > Estrie Region
- Lac-Mégantic (0.14)
- Costa Rica (0.04)
- Guatemala (0.04)
- Mexico (0.04)
- United States
- California
- Los Angeles County > Los Angeles (0.04)
- San Diego County > San Diego (0.04)
- Colorado (0.04)
- Missouri > Jasper County
- Joplin (0.04)
- Oklahoma (0.04)
- Texas (0.15)
- California
- Canada
- Oceania
- Australia
- New South Wales (0.05)
- Queensland (0.06)
- Vanuatu (0.04)
- Australia
- South America
- Asia
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Health & Medicine > Therapeutic Area (0.68)
- Information Technology (0.67)
- Transportation > Air (1.00)
- Technology: