AlbNews: A Corpus of Headlines for Topic Modeling in Albanian
–arXiv.org Artificial Intelligence
The scarcity of available text corpora for low-resource languages like Albanian is a serious hurdle for research in natural language processing tasks. This paper introduces AlbNews, a collection of 600 topically labeled news headlines and 2600 unlabeled ones in Albanian. The data can be freely used for conducting topic modeling research. We report the initial classification scores of some traditional machine learning classifiers trained with the AlbNews samples. These results show that basic models outrun the ensemble learning ones and can serve as a baseline for future experiments.
arXiv.org Artificial Intelligence
Feb-6-2024
- Country:
- North America
- United States
- Hawaii (0.04)
- District of Columbia > Washington (0.04)
- New York > New York County
- New York City (0.15)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Massachusetts > Suffolk County
- Boston (0.04)
- Canada > British Columbia
- United States
- Europe
- Austria > Vienna (0.14)
- Slovakia > Bratislava
- Bratislava (0.04)
- France > Île-de-France
- Faroe Islands > Streymoy
- Tórshavn (0.04)
- Estonia > Tartu County
- Tartu (0.04)
- Asia
- Singapore (0.04)
- Middle East > Jordan (0.04)
- Japan > Honshū
- Kantō > Tokyo Metropolis Prefecture > Tokyo (0.14)
- China > Heilongjiang Province
- Daqing (0.04)
- North America
- Genre:
- Research Report > New Finding (0.35)
- Technology: