AlbNER: A Corpus for Named Entity Recognition in Albanian
–arXiv.org Artificial Intelligence
Scarcity of resources such as annotated text corpora for under-resourced languages like Albanian is a serious impediment in computational linguistics and natural language processing research. This paper presents AlbNER, a corpus of 900 sentences with labeled named entities, collected from Albanian Wikipedia articles. Preliminary results with BERT and RoBERTa variants fine-tuned and tested with AlbNER data indicate that model size has slight impact on NER performance, whereas language transfer has a significant one. AlbNER corpus and these obtained results should serve as baselines for future experiments.
arXiv.org Artificial Intelligence
Sep-15-2023
- Country:
- Asia > Indonesia (0.04)
- Oceania > Australia
- North America > United States
- Oregon (0.04)
- District of Columbia > Washington (0.04)
- Ohio > Franklin County
- Columbus (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Michigan > Washtenaw County
- Ann Arbor (0.04)
- Europe
- Austria > Vienna (0.14)
- Switzerland (0.04)
- Spain
- Canary Islands (0.04)
- Catalonia > Barcelona Province
- Barcelona (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Albania > Korçë County
- Korçë (0.04)
- Genre:
- Research Report (0.40)
- Technology: