Goto

Collaborating Authors

 albanian


AlbNews: A Corpus of Headlines for Topic Modeling in Albanian

arXiv.org Artificial Intelligence

The scarcity of available text corpora for low-resource languages like Albanian is a serious hurdle for research in natural language processing tasks. This paper introduces AlbNews, a collection of 600 topically labeled news headlines and 2600 unlabeled ones in Albanian. The data can be freely used for conducting topic modeling research. We report the initial classification scores of some traditional machine learning classifiers trained with the AlbNews samples. These results show that basic models outrun the ensemble learning ones and can serve as a baseline for future experiments.


AlbNER: A Corpus for Named Entity Recognition in Albanian

arXiv.org Artificial Intelligence

Scarcity of resources such as annotated text corpora for under-resourced languages like Albanian is a serious impediment in computational linguistics and natural language processing research. This paper presents AlbNER, a corpus of 900 sentences with labeled named entities, collected from Albanian Wikipedia articles. Preliminary results with BERT and RoBERTa variants fine-tuned and tested with AlbNER data indicate that model size has slight impact on NER performance, whereas language transfer has a significant one. AlbNER corpus and these obtained results should serve as baselines for future experiments.


AlbMoRe: A Corpus of Movie Reviews for Sentiment Analysis in Albanian

arXiv.org Artificial Intelligence

Lack of available resources such as text corpora for low-resource languages seriously hinders research on natural language processing and computational linguistics. This paper presents AlbMoRe, a corpus of 800 sentiment annotated movie reviews in Albanian. Each text is labeled as positive or negative and can be used for sentiment analysis research. Preliminary results based on traditional machine learning classifiers trained with the AlbMoRe samples are also reported. They can serve as comparison baselines for future research experiments.


Language: People across cultures agree the word 'bouba' sounds round while 'kiki' sounds pointy

Daily Mail - Science & tech

From English to Chinese, Hungarian and Zulu, people who speak different languages make the same links between sounds and shapes, a new study shows. An international research team has conducted the largest cross-cultural test of the'bouba-kiki effect' – the tendency to associate made-up words'bouba' with a round shape and'kiki' with a spiky shape. The researchers surveyed 917 speakers of 25 different languages representing nine language families and 10 writing systems. They found the effect exists independently of the language that a person speaks or the writing system that they use, whether it's the Roman alphabet (A, B, C), the Greek alphabet (alpha, beta, gamma) or Chinese characters (北, 方, 话). Such universally-meaningful vocalisations may form a global basis for the creation of new words, such as terms that circulate on social media. Bouba and kiki shapes used in the experiment.