bulgaria
Germany warns of 'daily hybrid warfare' after explosive-laden drone found
Is the war entering a new phase? Germany warns of'daily hybrid warfare' after explosive-laden drone found Germany is facing daily "hybrid warfare" attacks from abroad, an official has said, days after an explosive-laden drone was discovered at Leipzig/Halle Airport. Interior Minister Alexander Dobrindt told the local newspaper Bild am Sonntag on Sunday that foreign powers wanted to subdue Germany politically and socially by stirring up fear. "Espionage, sabotage, cyberattacks, or covert operations by foreign powers aimed at destabilising Germany or inflicting direct harm are a constant reality." He did not identify any countries behind the alleged attacks.
Fireworks illuminate Barcelona's Sagrada Família during Pope visit
Pope Leo XIV has described Barcelona's Sagrada Família as a masterpiece of stones, colours and light as he inaugurated its newest - and tallest - tower. The giant Tower of Jesus Christ, completed in February, has brought the church to a soaring height of 172.5m (566ft) - cementing it as the tallest church in the world. His visit to the iconic basilica also marks 100 years since the death of its architect, Antoni Gaudí. Among those attending the service were Spanish royals King Felipe VI and Queen Letizia, as well as Prime Minister Pedro Sánchez. The pope's week-long visit to Spain, which began on Saturday, is the first by a pope in some 15 years.
Ants can be used to make 'tangy' yogurt
Environment Animals Insects Ants can be used to make'tangy' yogurt An old family recipe from Bulgaria goes under the microscope. Breakthroughs, discoveries, and DIY tips sent every weekday. Typically, we humans do our best to keep ants out of our kitchens and away from our food . But an old and almost forgotten recipe harnesses the power of these hard-working insects to make yogurt. The recipe, which was once common across Turkey (or Türkiye) and the Balkans, has been recreated in a study published today in the journal .
Rare cataclysmic exploding star spotted by citizen scientists
Breakthroughs, discoveries, and DIY tips sent every weekday. Two years ago, a team of astronomers requested help from citizen scientists around the world for the Kilonova Seekers Project. Launched in July 2023, the endeavor tasks volunteers with parsing through all-sky survey images captured daily by telescopes on opposite sides of the planet known as the Gravitational-wave Optical Transient Observer (GOTO). Within six months, Kilonova Seekers' over 2,000 volunteers contributed more than 600,000 classifications to researchers, resulting in a total of 20 new discoveries. Now, astronomers have announced the project's first major published find in Astronomy & Astrophysics: a brilliant exploding star observed in near real-time.
ParsiPy: NLP Toolkit for Historical Persian Texts in Python
Farsi, Farhan, Fazel, Parnian, Haghighi, Sepand, Sabouri, Sadra, Goshtasb, Farzaneh, Hajipour, Nadia, Asgari, Ehsaneddin, Sameti, Hossein
The study of historical languages presents unique challenges due to their complex orthographic systems, fragmentary textual evidence, and the absence of standardized digital representations of text in those languages. Tackling these challenges needs special NLP digital tools to handle phonetic transcriptions and analyze ancient texts. This work introduces ParsiPy, an NLP toolkit designed to facilitate the analysis of historical Persian languages by offering modules for tokenization, lemmatization, part-of-speech tagging, phoneme-to-transliteration conversion, and word embedding. We demonstrate the utility of our toolkit through the processing of Parsig (Middle Persian) texts, highlighting its potential for expanding computational methods in the study of historical languages. Through this work, we contribute to computational philology, offering tools that can be adapted for the broader study of ancient texts and their digital preservation.
BgGPT 1.0: Extending English-centric LLMs to other languages
Alexandrov, Anton, Raychev, Veselin, Dimitrov, Dimitar I., Zhang, Ce, Vechev, Martin, Toutanova, Kristina
We present BgGPT-Gemma-2-27B-Instruct and BgGPT-Gemma-2-9B-Instruct: continually pretrained and fine-tuned versions of Google's Gemma-2 models, specifically optimized for Bulgarian language understanding and generation. Leveraging Gemma-2's multilingual capabilities and over 100 billion tokens of Bulgarian and English text data, our models demonstrate strong performance in Bulgarian language tasks, setting a new standard for language-specific AI models. Our approach maintains the robust capabilities of the original Gemma-2 models, ensuring that the English language performance remains intact. To preserve the base model capabilities, we incorporate continual learning strategies based on recent Branch-and-Merge techniques as well as thorough curation and selection of training data. We provide detailed insights into our methodology, including the release of model weights with a commercial-friendly license, enabling broader adoption by researchers, companies, and hobbyists. Further, we establish a comprehensive set of benchmarks based on non-public educational data sources to evaluate models on Bulgarian language tasks as well as safety and chat capabilities. Our findings demonstrate the effectiveness of fine-tuning state-of-the-art models like Gemma 2 to enhance language-specific AI applications while maintaining cross-lingual capabilities.
Are Large Language Models Chameleons?
Geng, Mingmeng, He, Sihong, Trotta, Roberto
Do large language models (LLMs) have their own worldviews and personality tendencies? Simulations in which an LLM was asked to answer subjective questions were conducted more than 1 million times. Comparison of the responses from different LLMs with real data from the European Social Survey (ESS) suggests that the effect of prompts on bias and variability is fundamental, highlighting major cultural, age, and gender biases. Methods for measuring the difference between LLMs and survey data are discussed, such as calculating weighted means and a new proposed measure inspired by Jaccard similarity. We conclude that it is important to analyze the robustness and variability of prompts before using LLMs to model individual decisions or collective behavior, as their imitation abilities are approximate at best.
cantnlp@LT-EDI-2024: Automatic Detection of Anti-LGBTQ+ Hate Speech in Under-resourced Languages
Wong, Sidney G. -J., Durward, Matthew
This paper describes our homophobia/transphobia in social media comments detection system developed as part of the shared task at LT-EDI-2024. We took a transformer-based approach to develop our multiclass classification model for ten language conditions (English, Spanish, Gujarati, Hindi, Kannada, Malayalam, Marathi, Tamil, Tulu, and Telugu). We introduced synthetic and organic instances of script-switched language data during domain adaptation to mirror the linguistic realities of social media language as seen in the labelled training data. Our system ranked second for Gujarati and Telugu with varying levels of performance for other language conditions. The results suggest incorporating elements of paralinguistic behaviour such as script-switching may improve the performance of language detection systems especially in the cases of under-resourced languages conditions.
Challenges and Applications of Automated Extraction of Socio-political Events from Text (CASE 2023): Workshop and Shared Task Report
Hürriyetoğlu, Ali, Tanev, Hristo, Mutlu, Osman, Thapa, Surendrabikram, Tan, Fiona Anting, Yörük, Erdem
We provide a summary of the sixth edition of the CASE workshop that is held in the scope of RANLP 2023. The workshop consists of regular papers, three keynotes, working papers of shared task participants, and shared task overview papers. This workshop series has been bringing together all aspects of event information collection across technical and social science fields. In addition to contributing to the progress in text based event extraction, the workshop provides a space for the organization of a multimodal event information collection task.