Goto

Collaborating Authors

 Africa


GRASP: Grouped Regression with Adaptive Shrinkage Priors

arXiv.org Machine Learning

Group structures are common in regression analysis. They can appear in the form of categorical predictors represented by groups of dummy variables or in the context of additive modeling, where each predictor can be expressed as a set of basis functions forming a group; in applications such as gene expression analysis and financial market modeling, groupings exist naturally in the data. For instance, genes that influence similar traits form groups in gene expression data, while stocks from the same sector form groups in financial data. In these scenarios, group shrinkage plays an important role: when there is insufficient evidence to suggest the significance of predictors within a group, the entire group of predictors is shrunk towards zero. This reduces the noise from individual "spurious predictors", which tend to appear more frequently in high-dimensional settings, and decreases model complexity, thereby reducing the risk of overfitting. 1 Within the Bayesian framework, there has been extensive research focusing on the application of continuous shrinkage priors for linear regression problems involving group predictor variables. Traditional approaches, such as the group lasso[31, 24], the group bridge [16], and the group horseshoe [29] primarily apply shrinkage at the group level and do not consider within-group shrinkage.


Phase transition of \emph{descending} phase retrieval algorithms

arXiv.org Machine Learning

We study theoretical limits of \emph{descending} phase retrieval algorithms. Utilizing \emph{Random duality theory} (RDT) we develop a generic program that allows statistical characterization of various algorithmic performance metrics. Through these we identify the concepts of \emph{parametric manifold} and its \emph{funneling points} as key mathematical objects that govern the underlying algorithms' behavior. An isomorphism between single funneling point manifolds and global convergence of descending algorithms is established. The structure and shape of the parametric manifold as well as its dependence on the sample complexity are studied through both plain and lifted RDT. Emergence of a phase transition is observed. Namely, as sample complexity increases, parametric manifold transitions from a multi to a single funneling point structure. This in return corresponds to a transition from the scenarios where descending algorithms generically fail to the scenarios where they succeed in solving phase retrieval. We also develop and implement a practical algorithmic variant that in a hybrid alternating fashion combines a barrier and a plain gradient descent. Even though the theoretical results are obtained for infinite dimensional scenarios (and consequently non-jittery parametric manifolds), we observe a strong agrement between theoretical and simulated phase transitions predictions for fairly small dimensions on the order of a few hundreds.


HausaNLP at SemEval-2025 Task 11: Hausa Text Emotion Detection

arXiv.org Artificial Intelligence

This paper presents our approach to multi-label emotion detection in Hausa, a low-resource African language, for SemEval Track A. We fine-tuned AfriBERTa, a transformer-based model pre-trained on African languages, to classify Hausa text into six emotions: anger, disgust, fear, joy, sadness, and surprise. Our methodology involved data preprocessing, tokenization, and model fine-tuning using the Hugging Face Trainer API. The system achieved a validation accuracy of 74.00%, with an F1-score of 73.50%, demonstrating the effectiveness of transformer-based models for emotion detection in low-resource languages.


QuranMorph: Morphologically Annotated Quranic Corpus

arXiv.org Artificial Intelligence

We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process utilized lemmas from Qabas, an Arabic lexicographic database linked with 110 lexicons and corpora of 2 million tokens. The part-of-speech tagging was performed using the fine-grained SAMA/Qabas tagset, which encompasses 40 tags. As shown in this paper, this rich lemmatization and POS tagset enabled the QuranMorph corpus to be inter-linked with many linguistic resources. The corpus is open-source and publicly available as part of the SinaLab resources at (https://sina.birzeit.edu/quran)


AI based Content Creation and Product Recommendation Applications in E-commerce: An Ethical overview

arXiv.org Artificial Intelligence

As e-commerce rapidly integrates artificial intelligence for content creation and product recommendations, these technologies offer significant benefits in personalization and efficiency. AI-driven systems automate product descriptions, generate dynamic advertisements, and deliver tailored recommendations based on consumer behavior, as seen in major platforms like Amazon and Shopify. However, the widespread use of AI in e-commerce raises crucial ethical challenges, particularly around data privacy, algorithmic bias, and consumer autonomy. Bias -- whether cultural, gender-based, or socioeconomic -- can be inadvertently embedded in AI models, leading to inequitable product recommendations and reinforcing harmful stereotypes. This paper examines the ethical implications of AI-driven content creation and product recommendations, emphasizing the need for frameworks to ensure fairness, transparency, and need for more established and robust ethical standards. We propose actionable best practices to remove bias and ensure inclusivity, such as conducting regular audits of algorithms, diversifying training data, and incorporating fairness metrics into AI models. Additionally, we discuss frameworks for ethical conformance that focus on safeguarding consumer data privacy, promoting transparency in decision-making processes, and enhancing consumer autonomy. By addressing these issues, we provide guidelines for responsibly utilizing AI in e-commerce applications for content creation and product recommendations, ensuring that these technologies are both effective and ethically sound.


As Israel-Iran war escalates, Ukraine fears 'more losses' to Russia

Al Jazeera

Kyiv, Ukraine – There is a Persian word millions of Ukrainians fear. Shahed – also spelled as Shaheed or Shahid, originally a Quranic term for "martyr" or "witness" – is the name given to the triangular, explosives-laden, Iranian-designed drones that became a harrowing part of daily life and death in wartime Ukraine. These days, they are assembled in the Volga-region Russian city of Yelabuga and undergo constant modifications to make them faster, smarter and deadlier during each air raid that involves hundreds of drones. Their latest Russian versions shot down in Ukraine earlier this month have artificial intelligence modules to better recognise targets, video cameras and two-way radio communication with human operators. "The word'Shahed' will forever be cursed in Ukrainian next to'Moscow' and'Putin'," said Denys Kovalenko, referring to Russian President Vladimir Putin. Kovalenko's face and arms were cut by glass shards after a Shahed exploded above his northern Kyiv neighbourhood in 2023.


Taiwan Is Rushing to Make Its Own Drones Before It's Too Late

WIRED

In the span of just a few years, drones have become instrumental in warfare. Conflicts in Ukraine, Iran, Nagorno-Karabakh, Sudan, and elsewhere have shown how autonomous vehicles have become a quintessential part of modern combat. It's a fact that Taiwan knows all too well. The island nation, fearing imminent invasion from China, has both the need, know-how, and industry necessary to build a robust and advanced drone program. Yet Taiwan, which has set an ambitious target of producing 180,000 drones per year by 2028, is struggling to create this industry from scratch.


Four killed in Kyiv in new Russian aerial attack

BBC News

Four killed in Kyiv in new Russian aerial attack 12 minutes agoShareSaveJaroslav LukivBBC NewsShareSaveUkraine's emergencies service DSNSRescuers from Ukraine's emergencies service DSNS tackle fire in a residential building destroyed in the latest Russian attack on Kyiv At least four people have been killed in an overnight Russian missile and drone attack on Ukraine's capital Kyiv, the interior minister says. In a post on social media, Ihor Klymenko says residential areas, hospitals and sports infrastructure were hit. "An entire section of a residential high-rise building was destroyed" in the worst-hit Shevchenkivskyi district, he says, adding that some people are trapped under the rubble. In the Kyiv region, a woman was killed and another two people injured in the Russian aerial attack, regional head Mykola Kalashnyk says. The Russian military has not commented on the issue.


On Path to Multimodal Historical Reasoning: HistBench and HistAgent

arXiv.org Artificial Intelligence

Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored. Historical reasoning poses unique challenges for AI, involving multimodal source interpretation, temporal inference, and cross-linguistic analysis. While general-purpose agents perform well on many existing benchmarks, they lack the domain-specific expertise required to engage with historical materials and questions. To address this gap, we introduce HistBench, a new benchmark of 414 high-quality questions designed to evaluate AI's capacity for historical reasoning and authored by more than 40 expert contributors. The tasks span a wide range of historical problems-from factual retrieval based on primary sources to interpretive analysis of manuscripts and images, to interdisciplinary challenges involving archaeology, linguistics, or cultural history. Furthermore, the benchmark dataset spans 29 ancient and modern languages and covers a wide range of historical periods and world regions. Finding the poor performance of LLMs and other agents on HistBench, we further present HistAgent, a history-specific agent equipped with carefully designed tools for OCR, translation, archival search, and image understanding in History. On HistBench, HistAgent based on GPT-4o achieves an accuracy of 27.54% pass@1 and 36.47% pass@2, significantly outperforming LLMs with online search and generalist agents, including GPT-4o (18.60%), DeepSeek-R1(14.49%) and Open Deep Research-smolagents(20.29% pass@1 and 25.12% pass@2). These results highlight the limitations of existing LLMs and generalist agents and demonstrate the advantages of HistAgent for historical reasoning.


The Role of Explanation Styles and Perceived Accuracy on Decision Making in Predictive Process Monitoring

arXiv.org Artificial Intelligence

Predictive Process Monitoring (PPM) often uses deep learning models to predict the future behavior of ongoing processes, such as predicting process outcomes. While these models achieve high accuracy, their lack of interpretability undermines user trust and adoption. Explainable AI (XAI) aims to address this challenge by providing the reasoning behind the predictions. However, current evaluations of XAI in PPM focus primarily on functional metrics (such as fidelity), overlooking user-centered aspects such as their effect on task performance and decision-making. This study investigates the effects of explanation styles (feature importance, rule-based, and counterfactual) and perceived AI accuracy (low or high) on decision-making in PPM. We conducted a decision-making experiment, where users were presented with the AI predictions, perceived accuracy levels, and explanations of different styles. Users' decisions were measured both before and after receiving explanations, allowing the assessment of objective metrics (Task Performance and Agreement) and subjective metrics (Decision Confidence). Our findings show that perceived accuracy and explanation style have a significant effect.