AITopics | Alnumay, Yazeed

Collaborating Authors

Alnumay, Yazeed

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

If you are looking for an answer to the question What is Artificial Intelligence? and you only have a minute, then here's the definition the Association for the Advancement of Artificial Intelligence offers on its home page: "the scientific understanding of the mechanisms underlying thought and intelligent behavior and their embodiment in machines."

However, if you are fortunate enough to have more than a minute, then please get ready to embark upon an exciting journey exploring AI (but beware, it could last a lifetime) …

Command R7B Arabic: A Small, Enterprise Focused, Multilingual, and Culturally Aware Arabic LLM

Alnumay, Yazeed, Barbet, Alexandre, Bialas, Anna, Darling, William, Desai, Shaan, Devassy, Joan, Duffy, Kyle, Howe, Stephanie, Lasche, Olivia, Lee, Justin, Shrinivason, Anirudh, Tracey, Jennifer

arXiv.org Artificial IntelligenceMar-18-2025

Building high-quality large language models (LLMs) for enterprise Arabic applications remains challenging due to the limited availability of digitized Arabic data. In this work, we present a data synthesis and refinement strategy to help address this problem, namely, by leveraging synthetic data generation and human-in-the-loop annotation to expand our Arabic training corpus. We further present our iterative post training recipe that is essential to achieving state-of-the-art performance in aligning the model with human preferences, a critical aspect to enterprise use cases. The culmination of this effort is the release of a small, 7B, open-weight model that outperforms similarly sized peers in head-to-head comparisons and on Arabic-focused benchmarks covering cultural knowledge, instruction following, RAG, and contextual faithfulness.

arabic, arxiv, preprint, (16 more...)

arXiv.org Artificial Intelligence

2503.14603

Country:

Europe > Ukraine (0.14)
North America > United States (0.14)
Europe > France (0.14)
Asia > Thailand (0.14)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Large Language Model (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.47)

Add feedback

ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition

Alyahya, Hisham A., Khan, Haidar, Alnumay, Yazeed, Bari, M Saiful, Yener, Bülent

arXiv.org Artificial IntelligenceMar-10-2025

We introduce ZeroSumEval, a dynamic, competition-based, and evolving evaluation framework for Large Language Models (LLMs) that leverages competitive games. ZeroSumEval encompasses a diverse suite of games, including security challenges (Capture the Flag), classic board games (chess), and knowledge tests (MathQuiz). These games are designed to evaluate a range of capabilities such as strategic reasoning, planning, knowledge application, safety, and adaptability. Building upon recent studies that highlight the effectiveness of game-based evaluations for LLMs, ZeroSumEval enhances these approaches by providing a standardized and extensible framework for easily implementing games and leverages DSPy to provide a better abstraction for LLM player strategies.

large language model, machine learning, natural language, (19 more...)

arXiv.org Artificial Intelligence

2503.10673

Country: Asia > Middle East (0.14)

Genre: Research Report (0.84)

Industry: Leisure & Entertainment > Games > Chess (0.49)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Large Language Model (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.47)

Add feedback

When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Alzahrani, Norah, Alyahya, Hisham Abdullah, Alnumay, Yazeed, Alrashed, Sultan, Alsubaie, Shaykhah, Almushaykeh, Yusef, Mirza, Faisal, Alotaibi, Nouf, Altwairesh, Nora, Alowisheq, Areeb, Bari, M Saiful, Khan, Haidar

arXiv.org Artificial IntelligenceFeb-1-2024

Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are taken at face value - we show this is a (potentially costly) mistake. Under existing leaderboards, the relative performance of LLMs is highly sensitive to (often minute) details. We show that for popular multiple choice question benchmarks (e.g. MMLU) minor perturbations to the benchmark, such as changing the order of choices or the method of answer selection, result in changes in rankings up to 8 positions. We explain this phenomenon by conducting systematic experiments over three broad categories of benchmark perturbations and identifying the sources of this behavior. Our analysis results in several best-practice recommendations, including the advantage of a hybrid scoring method for answer selection. Our study highlights the dangers of relying on simple benchmark evaluations and charts the path for more robust evaluation schemes on the existing benchmarks.

answer choice, large language model, machine learning, (19 more...)

arXiv.org Artificial Intelligence

2402.01781

Country: Asia > Middle East > Saudi Arabia (0.28)

Genre: Research Report (1.00)

Industry: Education (0.67)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Large Language Model (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (1.00)

Add feedback