Goto

Collaborating Authors

 Africa


Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

arXiv.org Artificial Intelligence

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of tasks that enables a large language model (LLM) to respond to instructions. Instruction fine-tuning (IFT) requires specifically constructed and annotated datasets. However, existing datasets are almost all in the English language. In this work, our primary goal is to bridge the language gap by building a human-curated instruction-following dataset spanning 65 languages. We worked with fluent speakers of languages from around the world to collect natural instances of instructions and completions. Furthermore, we create the most extensive multilingual collection to date, comprising 513 million instances through templating and translating existing datasets across 114 languages. In total, we contribute four key resources: we develop and open-source the Aya Annotation Platform, the Aya Dataset, the Aya Collection, and the Aya Evaluation Suite. The Aya initiative also serves as a valuable case study in participatory research, involving collaborators from 119 countries. We see this as a valuable framework for future research collaborations that aim to bridge gaps in resources.


On the Out-Of-Distribution Generalization of Multimodal Large Language Models

arXiv.org Artificial Intelligence

We investigate the generalization boundaries of current Multimodal Large Language Models (MLLMs) via comprehensive evaluation under out-of-distribution scenarios and domain-specific tasks. We evaluate their zero-shot generalization across synthetic images, real-world distributional shifts, and specialized datasets like medical and molecular imagery. Empirical results indicate that MLLMs struggle with generalization beyond common training domains, limiting their direct application without adaptation. To understand the cause of unreliable performance, we analyze three hypotheses: semantic misinterpretation, visual feature extraction insufficiency, and mapping deficiency. Results identify mapping deficiency as the primary hurdle. To address this problem, we show that in-context learning (ICL) can significantly enhance MLLMs' generalization, opening new avenues for overcoming generalization barriers. We further explore the robustness of ICL under distribution shifts and show its vulnerability to domain shifts, label shifts, and spurious correlation shifts between in-context examples and test data.


Incorporating Taylor Series and Recursive Structure in Neural Networks for Time Series Prediction

arXiv.org Artificial Intelligence

Time series analysis is relevant in various disciplines such as physics, biology, chemistry, and Time series analysis plays a pivotal role in extracting valuable finance. In this paper, we present a novel neural insights from sequential data, uncovering patterns, network architecture that integrates elements trends, and underlying structures that drive temporal dynamics from ResNet structures, while introducing the innovative (Zhang, 2003; Tang et al., 1991). The ubiquity incorporation of the Taylor series framework. of time series data across diverse domains, including finance, This approach demonstrates notable enhancements healthcare, and environmental science, underscores in test accuracy across many of the the critical need for accurate and efficient analytical methods baseline datasets investigated.


Finding hardness reductions automatically using SAT solvers

arXiv.org Artificial Intelligence

In this article, we show that the completion problem, i.e. the decision problem whether a partial structure can be completed to a full structure, is NP-complete for many combinatorial structures. While the gadgets for most reductions in literature are found by hand, we present an algorithm to construct gadgets in a fully automated way. Using our framework which is based on SAT, we present the first thorough study of the completion problem on sign mappings with forbidden substructures by classifying thousands of structures for which the completion problem is NP-complete. Our list in particular includes interior triple systems, which were introduced by Knuth towards an axiomatization of planar point configurations. Last but not least, we give an infinite family of structures generalizing interior triple system to higher dimensions for which the completion problem is NP-complete.


Continual Learning on Graphs: A Survey

arXiv.org Artificial Intelligence

Recently, continual graph learning has been increasingly adopted for diverse graph-structured data processing tasks in non-stationary environments. Despite its promising learning capability, current studies on continual graph learning mainly focus on mitigating the catastrophic forgetting problem while ignoring continuous performance improvement. To bridge this gap, this article aims to provide a comprehensive survey of recent efforts on continual graph learning. Specifically, we introduce a new taxonomy of continual graph learning from the perspective of overcoming catastrophic forgetting. Moreover, we systematically analyze the challenges of applying these continual graph learning methods in improving performance continuously and then discuss the possible solutions. Finally, we present open issues and future directions pertaining to the development of continual graph learning and discuss how they impact continuous performance improvement.


Probabilistic Forecasting of Irregular Time Series via Conditional Flows

arXiv.org Artificial Intelligence

Probabilistic forecasting of irregularly sampled multivariate time series with missing values is an important problem in many fields, including health care, astronomy, and climate. State-of-the-art methods for the task estimate only marginal distributions of observations in single channels and at single timepoints, assuming a fixed-shape parametric distribution. In this work, we propose a novel model, ProFITi, for probabilistic forecasting of irregularly sampled time series with missing values using conditional normalizing flows. The model learns joint distributions over the future values of the time series conditioned on past observations and queried channels and times, without assuming any fixed shape of the underlying distribution. As model components, we introduce a novel invertible triangular attention layer and an invertible non-linear activation function on and onto the whole real line. We conduct extensive experiments on four datasets and demonstrate that the proposed model provides $4$ times higher likelihood over the previously best model.


Towards participatory multi-modeling for policy support across domains and scales: a systematic procedure for integral multi-model design

arXiv.org Artificial Intelligence

Policymaking for complex challenges such as pandemics necessitates the consideration of intricate implications across multiple domains and scales. Computational models can support policymaking, but a single model is often insufficient for such multidomain and scale challenges. Multi-models comprising several interacting computational models at different scales or relying on different modeling paradigms offer a potential solution. Such multi-models can be assembled from existing computational models (i.e., integrated modeling) or be designed conceptually as a whole before their computational implementation (i.e., integral modeling). Integral modeling is particularly valuable for novel policy problems, such as those faced in the early stages of a pandemic, where relevant models may be unavailable or lack standard documentation. Designing such multi-models through an integral approach is, however, a complex task requiring the collaboration of modelers and experts from various domains. In this collaborative effort, modelers must precisely define the domain knowledge needed from experts and establish a systematic procedure for translating such knowledge into a multi-model. Yet, these requirements and systematic procedures are currently lacking for multi-models that are both multiscale and multi-paradigm. We address this challenge by introducing a procedure for developing multi-models with an integral approach based on clearly defined domain knowledge requirements derived from literature. We illustrate this procedure using the case of school closure policies in the Netherlands during the COVID-19 pandemic, revealing their potential implications in the short and long term and across the healthcare and educational domains. The requirements and procedure provided in this article advance the application of integral multi-modeling for policy support in multiscale and multidomain contexts.


Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey

arXiv.org Artificial Intelligence

Knowledge Graphs (KGs) play a pivotal role in advancing various AI applications, with the semantic web community's exploration into multi-modal dimensions unlocking new avenues for innovation. In this survey, we carefully review over 300 articles, focusing on KG-aware research in two principal aspects: KG-driven Multi-Modal (KG4MM) learning, where KGs support multi-modal tasks, and Multi-Modal Knowledge Graph (MM4KG), which extends KG studies into the MMKG realm. We begin by defining KGs and MMKGs, then explore their construction progress. Our review includes two primary task categories: KG-aware multi-modal learning tasks, such as Image Classification and Visual Question Answering, and intrinsic MMKG tasks like Multi-modal Knowledge Graph Completion and Entity Alignment, highlighting specific research trajectories. For most of these tasks, we provide definitions, evaluation benchmarks, and additionally outline essential insights for conducting relevant research. Finally, we discuss current challenges and identify emerging trends, such as progress in Large Language Modeling and Multi-modal Pre-training strategies. This survey aims to serve as a comprehensive reference for researchers already involved in or considering delving into KG and multi-modal learning research, offering insights into the evolving landscape of MMKG research and supporting future work.


Long-term monitoring of bird flocks in the wild – interview with Kshitiz

AIHub

In work presented at the 32nd International Joint Conference on Artificial Intelligence (IJCAI 2023), Kshitiz, Sonu Shreshtha, Ramy Mounir, Mayank Vatsa, Richa Singh, Saket Anand, Sudeep Sarkar and Sevaram Mali Parihar investigate using computer vision techniques to monitor large flocks of birds. In this interview, Kshitiz tells us more about this research. In our work, Long-term Monitoring of Bird Flocks in the Wild, published in IJCAI 2023, we delve into developing and applying computer vision techniques and datasets tailored for non-invasive monitoring and analysis of migratory bird flocks in their natural habitats. The aim is to understand the behavior and ecology of migratory birds through automated video analysis with minimal human intervention, thereby bolstering conservation initiatives. The core technical challenges associated with wildlife monitoring arise from the uncontrolled, outdoor nature of the imagery (both images and videos) capturing large flocks of migratory birds over several months.


Improving Token-Based World Models with Parallel Observation Prediction

arXiv.org Artificial Intelligence

Motivated by the success of Transformers when applied to sequences of discrete symbols, token-based world models (TBWMs) were recently proposed as sample-efficient methods. In TBWMs, the world model consumes agent experience as a language-like sequence of tokens, where each observation constitutes a sub-sequence. However, during imagination, the sequential token-by-token generation of next observations results in a severe bottleneck, leading to long training times, poor GPU utilization, and limited representations. To resolve this bottleneck, we devise a novel Parallel Observation Prediction (POP) mechanism. POP augments a Retentive Network (RetNet) with a novel forward mode tailored to our reinforcement learning setting. We incorporate POP in a novel TBWM agent named REM (Retentive Environment Model), showcasing a 15.4x faster imagination compared to prior TBWMs. REM attains superhuman performance on 12 out of 26 games of the Atari 100K benchmark, while training in less than 12 hours. Our code is available at \url{https://github.com/leor-c/REM}.