Africa
Multimodality Representation Learning: A Survey on Evolution, Pretraining and Its Applications
Manzoor, Muhammad Arslan, Albarri, Sarah, Xian, Ziting, Meng, Zaiqiao, Nakov, Preslav, Liang, Shangsong
Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA), Natural Language for Visual Reasoning (NLVR), and Vision Language Retrieval (VLR). Among these applications, cross-modal interaction and complementary information from different modalities are crucial for advanced models to perform any multimodal task, e.g., understand, recognize, retrieve, or generate optimally. Researchers have proposed diverse methods to address these tasks. The different variants of transformer-based architectures performed extraordinarily on multiple modalities. This survey presents the comprehensive literature on the evolution and enhancement of deep learning multimodal architectures to deal with textual, visual and audio features for diverse cross-modal and modern multimodal tasks. This study summarizes the (i) recent task-specific deep learning methodologies, (ii) the pretraining types and multimodal pretraining objectives, (iii) from state-of-the-art pretrained multimodal approaches to unifying architectures, and (iv) multimodal task categories and possible future improvements that can be devised for better multimodal learning. Moreover, we prepare a dataset section for new researchers that covers most of the benchmarks for pretraining and finetuning. Finally, major challenges, gaps, and potential research topics are explored. A constantly-updated paperlist related to our survey is maintained at https://github.com/marslanm/multimodality-representation-learning.
Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models
Shao, Zhihong, Gong, Yeyun, Shen, Yelong, Huang, Minlie, Duan, Nan, Chen, Weizhu
Large language models can perform various reasoning tasks by using chain-of-thought prompting, which guides them to find answers through step-by-step demonstrations. However, the quality of the prompts depends on the demonstrations given to the models, and creating many of them by hand is costly. We introduce Synthetic prompting, a method that leverages a few handcrafted examples to prompt the model to generate more examples by itself, and selects effective demonstrations to elicit better reasoning. Our method alternates between a backward and forward process to generate new examples. The backward process generates a question that match a sampled reasoning chain, so that the question is solvable and clear. The forward process produces a more detailed reasoning chain for the question, improving the quality of the example. We evaluate our method on numerical, symbolic, and algorithmic reasoning tasks, and show that it outperforms existing prompting techniques.
Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization
Haas, Lukas, Alberti, Silas, Skreta, Michal
By understanding the hidden locational clues in images, entirely new approaches of analyzing the natural and built environment are being opened up with profound implications for a number of fields, ranging from the recognition of weather, season, and climate patterns to rural and urban scene understanding, and improvements in navigation and self-driving car technology. Since the beginning of 2022, image geolocalization has additionally garnered extensive media coverage for becoming an immediate priority of investigative journalists and open source intelligence (OSINT) researchers in their attempt to verify information and to document war atrocities in Ukraine, extracting geolocational information from social media content. Despite high academic and public interest, image geolocalization remains an extremely challenging problem. This is because training datasets are geographically sparse, often limited to specific countries, and biased towards urban or rural scenes. The task is further complicated by the fact that geolocalization requires reasoning on multiple levels of geographic granularity (e.g.
Inference of Partial Colexifications from Multilingual Wordlists
The past years have seen a drastic rise in studies devoted to the investigation of colexification patterns in individual languages families in particular and the languages of the world in specific. Specifically computational studies have profited from the fact that colexification as a scientific construct is easy to operationalize, enabling scholars to infer colexification patterns for large collections of cross-linguistic data. Studies devoted to partial colexifications -- colexification patterns that do not involve entire words, but rather various parts of words--, however, have been rarely conducted so far. This is not surprising, since partial colexifications are less easy to deal with in computational approaches and may easily suffer from all kinds of noise resulting from false positive matches. In order to address this problem, this study proposes new approaches to the handling of partial colexifications by (1) proposing new models with which partial colexification patterns can be represented, (2) developing new efficient methods and workflows which help to infer various types of partial colexification patterns from multilingual wordlists, and (3) illustrating how inferred patterns of partial colexifications can be computationally analyzed and interactively visualized.
Exploring Semantic Perturbations on Grover
Kulkarni, Pranav, Ji, Ziqing, Xu, Yan, Neskovic, Marko, Nolan, Kevin
With news and information being as easy to access as they currently are, it is more important than ever to ensure that people are not mislead by what they read. Recently, the rise of neural fake news (AI-generated fake news) and its demonstrated effectiveness at fooling humans has prompted the development of models to detect it. One such model is the Grover model, which can both detect neural fake news to prevent it, and generate it to demonstrate how a model could be misused to fool human readers. In this work we explore the Grover model's fake news detection capabilities by performing targeted attacks through perturbations on input news articles. Through this we test Grover's resilience to these adversarial attacks and expose some potential vulnerabilities which should be addressed in further iterations to ensure it can detect all types of fake news accurately.
Experimental observation on a low-rank tensor model for eigenvalue problems
Neural networks-based machine learning methods are rapidly developed for various numerical problems, such as physics-informed neural networks (PINNs) [10, 11, 12], the deep Ritz method [2], and the deep Galerkin method [15]. One of the advantages of these approaches is that they show the possibility for solving high-dimensional problems. In [2], the deep learning techniques as well as the Monte-Carlo integration are used to solve eigenvalue problems, which provides a feasible strategy for high-dimensional cases. For the same eigenvalue problems, [16] applies a neural network-based low-rank tensor model, i.e. the tensor neural network (TNN), with a quadrature scheme to perform efficient numerical integration, and thus it achieves a much better result than [2]. Furthermore, [17] employs the TNN to solve the manybody Schrödinger equation, which emerges the practical value of such low-rank approximation method.
Visually Grounded Keyword Detection and Localisation for Low-Resource Languages
This study investigates the use of Visually Grounded Speech (VGS) models for keyword localisation in speech. The study focusses on two main research questions: (1) Is keyword localisation possible with VGS models and (2) Can keyword localisation be done cross-lingually in a real low-resource setting? Four methods for localisation are proposed and evaluated on an English dataset, with the best-performing method achieving an accuracy of 57%. A new dataset containing spoken captions in Yoruba language is also collected and released for cross-lingual keyword localisation. The cross-lingual model obtains a precision of 16% in actual keyword localisation and this performance can be improved by initialising from a model pretrained on English data. The study presents a detailed analysis of the model's success and failure modes and highlights the challenges of using VGS models for keyword localisation in low-resource settings.
Will an AI be the first to discover alien life?
The Robert C. Byrd Green Bank Telescope in West Virginia is one of several helping to search for alien civilizations.Credit: Jim West/Alamy From the hills of West Virginia to the flats of rural Australia, some of the world's largest telescopes are listening for signals from distant alien civilizations. The search for extraterrestrial intelligence, known as SETI, is an effort to find artificial-looking electromagnetic radiation that might have come from a technologically advanced civilization in a far-away solar system. A study published today1 describes one of several efforts to use machine learning, a subset of artificial intelligence (AI), to help astronomers sift quickly through the reams of data such searches yield. As AI reshapes many scientific fields, what promise does it hold for the search for life beyond Earth? "It is a new era for SETI research that is opening up thanks to machine learning technology," says Franck Marchis, a planetary astronomer at the SETI Institute in Mountain View, California.
Universal Topological Regularities of Syntactic Structures: Decoupling Efficiency from Optimization
Martín, Fermín Moscoso del Prado
Human syntactic structures are usually represented as graphs. Much research has focused on the mapping between such graphs and linguistic sequences, but less attention has been paid to the shapes of the graphs themselves: their topologies. This study investigates how the topologies of syntactic graphs reveal traces of the processes that led to their emergence. I report a new universal regularity in syntactic structures: Their topology is communicatively efficient above chance. The pattern holds, without exception, for all 124 languages studied, across linguistic families and modalities (spoken, written, and signed). This pattern can arise from a process optimizing for communicative efficiency or, alternatively, by construction, as a by-effect of a sublinear preferential attachment process reflecting language production mechanisms known from psycholinguistics. This dual explanation shows how communicative efficiency, per se, does not require optimization. Among the two options, efficiency without optimization offers the better explanation for the new pattern.
Deep learning-based lung segmentation and automatic regional template in chest X-ray images for pediatric tuberculosis
Capellán-Martín, Daniel, Gómez-Valverde, Juan J., Sanchez-Jacob, Ramon, Bermejo-Peláez, David, García-Delgado, Lara, López-Varela, Elisa, Ledesma-Carbayo, Maria J.
Tuberculosis (TB) is still considered a leading cause of death and a substantial threat to global child health. Both TB infection and disease are curable using antibiotics. However, most children who die of TB are never diagnosed or treated. In clinical practice, experienced physicians assess TB by examining chest X-rays (CXR). Pediatric CXR has specific challenges compared to adult CXR, which makes TB diagnosis in children more difficult. Computer-aided diagnosis systems supported by Artificial Intelligence have shown performance comparable to experienced radiologist TB readings, which could ease mass TB screening and reduce clinical burden. We propose a multi-view deep learning-based solution which, by following a proposed template, aims to automatically regionalize and extract lung and mediastinal regions of interest from pediatric CXR images where key TB findings may be present. Experimental results have shown accurate region extraction, which can be used for further analysis to confirm TB finding presence and severity assessment.