Goto

Collaborating Authors

 Africa


ALIGN-MLM: Word Embedding Alignment is Crucial for Multilingual Pre-training

arXiv.org Artificial Intelligence

Multilingual pre-trained models exhibit zero-shot cross-lingual transfer, where a model fine-tuned on a source language achieves surprisingly good performance on a target language. While studies have attempted to understand transfer, they focus only on MLM, and the large number of differences between natural languages makes it hard to disentangle the importance of different properties. In this work, we specifically highlight the importance of word embedding alignment by proposing a pre-training objective (ALIGN-MLM) whose auxiliary loss guides similar words in different languages to have similar word embeddings. ALIGN-MLM either outperforms or matches three widely adopted objectives (MLM, XLM, DICT-MLM) when we evaluate transfer between pairs of natural languages and their counterparts created by systematically modifying specific properties like the script. In particular, ALIGN-MLM outperforms XLM and MLM by 35 and 30 F1 points on POS-tagging for transfer between languages that differ both in their script and word order (left-to-right v.s. right-to-left). We also show a strong correlation between alignment and transfer for all objectives (e.g., rho=0.727 for XNLI), which together with ALIGN-MLM's strong performance calls for explicitly aligning word embeddings for multilingual models.


Premonition Net, A Multi-Timeline Transformer Network Architecture Towards Strawberry Tabletop Yield Forecasting

arXiv.org Artificial Intelligence

Abstract--Yield forecasting is a critical first step necessary for yield optimisation, with important consequences for the broader food supply chain, procurement, price-negotiation, logistics, and supply. However yield forecasting is notoriously difficult, and oft-inaccurate. Premonition Net is a multi-timeline, time sequence ingesting approach towards processing the past, the present, and premonitions of the future. We show how this structure combined with transformers attains critical yield forecasting proficiency towards improving food security, lowering prices, and reducing waste. We find data availability to be a continued difficulty however using our premonition network and our own collected data we attain yield forecasts 3 weeks ahead with a a testing set RMSE loss of 0.08 across our latest season.


Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token

arXiv.org Artificial Intelligence

The pre-training of masked language models (MLMs) consumes massive computation to achieve good results on downstream NLP tasks, resulting in a large carbon footprint. In the vanilla MLM, the virtual tokens, [MASK]s, act as placeholders and gather the contextualized information from unmasked tokens to restore the corrupted information. It raises the question of whether we can append [MASK]s at a later layer, to reduce the sequence length for earlier layers and make the pre-training more efficient. We show: (1) [MASK]s can indeed be appended at a later layer, being disentangled from the word embedding; (2) The gathering of contextualized information from unmasked tokens can be conducted with a few layers. By further increasing the masking rate from 15% to 50%, we can pre-train RoBERTa-base and RoBERTa-large from scratch with only 78% and 68% of the original computational budget without any degradation on the GLUE benchmark. When pre-training with the original budget, our method outperforms RoBERTa for 6 out of 8 GLUE tasks, on average by 0.4%.


DuReader_retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine

arXiv.org Artificial Intelligence

In this paper, we present DuReader_retrieval, a large-scale Chinese dataset for passage retrieval. DuReader_retrieval contains more than 90K queries and over 8M unique passages from a commercial search engine. To alleviate the shortcomings of other datasets and ensure the quality of our benchmark, we (1) reduce the false negatives in development and test sets by manually annotating results pooled from multiple retrievers, and (2) remove the training queries that are semantically similar to the development and testing queries. Additionally, we provide two out-of-domain testing sets for cross-domain evaluation, as well as a set of human translated queries for for cross-lingual retrieval evaluation. The experiments demonstrate that DuReader_retrieval is challenging and a number of problems remain unsolved, such as the salient phrase mismatch and the syntactic mismatch between queries and paragraphs. These experiments also show that dense retrievers do not generalize well across domains, and cross-lingual retrieval is essentially challenging. DuReader_retrieval is publicly available at https://github.com/baidu/DuReader/tree/master/DuReader-Retrieval.


Predicting Treatment Adherence of Tuberculosis Patients at Scale

arXiv.org Artificial Intelligence

Tuberculosis (TB), an infectious bacterial disease, is a significant cause of death, especially in low-income countries, with an estimated ten million new cases reported globally in $2020$. While TB is treatable, non-adherence to the medication regimen is a significant cause of morbidity and mortality. Thus, proactively identifying patients at risk of dropping off their medication regimen enables corrective measures to mitigate adverse outcomes. Using a proxy measure of extreme non-adherence and a dataset of nearly $700,000$ patients from four states in India, we formulate and solve the machine learning (ML) problem of early prediction of non-adherence based on a custom rank-based metric. We train ML models and evaluate against baselines, achieving a $\sim 100\%$ lift over rule-based baselines and $\sim 214\%$ over a random classifier, taking into account country-wide large-scale future deployment. We deal with various issues in the process, including data quality, high-cardinality categorical data, low target prevalence, distribution shift, variation across cohorts, algorithmic fairness, and the need for robustness and explainability. Our findings indicate that risk stratification of non-adherent patients is a viable, deployable-at-scale ML solution. As the official AI partner of India's Central TB Division, we are working on multiple city and state-level pilots with the goal of pan-India deployment.



'I lie in the bath, imagining that I am wandering the Rialto in Venice': my obsession with Duolingo

The Guardian

This morning, before checking in on my young son or making a coffee, I opened the Duolingo app on my phone and translated "They love smelling meat" into Italian. I've been starting my days like this for a few months now: wake up, wash face, grapple with the gerund. I usually spend between 10 and 20 minutes on it while the kettle boils or I load CBeebies or write some emails. Duolingo is a language learning app and pretty simple to use. After you've chosen which language you want to learn, you are presented with about 100 skill-sets divided by scenario or grammar (grocery shopping, the future tense and so on).


Huawei Calls for Network Evolution at COP27 to Enable Green Development

#artificialintelligence

A Huawei executive said Thursday information and communications technologies, or ICT, will enable the digitalization of industry, spark innovation and make other industries green. The remarks were made at a session organized by the Global Innovation Hub (UGIH) of the United Nations Framework Convention on Climate Change (UNFCCC) at the ongoing 27th Conference of the Parties, or COP27, in Sharm El-Sheikh of Egypt. Referring to what is known as the "enabling effect", Philippe Wang, Huawei's Executive Vice President for the Northern Africa region, said ICT is "making other industries greener". "5G, Artificial Intelligence, data analytics, cloud computing – all these things will improve industrial processes in a way that cuts energy use, and lowers carbon emissions," he said. According to Philippe Wang, in the same way that ICT enables a smart streetlight to turn itself off when no one is around, 5G wireless base stations can automatically shut down when there is no data traffic, which saves energy.


The AI Image Generator: The Limits of the Algorithm and Human Biases

#artificialintelligence

Over the past few years, these machine learning systems have been tweaked and refined, undergoing multiple iterations to find their present popularity with the everyday internet user. These image generators--DALL-E and Midjourney arguably the most prominent--generate imagery from a variety of text prompts, for instance allowing people to create conceptual renditions of architectures of the future, present, and past. But as we exist in a digital landscape filled with human biases--navigating these image generators requires careful reflection. Midjourney is a particularly interesting Artificial Intelligence tool, proving popular amongst artists and designers alike for its painting-like, imaginative images created from sometimes very minimal text prompts. But the results fed back using this tool also raise complicated questions surrounding image-making and design, questions brought to the forefront when using prompts like "African architecture" to produce images.


Towards Robust Numerical Question Answering: Diagnosing Numerical Capabilities of NLP Systems

arXiv.org Artificial Intelligence

Numerical Question Answering is the task of answering questions that require numerical capabilities. Previous works introduce general adversarial attacks to Numerical Question Answering, while not systematically exploring numerical capabilities specific to the topic. In this paper, we propose to conduct numerical capability diagnosis on a series of Numerical Question Answering systems and datasets. A series of numerical capabilities are highlighted, and corresponding dataset perturbations are designed. Empirical results indicate that existing systems are severely challenged by these perturbations. E.g., Graph2Tree experienced a 53.83% absolute accuracy drop against the ``Extra'' perturbation on ASDiv-a, and BART experienced 13.80% accuracy drop against the ``Language'' perturbation on the numerical subset of DROP. As a counteracting approach, we also investigate the effectiveness of applying perturbations as data augmentation to relieve systems' lack of robust numerical capabilities. With experiment analysis and empirical studies, it is demonstrated that Numerical Question Answering with robust numerical capabilities is still to a large extent an open question. We discuss future directions of Numerical Question Answering and summarize guidelines on future dataset collection and system design.