Oceania
Speaker- and Age-Invariant Training for Child Acoustic Modeling Using Adversarial Multi-Task Learning
Shahin, Mostafa, Ahmed, Beena, Epps, Julien
One of the major challenges in acoustic modelling of child speech is the rapid changes that occur in the children's articulators as they grow up, their differing growth rates and the subsequent high variability in the same age group. These high acoustic variations along with the scarcity of child speech corpora have impeded the development of a reliable speech recognition system for children. In this paper, a speaker- and age-invariant training approach based on adversarial multi-task learning is proposed. The system consists of one generator shared network that learns to generate speaker- and age-invariant features connected to three discrimination networks, for phoneme, age, and speaker. The generator network is trained to minimize the phoneme-discrimination loss and maximize the speaker- and age-discrimination losses in an adversarial multi-task learning fashion. The generator network is a Time Delay Neural Network (TDNN) architecture while the three discriminators are feed-forward networks. The system was applied to the OGI speech corpora and achieved a 13% reduction in the WER of the ASR.
Confidence Intervals for Unobserved Events
Consider a finite sample from an unknown distribution over a countable alphabet. Unobserved events are alphabet symbols which do not appear in the sample. Estimating the probabilities of unobserved events is a basic problem in statistics and related fields, which was extensively studied in the context of point estimation. In this work we introduce a novel interval estimation scheme for unobserved events. Our proposed framework applies selective inference, as we construct confidence intervals (CIs) for the desired set of parameters. Interestingly, we show that obtained CIs are dimension-free, as they do not grow with the alphabet size. Further, we show that these CIs are (almost) tight, in the sense that they cannot be further improved without violating the prescribed coverage rate. We demonstrate the performance of our proposed scheme in synthetic and real-world experiments, showing a significant improvement over the alternatives. Finally, we apply our proposed scheme to large alphabet modeling. We introduce a novel simultaneous CI scheme for large alphabet distributions which outperforms currently known methods while maintaining the prescribed coverage rate.
HumSet: Dataset of Multilingual Information Extraction and Classification for Humanitarian Crisis Response
Fekih, Selim, Tamagnone, Nicolรฒ, Minixhofer, Benjamin, Shrestha, Ranjan, Contla, Ximena, Oglethorpe, Ewan, Rekabsaz, Navid
Timely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data - a process that can highly benefit from expert-assisted NLP systems trained on validated and annotated data in the humanitarian response domain. To enable creation of such NLP systems, we introduce and release HumSet, a novel and rich multilingual dataset of humanitarian response documents annotated by experts in the humanitarian response community. The dataset provides documents in three languages (English, French, Spanish) and covers a variety of humanitarian crises from 2018 to 2021 across the globe. For each document, HUMSET provides selected snippets (entries) as well as assigned classes to each entry annotated using common humanitarian information analysis frameworks. HUMSET also provides novel and challenging entry extraction and multi-label entry classification tasks. In this paper, we take a first step towards approaching these tasks and conduct a set of experiments on Pre-trained Language Models (PLM) to establish strong baselines for future research in this domain. The dataset is available at https://blog.thedeep.io/humset/.
Sparse Horseshoe Estimation via Expectation-Maximisation
Tew, Shu Yu, Schmidt, Daniel F., Makalic, Enes
The horseshoe prior is known to possess many desirable properties for Bayesian estimation of sparse parameter vectors, yet its density function lacks an analytic form. As such, it is challenging to find a closed-form solution for the posterior mode. Conventional horseshoe estimators use the posterior mean to estimate the parameters, but these estimates are not sparse. We propose a novel expectation-maximisation (EM) procedure for computing the MAP estimates of the parameters in the case of the standard linear model. A particular strength of our approach is that the M-step depends only on the form of the prior and it is independent of the form of the likelihood. We introduce several simple modifications of this EM procedure that allow for straightforward extension to generalised linear models. In experiments performed on simulated and real data, our approach performs comparable, or superior to, state-of-the-art sparse estimation methods in terms of statistical performance and computational cost.
Complex Reading Comprehension Through Question Decomposition
Guo, Xiao-Yu, Li, Yuan-Fang, Haffari, Gholamreza
Multi-hop reading comprehension requires not only the ability to reason over raw text but also the ability to combine multiple evidence. We propose a novel learning approach that helps language models better understand difficult multi-hop questions and perform "complex, compositional" reasoning. Our model first learns to decompose each multi-hop question into several sub-questions by a trainable question decomposer. Instead of answering these sub-questions, we directly concatenate them with the original question and context, and leverage a reading comprehension model to predict the answer in a sequence-to-sequence manner. By using the same language model for these two components, our best seperate/unified t5-base variants outperform the baseline by 7.2/6.1 absolute F1 points on a hard subset of DROP dataset.
On the Domain Adaptation and Generalization of Pretrained Language Models: A Survey
Recent advances in NLP are brought by a range of large-scale pretrained language models (PLMs). These PLMs have brought significant performance gains for a range of NLP tasks, circumventing the need to customize complex designs for specific tasks. However, most current work focus on finetuning PLMs on a domain-specific datasets, ignoring the fact that the domain gap can lead to overfitting and even performance drop. Therefore, it is practically important to find an appropriate method to effectively adapt PLMs to a target domain of interest. Recently, a range of methods have been proposed to achieve this purpose. Early surveys on domain adaptation are not suitable for PLMs due to the sophisticated behavior exhibited by PLMs from traditional models trained from scratch and that domain adaptation of PLMs need to be redesigned to take effect. This paper aims to provide a survey on these newly proposed methods and shed light in how to apply traditional machine learning methods to newly evolved and future technologies. By examining the issues of deploying PLMs for downstream tasks, we propose a taxonomy of domain adaptation approaches from a machine learning system view, covering methods for input augmentation, model optimization and personalization. We discuss and compare those methods and suggest promising future research directions.
AAN+: Generalized Average Attention Network for Accelerating Neural Transformer
Zhang, Biao (a:1:{s:5:"en_US";s:23:"University of Edinburgh";}) | Xiong, Deyi | Ge, Yubin | Yao, Junfeng | Yue, Hao | Su, Jinsong
Transformer benefits from the high parallelization of attention networks in fast training, but it still suffers from slow decoding partially due to the linear dependency O(m) of the decoder self-attention on previous target words at inference. In this paper, we propose a generalized average attention network (AAN+) aiming at speeding up decoding by reducing the dependency from O(m) to O(1). We find that the learned self-attention weights in the decoder follow some patterns which can be approximated via a dynamic structure. Based on this insight, we develop AAN+, extending our previously proposed average attention (Zhang et al., 2018a, AAN) to support more general position- and content-based attention patterns. AAN+ only requires to maintain a small constant number of hidden states during decoding, ensuring its O(1) dependency. We apply AAN+ as a drop-in replacement of the decoder selfattention and conduct experiments on machine translation (with diverse language pairs), table-to-text generation and document summarization. With masking tricks and dynamic programming, AAN+ enables Transformer to decode sentences around 20% faster without largely compromising in the training speed and the generation performance. Our results further reveal the importance of the localness (neighboring words) in AAN+ and its capability in modeling long-range dependency.
AI's impact spreads to myriad sectors, from the arts to transport to social media management
After turning heads with her paintings of icons like Queen Elizabeth II and Sir Paul McCartney, artificial intelligence robot Ai-Da made a historic appearance in the UK's House of Lords on October 11, where the sophisticated automaton answered questions as part of a wider inquiry into the relationship between technology and creativity in the modern world. Ai-Da's parliamentary cameo is a reminder that emerging AI technologies have a more diverse range of positive applications than many people might think. As well as its extensive use in the creative sector (which is often heralded as the defining distinction between human and robot), AI also has much to offer in terms of improving safety, efficiency, and the overall user experience in a variety of different industries and areas. Ai-Da might be commanding much of the headlines surrounding AI art at the moment, but she represents the merest tip of the iceberg when it comes to AI's contribution to the creative sector. Indeed, the gaming industry has been one of the driving forces and guinea pigs behind the development of the technology, with the first chess algorithms coming out in the 1950s and IBM's Deep Blue creating quite the storm in 1997 when it beat then-Grandmaster Garry Kasparov.
Google Expands Flood and Wildfire Tracking to More Countries
A gaggle of new AI projects are coming soon from Google, including disaster monitoring tools and a service that uses machine intelligence to generate custom videos. The company announced the array of initiatives at its AI@ event this week. The most practical development: Google is expanding its AI-powered disaster tracking and response systems. The company rolled out a wildfire tracking tool during the apocalyptic 2020 fire season. The tool aims to track wildfire movements in real time using satellite imagery, on-the-ground data, and AI predictions.
How AI can help the public health sector face future crises
Join us on November 9 to learn how to successfully innovate and achieve efficiency by upskilling and scaling citizen developers at the Low-Code/No-Code Summit. From COVID--19 to monkeypox and intermittent polio scares, concerns around public health crises have significantly increased over the past several years. Living in a globally connected world amidst climate change and a growing population has enabled the emergence of more frequent viruses and fostered their spread. A research study last year estimated that the probability of novel disease outbreaks will grow three-fold in the next few decades. Fortunately, there have been significant technological developments that can help minimize the impact of these global health issues.