Oceania
The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages
Akerman, Vesa, Baines, David, Daspit, Damien, Hermjakob, Ulf, Jang, Taeho, Leong, Colin, Martin, Michael, Mathew, Joel, Robie, Jonathan, Schwarting, Marcus
Efficiently and accurately translating a corpus into a low-resource language remains a challenge, regardless of the strategies employed, whether manual, automated, or a combination of the two. Many Christian organizations are dedicated to the task of translating the Holy Bible into languages that lack a modern translation. Bible translation (BT) work is currently underway for over 3000 extremely low resource languages. We introduce the eBible corpus: a dataset containing 1009 translations of portions of the Bible with data in 833 different languages across 75 language families. In addition to a BT benchmarking dataset, we introduce model performance benchmarks built on the No Language Left Behind (NLLB) neural machine translation (NMT) models. Finally, we describe several problems specific to the domain of BT and consider how the established data and model benchmarks might be used for future translation efforts. For a BT task trained with NLLB, Austronesian and Trans-New Guinea language families achieve 35.1 and 31.6 BLEU scores respectively, which spurs future innovations for NMT for low-resource languages in Papua New Guinea.
Affective social anthropomorphic intelligent system
Mamun, Md. Adyelullahil, Abdullah, Hasnat Md., Alam, Md. Golam Rabiul, Hassan, Muhammad Mehedi, Uddin, Md. Zia
Human conversational styles are measured by the sense of humor, personality, and tone of voice. These characteristics have become essential for conversational intelligent virtual assistants. However, most of the state-of-the-art intelligent virtual assistants (IVAs) are failed to interpret the affective semantics of human voices. This research proposes an anthropomorphic intelligent system that can hold a proper human-like conversation with emotion and personality. A voice style transfer method is also proposed to map the attributes of a specific emotion. Initially, the frequency domain data (Mel-Spectrogram) is created by converting the temporal audio wave data, which comprises discrete patterns for audio features such as notes, pitch, rhythm, and melody. A collateral CNN-Transformer-Encoder is used to predict seven different affective states from voice. The voice is also fed parallelly to the deep-speech, an RNN model that generates the text transcription from the spectrogram. Then the transcripted text is transferred to the multi-domain conversation agent using blended skill talk, transformer-based retrieve-and-generate generation strategy, and beam-search decoding, and an appropriate textual response is generated. The system learns an invertible mapping of data to a latent space that can be manipulated and generates a Mel-spectrogram frame based on previous Mel-spectrogram frames to voice synthesize and style transfer. Finally, the waveform is generated using WaveGlow from the spectrogram. The outcomes of the studies we conducted on individual models were auspicious. Furthermore, users who interacted with the system provided positive feedback, demonstrating the system's effectiveness.
Fruit Picker Activity Recognition with Wearable Sensors and Machine Learning
Dabrowski, Joel Janek, Rahman, Ashfaqur
In this paper we present a novel application of detecting fruit picker activities based on time series data generated from wearable sensors. During harvesting, fruit pickers pick fruit into wearable bags and empty these bags into harvesting bins located in the orchard. Once full, these bins are quickly transported to a cooled pack house to improve the shelf life of picked fruits. For farmers and managers, the knowledge of when a picker bag is emptied is important for managing harvesting bins more effectively to minimise the time the picked fruit is left out in the heat (resulting in reduced shelf life). We propose a means to detect these bag-emptying events using human activity recognition with wearable sensors and machine learning methods. We develop a semi-supervised approach to labelling the data. A feature-based machine learning ensemble model and a deep recurrent convolutional neural network are developed and tested on a real-world dataset. When compared, the neural network achieves 86% detection accuracy.
Radar de Parit\'e: An NLP system to measure gender representation in French news stories
Soumah, Valentin-Gabriel, Rao, Prashanth, Eibl, Philipp, Taboada, Maite
We present the Radar de Paritรฉ, an automated Natural Language Processing (NLP) system that measures the proportion of women and men quoted daily in six Canadian French-language media outlets. We outline the system's architecture and detail the challenges we overcame to address French-specific issues, in particular regarding coreference resolution, a new contribution to the NLP literature on French. Our results highlight the underrepresentation of women in news stories, while also illustrating the application of modern NLP methods to measure gender representation and address societal issues. The commonality in most applied NLP research projects is the need to reliably and scalably extract information from unstructured text data. In this paper, we describe one such application: extracting quotes from news stories to quantify gender representation. Gender representation in the media is a long debated topic. From the 1970s, there have been studies into how much women and gender-diverse people are portrayed in news stories, with the general hypothesis that they tend to be underrepresented [1, 2]. There is also research studying how they are represented, i.e., whether sexist or homophobic tropes are present when we discuss women and gender-diverse people [3, 4]. In this work, we tackle one specific aspect of representation: who is quoted and in what proportions. Our starting hypothesis is that we hear less from women than from men in news stories, that is, that men are quoted more often than is to be expected from their proportion in the general population. To fully answer this question, we formulate a quantitative approach, collecting large amounts of representative data and extracting quotes from the unstructured text. This is the goal of the Radar de Paritรฉ. We define quotes as either direct or indirect reproductions of what a person said, and we define that person as a source in news articles. In order to extract quotes, we employ a full NLP pipeline, focusing on parsing to identify speakers, verbs, and quotes, in each news story. We then predict the gender of the speaker (or source), using external genderprediction services.
Low-resource Bilingual Dialect Lexicon Induction with Large Language Models
Artemova, Ekaterina, Plank, Barbara
Bilingual word lexicons are crucial tools for multilingual natural language understanding and machine translation tasks, as they facilitate the mapping of words in one language to their synonyms in another language. To achieve this, numerous papers have explored bilingual lexicon induction (BLI) in high-resource scenarios, using a typical pipeline consisting of two unsupervised steps: bitext mining and word alignment, both of which rely on pre-trained large language models~(LLMs). In this paper, we present an analysis of the BLI pipeline for German and two of its dialects, Bavarian and Alemannic. This setup poses several unique challenges, including the scarcity of resources, the relatedness of the languages, and the lack of standardization in the orthography of dialects. To evaluate the BLI outputs, we analyze them with respect to word frequency and pairwise edit distance. Additionally, we release two evaluation datasets comprising 1,500 bilingual sentence pairs and 1,000 bilingual word pairs. They were manually judged for their semantic similarity for each Bavarian-German and Alemannic-German language pair.
Decadal Temperature Prediction via Chaotic Behavior Tracking
Ren, Jinfu, Liu, Yang, Liu, Jiming
Decadal temperature prediction provides crucial information for quantifying the expected effects of future climate changes and thus informs strategic planning and decision-making in various domains. However, such long-term predictions are extremely challenging, due to the chaotic nature of temperature variations. Moreover, the usefulness of existing simulation-based and machine learning-based methods for this task is limited because initial simulation or prediction errors increase exponentially over time. To address this challenging task, we devise a novel prediction method involving an information tracking mechanism that aims to track and adapt to changes in temperature dynamics during the prediction phase by providing probabilistic feedback on the prediction error of the next step based on the current prediction. We integrate this information tracking mechanism, which can be considered as a model calibrator, into the objective function of our method to obtain the corrections needed to avoid error accumulation. Our results show the ability of our method to accurately predict global land-surface temperatures over a decadal range. Furthermore, we demonstrate that our results are meaningful in a real-world context: the temperatures predicted using our method are consistent with and can be used to explain the well-known teleconnections within and between different continents.
'AI isn't a threat' โ Boris Eldagsen, whose fake photo duped the Sony judges, hits back
Since 52-year-old German artist Boris Eldagsen went public with the fact that he won a Sony world photography award with an AI-generated image, relations between him and the award body have soured. Sony have issued a statement, saying: "We no longer feel we are able to engage in a meaningful and constructive dialogue with him." His website reads: "Sony: Stop saying nonsense!" "I don't know why they behaved like this," he says, speaking to me from Berlin on the morning after the controversy broke. But I have a fair idea: plainly, they feel like they were conned, and had their aesthetic discernment called into question.
The Digital Insider
This article is brought to you by Retail Technology Review: Retail Technology Show 2023: Retail Express to join leading tech innovators. Q&A with Ed Betts, Retail Lead Europe at Retail Express, who considers ahead of the show the role of intelligent merchandising technology and its critical place in the future of retail. The critical role that technology plays in retail is unquestionable. Within the industry we are now talking more and more about smart retail technology which is any technology that is used to improve efficiency and effectiveness of operations, such as Artificial Intelligence (AI). So, in short, the Retail Technology Show on 26-27 April in London is where all the leading vendors get together and Retail Express is excited to be there.
Data Modeler at TAL - Sydney, Australia
TAL has recently completed the acquisition of Westpac Banking Corp's life insurance business. The acquisition allows TAL to strengthen its position as the largest insurer in Australia and will provide expanded opportunities for its business and its people. Importantly, it will also ensure we continue to deliver the best possible outcomes for our customers, partners and stakeholders. The deal also includes an exclusive long-term 20-year strategic alliance through which TAL will provide life insurance products and services to Westpac's existing customers and partners. Every Australian life is different.