Africa
Syntactic Variation Across the Grammar: Modelling a Complex Adaptive System
While language is a complex adaptive system, most work on syntactic variation observes a few individual constructions in isolation from the rest of the grammar. This means that the grammar, a network which connects thousands of structures at different levels of abstraction, is reduced to a few disconnected variables. This paper quantifies the impact of such reductions by systematically modelling dialectal variation across 49 local populations of English speakers in 16 countries. We perform dialect classification with both an entire grammar as well as with isolated nodes within the grammar in order to characterize the syntactic differences between these dialects. The results show, first, that many individual nodes within the grammar are subject to variation but, in isolation, none perform as well as the grammar as a whole. This indicates that an important part of syntactic variation consists of interactions between different parts of the grammar. Second, the results show that the similarity between dialects depends heavily on the sub-set of the grammar being observed: for example, New Zealand English could be more similar to Australian English in phrasal verbs but at the same time more similar to UK English in dative phrases.
Knowledge Sanitization of Large Language Models
Ishibashi, Yoichi, Shimodaira, Hidetoshi
We explore a knowledge sanitization approach to mitigate the privacy concerns associated with large language models (LLMs). LLMs trained on a large corpus of Web data can memorize and potentially reveal sensitive or confidential information, raising critical security concerns. Our technique fine-tunes these models, prompting them to generate harmless responses such as ``I don't know'' when queried about specific information. Experimental results in a closed-book question-answering task show that our straightforward method not only minimizes particular knowledge leakage but also preserves the overall performance of LLM. These two advantages strengthen the defense against extraction attacks and reduces the emission of harmful content such as hallucinations.
Multimodal Transformers for Wireless Communications: A Case Study in Beam Prediction
Tian, Yu, Zhao, Qiyang, Kherroubi, Zine el abidine, Boukhalfa, Fouzi, Wu, Kebin, Bader, Faouzi
Wireless communications at high-frequency bands with large antenna arrays face challenges in beam management, which can potentially be improved by multimodality sensing information from cameras, LiDAR, radar, and GPS. In this paper, we present a multimodal transformer deep learning framework for sensing-assisted beam prediction. We employ a convolutional neural network to extract the features from a sequence of images, point clouds, and radar raw data sampled over time. At each convolutional layer, we use transformer encoders to learn the hidden relations between feature tokens from different modalities and time instances over abstraction space and produce encoded vectors for the next-level feature extraction. We train the model on a combination of different modalities with supervised learning. We try to enhance the model over imbalanced data by utilizing focal loss and exponential moving average. We also evaluate data processing and augmentation techniques such as image enhancement, segmentation, background filtering, multimodal data flipping, radar signal transformation, and GPS angle calibration. Experimental results show that our solution trained on image and GPS data produces the best distance-based accuracy of predicted beams at 78.44%, with effective generalization to unseen day scenarios near 73% and night scenarios over 84%. This outperforms using other modalities and arbitrary data processing techniques, which demonstrates the effectiveness of transformers with feature fusion in performing radio beam prediction from images and GPS. Furthermore, our solution could be pretrained from large sequences of multimodality wireless data, on fine-tuning for multiple downstream radio network tasks.
L1-aware Multilingual Mispronunciation Detection Framework
Kheir, Yassine El, Chowdhury, Shammur Absar, Ali, Ahmed
The phonological discrepancies between a speaker's native (L1) and the non-native language (L2) serves as a major factor for mispronunciation. This paper introduces a novel multilingual MDD architecture, L1-MultiMDD, enriched with L1-aware speech representation. An end-to-end speech encoder is trained on the input signal and its corresponding reference phoneme sequence. First, an attention mechanism is deployed to align the input audio with the reference phoneme sequence. Afterwards, the L1-L2-speech embedding are extracted from an auxiliary model, pretrained in a multi-task setup identifying L1 and L2 language, and are infused with the primary network. Finally, the L1-MultiMDD is then optimized for a unified multilingual phoneme recognition task using connectionist temporal classification (CTC) loss for the target languages: English, Arabic, and Mandarin. Our experiments demonstrate the effectiveness of the proposed L1-MultiMDD framework on both seen -- L2-ARTIC, LATIC, and AraVoiceL2v2; and unseen -- EpaDB and Speechocean762 datasets. The consistent gains in PER, and false rejection rate (FRR) across all target languages confirm our approach's robustness, efficacy, and generalizability.
Contextual Biasing of Named-Entities with Large Language Models
Sun, Chuanneng, Ahmed, Zeeshan, Ma, Yingyi, Liu, Zhe, Kabela, Lucas, Pang, Yutong, Kalinli, Ozlem
This paper studies contextual biasing with Large Language Models (LLMs), where during second-pass rescoring additional contextual information is provided to a LLM to boost Automatic Speech Recognition (ASR) performance. We propose to leverage prompts for a LLM without fine tuning during rescoring which incorporate a biasing list and few-shot examples to serve as additional information when calculating the score for the hypothesis. In addition to few-shot prompt learning, we propose multi-task training of the LLM to predict both the entity class and the next token. To improve the efficiency for contextual biasing and to avoid exceeding LLMs' maximum sequence lengths, we propose dynamic prompting, where we select the most likely class using the class tag prediction, and only use entities in this class as contexts for next token prediction. Word Error Rate (WER) evaluation is performed on i) an internal calling, messaging, and dictation dataset, and ii) the SLUE-Voxpopuli dataset. Results indicate that biasing lists and few-shot examples can achieve 17.8% and 9.6% relative improvement compared to first pass ASR, and that multi-task training and dynamic prompting can achieve 20.0% and 11.3% relative WER improvement, respectively.
On the different regimes of Stochastic Gradient Descent
Sclocchi, Antonio, Wyart, Matthieu
Modern deep networks are trained with stochastic gradient descent (SGD) whose key parameters are the number of data considered at each step or batch size $B$, and the step size or learning rate $\eta$. For small $B$ and large $\eta$, SGD corresponds to a stochastic evolution of the parameters, whose noise amplitude is governed by the `temperature' $T\equiv \eta/B$. Yet this description is observed to break down for sufficiently large batches $B\geq B^*$, or simplifies to gradient descent (GD) when the temperature is sufficiently small. Understanding where these cross-overs take place remains a central challenge. Here we resolve these questions for a teacher-student perceptron classification model, and show empirically that our key predictions still apply to deep networks. Specifically, we obtain a phase diagram in the $B$-$\eta$ plane that separates three dynamical phases: $\textit{(i)}$ a noise-dominated SGD governed by temperature, $\textit{(ii)}$ a large-first-step-dominated SGD and $\textit{(iii)}$ GD. These different phases also corresponds to different regimes of generalization error. Remarkably, our analysis reveals that the batch size $B^*$ separating regimes $\textit{(i)}$ and $\textit{(ii)}$ scale with the size $P$ of the training set, with an exponent that characterizes the hardness of the classification problem.
Scientists sound alarm as NASA says small chance asteroid 'Bennu' the size of the Empire State Building could smash into earth: 'It would be like unleashing 24 atomic bombs'
NASA has spent seven years trying to prevent Bennu -- an asteroid taller than the Empire State Building and named after ancient Egypt's fiery bird-god -- from crashing cataclysmically into Earth. While Bennu's chances of impact are just 1-in-2,700, more than five times a person's chance of being struck by lightning, NASA's team nevertheless has categorized it as one of the two'most hazardous known asteroids.' In a worst-case scenario, the roughly 510-meter wide, carbon-based behemoth would smash into Earth with 1,200 megatons of energy: 24 times the power of the largest nuclear bomb ever detonated (the Soviet Union's'Tsar Bomba'). If it happens, Bennu's impact would unleash its 1.2 gigaton impact 159 years from this Sunday, on September 24, 2182. While Bennu is nowhere near the size of the dino-killing, six-mile across space rock that hit the Yucatan 66 million years ago, astronomers believe that the asteroid'could cause continental devastation if it became an Earth impactor.'
Do YOU speak chicken? Scientists say you can tell how birds are feeling based on their noises - so can you decipher these clucks?
From clucks to squawks and even'growling', the meanings behind chicken sounds have always been a mystery, even to farmers. Not any more, however, because artificial intelligence (AI) technology from Japan has finally been able to translate them – giving a unique insight into a chicken's wellbeing. Experts trained an AI model with about 100 hours of chicken recordings until it could identify with 80 per cent accuracy if a bird was happy, sad or frightened. The scientists used machine learning (ML), a specific subset of AI which allows systems to learn and come to informed conclusions. Audio clips released by the experts show the wide range of noises that the birds make – but can you identify a chicken's emotion as efficiently as an AI? Humans can look forward to more'meaningful' interactions with chickens thanks to the study results, according to researchers Mother hens are such caring parents that they'feel' their chicks' pain The research was led by Professor Adrian David Cheok at the the University of Tokyo, who is known for his expertise in the area of sex robots. 'It's a cluckin' great leap for science and this is just the beginning,' Professor Cheok said.
Full text: Zelenskyy's speech to the UN General Assembly
Ukrainian President Volodymyr Zelenskyy travelled to New York to address the United Nations General Assembly in person for the first time since Moscow began its full-scale invasion of his country in February 2022. Dressed in his trademark khaki green shirt, he urged member states to come together to oppose Russian aggression and stressed the need for a peace recognising Ukraine's territorial integrity. Here is the full text of Zelenskyy's speech from September 19. I welcome all who stand for common efforts! And I promise – being really united we can guarantee fair peace for all nations.
Unveiling Optimal SDG Pathways: An Innovative Approach Leveraging Graph Pruning and Intent Graph for Effective Recommendations
Yu, Zhihang, Wang, Shu, Zhu, Yunqiang, Yuan, Wen, Dai, Xiaoliang, Zou, Zhiqiang
The recommendation of appropriate development pathways, also known as ecological civilization patterns for achieving Sustainable Development Goals (namely, sustainable development patterns), are of utmost importance for promoting ecological, economic, social, and resource sustainability in a specific region. To achieve this, the recommendation process must carefully consider the region's natural, environmental, resource, and economic characteristics. However, current recommendation algorithms in the field of computer science fall short in adequately addressing the spatial heterogeneity related to environment and sparsity of regional historical interaction data, which limits their effectiveness in recommending sustainable development patterns. To overcome these challenges, this paper proposes a method called User Graph after Pruning and Intent Graph (UGPIG). Firstly, we utilize the high-density linking capability of the pruned User Graph to address the issue of spatial heterogeneity neglect in recommendation algorithms. Secondly, we construct an Intent Graph by incorporating the intent network, which captures the preferences for attributes including environmental elements of target regions. This approach effectively alleviates the problem of sparse historical interaction data in the region. Through extensive experiments, we demonstrate that UGPIG outperforms state-of-the-art recommendation algorithms like KGCN, KGAT, and KGIN in sustainable development pattern recommendations, with a maximum improvement of 9.61% in Top-3 recommendation performance.