Goto

Collaborating Authors

 Oceania


Transformers in Speech Processing: A Survey

arXiv.org Artificial Intelligence

The remarkable success of transformers in the field of natural language processing has sparked the interest of the speech-processing community, leading to an exploration of their potential for modeling long-range dependencies within speech sequences. Recently, transformers have gained prominence across various speech-related domains, including automatic speech recognition, speech synthesis, speech translation, speech para-linguistics, speech enhancement, spoken dialogue systems, and numerous multimodal applications. In this paper, we present a comprehensive survey that aims to bridge research studies from diverse subfields within speech technology. By consolidating findings from across the speech technology landscape, we provide a valuable resource for researchers interested in harnessing the power of transformers to advance the field. We identify the challenges encountered by transformers in speech processing while also offering insights into potential solutions to address these issues.


Lidar Line Selection with Spatially-Aware Shapley Value for Cost-Efficient Depth Completion

arXiv.org Artificial Intelligence

Lidar is a vital sensor for estimating the depth of a scene. Typical spinning lidars emit pulses arranged in several horizontal lines and the monetary cost of the sensor increases with the number of these lines. In this work, we present the new problem of optimizing the positioning of lidar lines to find the most effective configuration for the depth completion task. We propose a solution to reduce the number of lines while retaining the up-to-the-mark quality of depth completion. Our method consists of two components, (1) line selection based on the marginal contribution of a line computed via the Shapley value and (2) incorporating line position spread to take into account its need to arrive at image-wide depth completion. Spatially-aware Shapley values (SaS) succeed in selecting line subsets that yield a depth accuracy comparable to the full lidar input while using just half of the lines.


Unlocking Layer-wise Relevance Propagation for Autoencoders

arXiv.org Artificial Intelligence

Autoencoders are a powerful and versatile tool often used for various problems such as anomaly detection, image processing and machine translation. However, their reconstructions are not always trivial to explain. Therefore, we propose a fast explainability solution by extending the Layer-wise Relevance Propagation method with the help of Deep Taylor Decomposition framework. Furthermore, we introduce a novel validation technique for comparing our explainability approach with baseline methods in the case of missing ground-truth data. Our results highlight computational as well as qualitative advantages of the proposed explainability solution with respect to existing methods.


Dens-PU: PU Learning with Density-Based Positive Labeled Augmentation

arXiv.org Artificial Intelligence

Labeled data are often scarce and expensive to obtain in many real-world applications, making training machine-learning models a challenging task [1]. In traditional supervised learning, the goal is to train a model to predict the correct class label for every sample in a training dataset [2]. The training data consist of labeled examples associated with a known class label. Typically, the class distribution of labeled data is assumed to be representative of the class distribution of unlabeled ones. Prior knowledge of the labels makes it easy to train a model to accurately predict the class labels for unseen samples. As Figure 1 shows, in the case of PU learning, the class label is known only for data belonging to a single class; thus, for negative samples, the label is unknown [3]. The lack of this knowledge makes it impossible to effectively train a typical binary classification model to distinguish between positive and negative classes.


A Survey on Class Imbalance in Federated Learning

arXiv.org Artificial Intelligence

Federated learning, which allows multiple client devices in a network to jointly train a machine learning model without direct exposure of clients' data, is an emerging distributed learning technique due to its nature of privacy preservation. However, it has been found that models trained with federated learning usually have worse performance than their counterparts trained in the standard centralized learning mode, especially when the training data is imbalanced. In the context of federated learning, data imbalance may occur either locally one one client device, or globally across many devices. The complexity of different types of data imbalance has posed challenges to the development of federated learning technique, especially considering the need of relieving data imbalance issue and preserving data privacy at the same time. Therefore, in the literature, many attempts have been made to handle class imbalance in federated learning. In this paper, we present a detailed review of recent advancements along this line. We first introduce various types of class imbalance in federated learning, after which we review existing methods for estimating the extent of class imbalance without the need of knowing the actual data to preserve data privacy. After that, we discuss existing methods for handling class imbalance in FL, where the advantages and disadvantages of the these approaches are discussed. We also summarize common evaluation metrics for class imbalanced tasks, and point out potential future directions.


Fundamentals of Generative Large Language Models and Perspectives in Cyber-Defense

arXiv.org Artificial Intelligence

Generative Language Models gained significant attention in late 2022 / early 2023, notably with the introduction of models refined to act consistently with users' expectations of interactions with AI (conversational models). Arguably the focal point of public attention has been such a refinement of the GPT3 model -- the ChatGPT and its subsequent integration with auxiliary capabilities, including search as part of Microsoft Bing. Despite extensive prior research invested in their development, their performance and applicability to a range of daily tasks remained unclear and niche. However, their wider utilization without a requirement for technical expertise, made in large part possible through conversational fine-tuning, revealed the extent of their true capabilities in a real-world environment. This has garnered both public excitement for their potential applications and concerns about their capabilities and potential malicious uses. This review aims to provide a brief overview of the history, state of the art, and implications of Generative Language Models in terms of their principles, abilities, limitations, and future prospects -- especially in the context of cyber-defense, with a focus on the Swiss operational environment.


How AI fooled Centrelink, and could fool you

#artificialintelligence

Thanks to artificial intelligence, faking someone's voice is easier than ever - all you need is a few minutes of audio. An investigation by Guardian Australia has found that this technology is able to fool a voice identification system that's used by the Australian government to secure the private information of millions of people. Data and interactives editor Nick Evershed explains how he discovered this security flaw and AI expert Toby Walsh explores how this technology could potentially make it easier than ever to steal someone's identity or commit scams


Apple HomePod review: a Siri speaker with a bass problem

The Guardian

Apple's big, high-quality smart speaker is back for a surprise second generation. But five years since the first model was launched, a lot has changed in the world of voice-controlled home hi-fi. Can the HomePod still cut it? The new HomePod has the same design as the old version: a marshmallow-like shape with a light-up disc at the top, fabric-covered body and a small silicone foot. The detachable power cable slots in the back but otherwise there are no ports or recesses. As with other HomePods, this speaker is for Apple users only.


Cybersecurity funds should go towards beefing up Centrelink voice authentication, Greens say

The Guardian

The federal government should be using some of the $10bn allocated in the budget to cybersecurity defences to combat people using AI to bypass biometric securities including voice authentication, a Greens senator has said. On Friday Guardian Australia reported that Centrelink's voice authentication system can be tricked using a free online AI cloning service and just four minutes of audio of the user's voice. After the Guardian Australia journalist Nick Evershed cloned his own voice, he was able to access his account using his cloned voice and his customer reference number. The voiceprint service, provided by the Microsoft-owned voice software company Nuance, was being used by 3.8 million Centrelink clients at the end of February, and more than 7.1 million people had verified their voice using the same system with the Australian Taxation Office. Despite being alerted to the vulnerability last week, Services Australia has not indicated it will change its use of voice ID, saying the technology is a "highly secure authentication method" and the agency "continually scans for potential threats and make ongoing enhancements to ensure customer security".


NASA Science Mission Directorate Knowledge Graph Discovery

arXiv.org Artificial Intelligence

The size of the National Aeronautics and Space Administration (NASA) Science Mission Directorate (SMD) is growing exponentially, allowing researchers to make discoveries. However, making discoveries is challenging and time-consuming due to the size of the data catalogs, and as many concepts and data are indirectly connected. This paper proposes a pipeline to generate knowledge graphs (KGs) representing different NASA SMD domains. These KGs can be used as the basis for dataset search engines, saving researchers time and supporting them in finding new connections. We collected textual data and used several modern natural language processing (NLP) methods to create the nodes and the edges of the KGs. We explore the cross-domain connections, discuss our challenges, and provide future directions to inspire researchers working on similar challenges.