Goto

Collaborating Authors

 Oceania


Unlocking Layer-wise Relevance Propagation for Autoencoders

arXiv.org Artificial Intelligence

Autoencoders are a powerful and versatile tool often used for various problems such as anomaly detection, image processing and machine translation. However, their reconstructions are not always trivial to explain. Therefore, we propose a fast explainability solution by extending the Layer-wise Relevance Propagation method with the help of Deep Taylor Decomposition framework. Furthermore, we introduce a novel validation technique for comparing our explainability approach with baseline methods in the case of missing ground-truth data. Our results highlight computational as well as qualitative advantages of the proposed explainability solution with respect to existing methods.


Dens-PU: PU Learning with Density-Based Positive Labeled Augmentation

arXiv.org Artificial Intelligence

Labeled data are often scarce and expensive to obtain in many real-world applications, making training machine-learning models a challenging task [1]. In traditional supervised learning, the goal is to train a model to predict the correct class label for every sample in a training dataset [2]. The training data consist of labeled examples associated with a known class label. Typically, the class distribution of labeled data is assumed to be representative of the class distribution of unlabeled ones. Prior knowledge of the labels makes it easy to train a model to accurately predict the class labels for unseen samples. As Figure 1 shows, in the case of PU learning, the class label is known only for data belonging to a single class; thus, for negative samples, the label is unknown [3]. The lack of this knowledge makes it impossible to effectively train a typical binary classification model to distinguish between positive and negative classes.


A Survey on Class Imbalance in Federated Learning

arXiv.org Artificial Intelligence

Federated learning, which allows multiple client devices in a network to jointly train a machine learning model without direct exposure of clients' data, is an emerging distributed learning technique due to its nature of privacy preservation. However, it has been found that models trained with federated learning usually have worse performance than their counterparts trained in the standard centralized learning mode, especially when the training data is imbalanced. In the context of federated learning, data imbalance may occur either locally one one client device, or globally across many devices. The complexity of different types of data imbalance has posed challenges to the development of federated learning technique, especially considering the need of relieving data imbalance issue and preserving data privacy at the same time. Therefore, in the literature, many attempts have been made to handle class imbalance in federated learning. In this paper, we present a detailed review of recent advancements along this line. We first introduce various types of class imbalance in federated learning, after which we review existing methods for estimating the extent of class imbalance without the need of knowing the actual data to preserve data privacy. After that, we discuss existing methods for handling class imbalance in FL, where the advantages and disadvantages of the these approaches are discussed. We also summarize common evaluation metrics for class imbalanced tasks, and point out potential future directions.


Fundamentals of Generative Large Language Models and Perspectives in Cyber-Defense

arXiv.org Artificial Intelligence

Generative Language Models gained significant attention in late 2022 / early 2023, notably with the introduction of models refined to act consistently with users' expectations of interactions with AI (conversational models). Arguably the focal point of public attention has been such a refinement of the GPT3 model -- the ChatGPT and its subsequent integration with auxiliary capabilities, including search as part of Microsoft Bing. Despite extensive prior research invested in their development, their performance and applicability to a range of daily tasks remained unclear and niche. However, their wider utilization without a requirement for technical expertise, made in large part possible through conversational fine-tuning, revealed the extent of their true capabilities in a real-world environment. This has garnered both public excitement for their potential applications and concerns about their capabilities and potential malicious uses. This review aims to provide a brief overview of the history, state of the art, and implications of Generative Language Models in terms of their principles, abilities, limitations, and future prospects -- especially in the context of cyber-defense, with a focus on the Swiss operational environment.


How AI fooled Centrelink, and could fool you

#artificialintelligence

Thanks to artificial intelligence, faking someone's voice is easier than ever - all you need is a few minutes of audio. An investigation by Guardian Australia has found that this technology is able to fool a voice identification system that's used by the Australian government to secure the private information of millions of people. Data and interactives editor Nick Evershed explains how he discovered this security flaw and AI expert Toby Walsh explores how this technology could potentially make it easier than ever to steal someone's identity or commit scams


Apple HomePod review: a Siri speaker with a bass problem

The Guardian

Apple's big, high-quality smart speaker is back for a surprise second generation. But five years since the first model was launched, a lot has changed in the world of voice-controlled home hi-fi. Can the HomePod still cut it? The new HomePod has the same design as the old version: a marshmallow-like shape with a light-up disc at the top, fabric-covered body and a small silicone foot. The detachable power cable slots in the back but otherwise there are no ports or recesses. As with other HomePods, this speaker is for Apple users only.


Cybersecurity funds should go towards beefing up Centrelink voice authentication, Greens say

The Guardian

The federal government should be using some of the $10bn allocated in the budget to cybersecurity defences to combat people using AI to bypass biometric securities including voice authentication, a Greens senator has said. On Friday Guardian Australia reported that Centrelink's voice authentication system can be tricked using a free online AI cloning service and just four minutes of audio of the user's voice. After the Guardian Australia journalist Nick Evershed cloned his own voice, he was able to access his account using his cloned voice and his customer reference number. The voiceprint service, provided by the Microsoft-owned voice software company Nuance, was being used by 3.8 million Centrelink clients at the end of February, and more than 7.1 million people had verified their voice using the same system with the Australian Taxation Office. Despite being alerted to the vulnerability last week, Services Australia has not indicated it will change its use of voice ID, saying the technology is a "highly secure authentication method" and the agency "continually scans for potential threats and make ongoing enhancements to ensure customer security".


NASA Science Mission Directorate Knowledge Graph Discovery

arXiv.org Artificial Intelligence

The size of the National Aeronautics and Space Administration (NASA) Science Mission Directorate (SMD) is growing exponentially, allowing researchers to make discoveries. However, making discoveries is challenging and time-consuming due to the size of the data catalogs, and as many concepts and data are indirectly connected. This paper proposes a pipeline to generate knowledge graphs (KGs) representing different NASA SMD domains. These KGs can be used as the basis for dataset search engines, saving researchers time and supporting them in finding new connections. We collected textual data and used several modern natural language processing (NLP) methods to create the nodes and the edges of the KGs. We explore the cross-domain connections, discuss our challenges, and provide future directions to inspire researchers working on similar challenges.


A reproducible approach to merging behavior analysis based on High Definition Map

arXiv.org Artificial Intelligence

Existing research on merging behavior generally prioritize the application of various algorithms, but often overlooks the fine-grained process and analysis of trajectories. This leads to the neglect of surrounding vehicle matching, the opaqueness of indicators definition, and reproducible crisis. To address these gaps, this paper presents a reproducible approach to merging behavior analysis. Specifically, we outline the causes of subjectivity and irreproducibility in existing studies. Thereafter, we employ lanelet2 High Definition (HD) map to construct a reproducible framework, that minimizes subjectivities, defines standardized indicators, identifies alongside vehicles, and divides scenarios. A comparative macroscopic and microscopic analysis is subsequently conducted. More importantly, this paper adheres to the Reproducible Research concept, providing all the source codes and reproduction instructions. Our results demonstrate that although scenarios with alongside vehicles occur in less than 6% of cases, their characteristics are significantly different from others, and these scenarios are often accompanied by high risk. This paper refines the understanding of merging behavior, raises awareness of reproducible studies, and serves as a watershed moment.


A fuzzy adaptive evolutionary-based feature selection and machine learning framework for single and multi-objective body fat prediction

arXiv.org Artificial Intelligence

Predicting body fat can provide medical practitioners and users with essential information for preventing and diagnosing heart diseases. Hybrid machine learning models offer better performance than simple regression analysis methods by selecting relevant body measurements and capturing complex nonlinear relationships among selected features in modelling body fat prediction problems. There are, however, some disadvantages to them. Current machine learning. Modelling body fat prediction as a combinatorial single- and multi-objective optimisation problem often gets stuck in local optima. When multiple feature subsets produce similar or close predictions, avoiding local optima becomes more complex. Evolutionary feature selection has been used to solve several machine-learning-based optimisation problems. A fuzzy set theory determines appropriate levels of exploration and exploitation while managing parameterisation and computational costs. A weighted-sum body fat prediction approach was explored using evolutionary feature selection, fuzzy set theory, and machine learning algorithms, integrating contradictory metrics into a single composite goal optimised by fuzzy adaptive evolutionary feature selection. Hybrid fuzzy adaptive global learning local search universal diversity-based feature selection is applied to this single-objective feature selection-machine learning framework (FAGLSUD-based FS-ML). While using fewer features, this model achieved a more accurate and stable estimate of body fat percentage than other hybrid and state-of-the-art machine learning models. A multi-objective FAGLSUD-based FS-MLP is also proposed to analyse accuracy, stability, and dimensionality conflicts simultaneously. To make informed decisions about fat deposits in the most vital body parts and blood lipid levels, medical practitioners and users can use a well-distributed Pareto set of trade-off solutions.