Oceania
A Survey on Dialogue Management in Human-Robot Interaction
Reimann, Merle M., Kunneman, Florian A., Oertel, Catharine, Hindriks, Koen V.
Social robots are robots that are designed specifically to interact with their human users [14] for example by using spoken dialogue. For social robots, the interaction with humans plays a crucial role [7, 27], for example in the context of elderly care [15] or education [9]. Robots that use speech as a main mode of interaction do not only need to understand the user's utterances, but also need to select appropriate responses given the context. Dialogue management (DM), according to Traum and Larsson [88], is the part of a dialogue system that performs four key functions: 1) it maintains and updates the context of the dialogue, 2) it includes the context of the utterance for interpretation of input, 3) it selects the timing and content of the next utterance, and 4) it coordinates with (non-)dialogue modules. In spoken dialogue systems, the dialogue manager receives its input from a natural language understanding (NLU) module and forwards its results to a natural language generation (NLG) module, which then generates the output (see Figure 1). In contrast to general DM, DM in human-robot interaction (HRI) has to also consider and manage the complexity added by social robots (see Figure 1). The concentric circles of the figure describe decisions that have to be made when designing a dialogue manager for human-robot interaction. From each circle, one or more options can be chosen and combined with each other.
Music Genre Classification with ResNet and Bi-GRU Using Visual Spectrograms
Music recommendation systems have emerged as a vital component to enhance user experience and satisfaction for the music streaming services, which dominates music consumption. The key challenge in improving these recommender systems lies in comprehending the complexity of music data, specifically for the underpinning music genre classification. The limitations of manual genre classification have highlighted the need for a more advanced system, namely the Automatic Music Genre Classification (AMGC) system. While traditional machine learning techniques have shown potential in genre classification, they heavily rely on manually engineered features and feature selection, failing to capture the full complexity of music data. On the other hand, deep learning classification architectures like the traditional Convolutional Neural Networks (CNN) are effective in capturing the spatial hierarchies but struggle to capture the temporal dynamics inherent in music data. To address these challenges, this study proposes a novel approach using visual spectrograms as input, and propose a hybrid model that combines the strength of the Residual neural Network (ResNet) and the Gated Recurrent Unit (GRU). This model is designed to provide a more comprehensive analysis of music data, offering the potential to improve the music recommender systems through achieving a more comprehensive analysis of music data and hence potentially more accurate genre classification.
Differences Between Hard and Noisy-labeled Samples: An Empirical Study
Forouzesh, Mahsa, Thiran, Patrick
Extracting noisy or incorrectly labeled samples from a labeled dataset with hard/difficult samples is an important yet under-explored topic. Two general and often independent lines of work exist, one focuses on addressing noisy labels, and another deals with hard samples. However, when both types of data are present, most existing methods treat them equally, which results in a decline in the overall performance of the model. In this paper, we first design various synthetic datasets with custom hardness and noisiness levels for different samples. Our proposed systematic empirical study enables us to better understand the similarities and more importantly the differences between hard-to-learn samples and incorrectly-labeled samples. These controlled experiments pave the way for the development of methods that distinguish between hard and noisy samples. Through our study, we introduce a simple yet effective metric that filters out noisy-labeled samples while keeping the hard samples. We study various data partitioning methods in the presence of label noise and observe that filtering out noisy samples from hard samples with this proposed metric results in the best datasets as evidenced by the high test accuracy achieved after models are trained on the filtered datasets. We demonstrate this for both our created synthetic datasets and for datasets with real-world label noise. Furthermore, our proposed data partitioning method significantly outperforms other methods when employed within a semi-supervised learning framework.
Towards an architectural framework for intelligent virtual agents using probabilistic programming
Andreev, Anton, Cattan, Grégoire
We present a new framework called KorraAI for conceiving and building embodied conversational agents (ECAs). Our framework models ECAs' behavior considering contextual information, for example, about environment and interaction time, and uncertain information provided by the human interaction partner. Moreover, agents built with KorraAI can show proactive behavior, as they can initiate interactions with human partners. For these purposes, KorraAI exploits probabilistic programming. Probabilistic models in KorraAI are used to model its behavior and interactions with the user. They enable adaptation to the user's preferences and a certain degree of indeterminism in the ECAs to achieve more natural behavior. Human-like internal states, such as moods, preferences, and emotions (e.g., surprise), can be modeled in KorraAI with distributions and Bayesian networks. These models can evolve over time, even without interaction with the user. ECA models are implemented as plugins and share a common interface. This enables ECA designers to focus more on the character they are modeling and less on the technical details, as well as to store and exchange ECA models. Several applications of KorraAI ECAs are possible, such as virtual sales agents, customer service agents, virtual companions, entertainers, or tutors.
A Survey of What to Share in Federated Learning: Perspectives on Model Utility, Privacy Leakage, and Communication Efficiency
Shao, Jiawei, Li, Zijian, Sun, Wenqiang, Zhou, Tailin, Sun, Yuchang, Liu, Lumin, Lin, Zehong, Zhang, Jun
Federated learning (FL) has emerged as a highly effective paradigm for privacy-preserving collaborative training among different parties. Unlike traditional centralized learning, which requires collecting data from each party, FL allows clients to share privacy-preserving information without exposing private datasets. This approach not only guarantees enhanced privacy protection but also facilitates more efficient and secure collaboration among multiple participants. Therefore, FL has gained considerable attention from researchers, promoting numerous surveys to summarize the related works. However, the majority of these surveys concentrate on methods sharing model parameters during the training process, while overlooking the potential of sharing other forms of local information. In this paper, we present a systematic survey from a new perspective, i.e., what to share in FL, with an emphasis on the model utility, privacy leakage, and communication efficiency. This survey differs from previous ones due to four distinct contributions. First, we present a new taxonomy of FL methods in terms of the sharing methods, which includes three categories of shared information: model sharing, synthetic data sharing, and knowledge sharing. Second, we analyze the vulnerability of different sharing methods to privacy attacks and review the defense mechanisms that provide certain privacy guarantees. Third, we conduct extensive experiments to compare the performance and communication overhead of various sharing methods in FL. Besides, we assess the potential privacy leakage through model inversion and membership inference attacks, while comparing the effectiveness of various defense approaches. Finally, we discuss potential deficiencies in current methods and outline future directions for improvement.
Challenges and Solutions in AI for All
Shams, Rifat Ara, Zowghi, Didar, Bano, Muneera
Yet, these considerations are often overlooked, leading to issues of bias, discrimination, and perceived untrustworthiness. In response, we conducted a Systematic Review to unearth challenges and solutions relating to D&I in AI. Our rigorous search yielded 48 research articles published between 2017 and 2022. Open coding of these papers revealed 55 unique challenges and 33 solutions for D&I in AI, as well as 24 unique challenges and 23 solutions for enhancing such practices using AI. This study, by offering a deeper understanding of these issues, will enlighten researchers and practitioners seeking to integrate these principles into future AI systems.
Ensemble Learning based Anomaly Detection for IoT Cybersecurity via Bayesian Hyperparameters Sensitivity Analysis
Lai, Tin, Farid, Farnaz, Bello, Abubakar, Sabrina, Fariza
The Internet of Things (IoT) integrates more than billions of intelligent devices over the globe with the capability of communicating with other connected devices with little to no human intervention. IoT enables data aggregation and analysis on a large scale to improve life quality in many domains. In particular, data collected by IoT contain a tremendous amount of information for anomaly detection. The heterogeneous nature of IoT is both a challenge and an opportunity for cybersecurity. Traditional approaches in cybersecurity monitoring often require different kinds of data pre-processing and handling for various data types, which might be problematic for datasets that contain heterogeneous features. However, heterogeneous types of network devices can often capture a more diverse set of signals than a single type of device readings, which is particularly useful for anomaly detection. In this paper, we present a comprehensive study on using ensemble machine learning methods for enhancing IoT cybersecurity via anomaly detection. Rather than using one single machine learning model, ensemble learning combines the predictive power from multiple models, enhancing their predictive accuracy in heterogeneous datasets rather than using one single machine learning model. We propose a unified framework with ensemble learning that utilises Bayesian hyperparameter optimisation to adapt to a network environment that contains multiple IoT sensor readings. Experimentally, we illustrate their high predictive power when compared to traditional methods.
Intelligent model for offshore China sea fog forecasting
Xiang, Yanfei, Zhang, Qinghong, Wang, Mingqing, Xia, Ruixue, Kong, Yang, Huang, Xiaomeng
Accurate and timely prediction of sea fog is very important for effectively managing maritime and coastal economic activities. Given the intricate nature and inherent variability of sea fog, traditional numerical and statistical forecasting methods are often proven inadequate. This study aims to develop an advanced sea fog forecasting method embedded in a numerical weather prediction model using the Yangtze River Estuary (YRE) coastal area as a case study. Prior to training our machine learning model, we employ a time-lagged correlation analysis technique to identify key predictors and decipher the underlying mechanisms driving sea fog occurrence. In addition, we implement ensemble learning and a focal loss function to address the issue of imbalanced data, thereby enhancing the predictive ability of our model. To verify the accuracy of our method, we evaluate its performance using a comprehensive dataset spanning one year, which encompasses both weather station observations and historical forecasts. Remarkably, our machine learning-based approach surpasses the predictive performance of two conventional methods, the weather research and forecasting nonhydrostatic mesoscale model (WRF-NMM) and the algorithm developed by the National Oceanic and Atmospheric Administration (NOAA) Forecast Systems Laboratory (FSL). Specifically, in regard to predicting sea fog with a visibility of less than or equal to 1 km with a lead time of 60 hours, our methodology achieves superior results by increasing the probability of detection (POD) while simultaneously reducing the false alarm ratio (FAR).
Towards Robust Aspect-based Sentiment Analysis through Non-counterfactual Augmentations
Liu, Xinyu, Ding, Yan, An, Kaikai, Xiao, Chunyang, Madhyastha, Pranava, Xiao, Tong, Zhu, Jingbo
While state-of-the-art NLP models have demonstrated excellent performance for aspect based sentiment analysis (ABSA), substantial evidence has been presented on their lack of robustness. This is especially manifested as significant degradation in performance when faced with out-of-distribution data. Recent solutions that rely on counterfactually augmented datasets show promising results, but they are inherently limited because of the lack of access to explicit causal structure. In this paper, we present an alternative approach that relies on non-counterfactual data augmentation. Our proposal instead relies on using noisy, cost-efficient data augmentations that preserve semantics associated with the target aspect. Our approach then relies on modelling invariances between different versions of the data to improve robustness. A comprehensive suite of experiments shows that our proposal significantly improves upon strong pre-trained baselines on both standard and robustness-specific datasets. Our approach further establishes a new state-of-the-art on the ABSA robustness benchmark and transfers well across domains.
Factoring the Matrix of Domination: A Critical Review and Reimagination of Intersectionality in AI Fairness
Ovalle, Anaelia, Subramonian, Arjun, Gautam, Vagrant, Gee, Gilbert, Chang, Kai-Wei
These notions vary across conceptualization Intersectionality is a critical framework that, through inquiry and (e.g., group, individual fairness [8]) and operationalization (e.g., praxis, allows us to examine how social inequalities persist through pre/in/post-processing [2]) [54]; nevertheless, the literature generally domains of structure and discipline. Given AI fairness' raison d'être agrees on the goal of minimizing negative outcomes across of "fairness," we argue that adopting intersectionality as an analytical demographic groups, including groups associated with multiple, framework is pivotal to effectively operationalizing fairness. "intersectional" demographic attributes (e.g., Black women) [92]. Through a critical review of how intersectionality is discussed in However, Kong [66] observes that AI fairness papers often narrowly 30 papers from the AI fairness literature, we deductively and inductively: interpret intersectional subgroup fairness as intersectionality, the 1) map how intersectionality tenets operate within the critical framework from which the term originates [29, 67]. This AI fairness paradigm and 2) uncover gaps between the conceptualization myopic conceptualization of intersectionality has non-trivial consequences and operationalization of intersectionality. We find that for just AI design and epistemology (i.e., ways of knowing).