Goto

Collaborating Authors

 Oceania


Dialogue Agents 101: A Beginner's Guide to Critical Ingredients for Designing Effective Conversational Systems

arXiv.org Artificial Intelligence

Sharing ideas through communication with peers is the primary mode of human interaction. Consequently, extensive research has been conducted in the area of conversational AI, leading to an increase in the availability and diversity of conversational tasks, datasets, and methods. However, with numerous tasks being explored simultaneously, the current landscape of conversational AI becomes fragmented. Therefore, initiating a well-thought-out model for a dialogue agent can pose significant challenges for a practitioner. Towards highlighting the critical ingredients needed for a practitioner to design a dialogue agent from scratch, the current study provides a comprehensive overview of the primary characteristics of a dialogue agent, the supporting tasks, their corresponding open-domain datasets, and the methods used to benchmark these datasets. We observe that different methods have been used to tackle distinct dialogue tasks. However, building separate models for each task is costly and does not leverage the correlation among the several tasks of a dialogue agent. As a result, recent trends suggest a shift towards building unified foundation models. To this end, we propose UNIT, a UNified dIalogue dataseT constructed from conversations of existing datasets for different dialogue tasks capturing the nuances for each of them. We also examine the evaluation strategies used to measure the performance of dialogue agents and highlight the scope for future research in the area of conversational AI.


Rigorous Runtime Analysis of Diversity Optimization with GSEMO on OneMinMax

arXiv.org Artificial Intelligence

The evolutionary diversity optimization aims at finding a diverse set of solutions which satisfy some constraint on their fitness. In the context of multi-objective optimization this constraint can require solutions to be Pareto-optimal. In this paper we study how the GSEMO algorithm with additional diversity-enhancing heuristic optimizes a diversity of its population on a bi-objective benchmark problem OneMinMax, for which all solutions are Pareto-optimal. We provide a rigorous runtime analysis of the last step of the optimization, when the algorithm starts with a population with a second-best diversity, and prove that it finds a population with optimal diversity in expected time $O(n^2)$, when the problem size $n$ is odd. For reaching our goal, we analyse the random walk of the population, which reflects the frequency of changes in the population and their outcomes.


TriFormer: A Multi-modal Transformer Framework For Mild Cognitive Impairment Conversion Prediction

arXiv.org Artificial Intelligence

Magnetic resonance imaging (MRI) and Positron emission tomography (PET) could help more accurately predict MCI The prediction of mild cognitive impairment (MCI) conversion conversion [2]. to Alzheimer's disease (AD) is important for early Convolutional neural networks (CNNs) have been widely treatment to prevent or slow the progression of AD. To accurately applied to AD classification and prediction from imaging predict the MCI conversion to stable MCI or progressive data. Valliani et al. [3] fine-tuned a pretrained ResNet-50 MCI, we propose TriFormer, a novel transformer-based to classify AD and CN based on 2D axial slices. Wen et framework with three specialized transformers to incorporate al. [4] leveraged 3D spatial information by using a 3D CNN multi-modal data. TriFormer uses I) an image transformer to and outperformed previous 2D-based methods in AD classification extract multi-view image features from medical scans, II) a and MCI conversion prediction. However, both clinical transformer to embed and correlate multi-modal clinical 2D and 3D CNNs have a strong inductive bias towards local data, and III) a modality fusion transformer that produces receptive fields, which could limit the performance on an accurate prediction based on fusing the outputs from the high dimensional data [5]. Recently, transformers have been image and clinical transformers. Triformer is evaluated on the shown to be effective in capturing global long-range dependency Alzheimer's Disease Neuroimaging Initiative (ADNI) 1 and within imaging [6] and sequential data [7]. They also ADNI2 datasets and outperforms previous state-of-the-art have no indictive bias compared with CNNs.


Cross-Language Speech Emotion Recognition Using Multimodal Dual Attention Transformers

arXiv.org Artificial Intelligence

Despite the recent progress in speech emotion recognition (SER), state-of-the-art systems are unable to achieve improved performance in cross-language settings. In this paper, we propose a Multimodal Dual Attention Transformer (MDAT) model to improve cross-language SER. Our model utilises pre-trained models for multimodal feature extraction and is equipped with a dual attention mechanism including graph attention and co-attention to capture complex dependencies across different modalities and achieve improved cross-language SER results using minimal target language data. In addition, our model also exploits a transformer encoder layer for high-level feature representation to improve emotion classification accuracy. In this way, MDAT performs refinement of feature representation at various stages and provides emotional salient features to the classification layer. This novel approach also ensures the preservation of modality-specific emotional information while enhancing cross-modality and cross-language interactions. We assess our model's performance on four publicly available SER datasets and establish its superior effectiveness compared to recent approaches and baseline models.


Navigating Explanatory Multiverse Through Counterfactual Path Geometry

arXiv.org Artificial Intelligence

Counterfactual explanations are the de facto standard when tasked with interpreting decisions of (opaque) predictive models. Their generation is often subject to algorithmic and domain-specific constraints -- such as density-based feasibility and attribute (im)mutability or directionality of change -- that aim to maximise their real-life utility. In addition to desiderata with respect to the counterfactual instance itself, existence of a viable path connecting it with the factual data point, known as algorithmic recourse, has become an important technical consideration. While both of these requirements ensure that the steps of the journey as well as its destination are admissible, current literature neglects the multiplicity of such counterfactual paths. To address this shortcoming we introduce the novel concept of explanatory multiverse that encompasses all the possible counterfactual journeys; we then show how to navigate, reason about and compare the geometry of these trajectories -- their affinity, branching, divergence and possible future convergence -- with two methods: vector spaces and graphs. Implementing this (interactive) explanatory process grants explainees agency by allowing them to select counterfactuals based on the properties of the journey leading to them in addition to their absolute differences.


UNITE: A Unified Benchmark for Text-to-SQL Evaluation

arXiv.org Artificial Intelligence

A practical text-to-SQL system should generalize well on a wide variety of natural language questions, unseen database schemas, and novel SQL query structures. To comprehensively evaluate text-to-SQL systems, we introduce a UNIfied benchmark for Text-to-SQL Evaluation (UNITE). It is composed of publicly available text-to-SQL datasets, containing natural language questions from more than 12 domains, SQL queries from more than 3.9K patterns, and 29K databases. Compared to the widely used Spider benchmark, we introduce $\sim$120K additional examples and a threefold increase in SQL patterns, such as comparative and boolean questions. We conduct a systematic study of six state-of-the-art (SOTA) text-to-SQL parsers on our new benchmark and show that: 1) Codex performs surprisingly well on out-of-domain datasets; 2) specially designed decoding methods (e.g. constrained beam search) can improve performance for both in-domain and out-of-domain settings; 3) explicitly modeling the relationship between questions and schemas further improves the Seq2Seq models. More importantly, our benchmark presents key challenges towards compositional generalization and robustness issues -- which these SOTA models cannot address well. Our code and data processing script are available at https://github.com/awslabs/unified-text2sql-benchmark


Vision Transformer Based Model for Describing a Set of Images as a Story

arXiv.org Artificial Intelligence

Visual Story-Telling is the process of forming a multi-sentence story from a set of images. Appropriately including visual variation and contextual information captured inside the input images is one of the most challenging aspects of visual storytelling. Consequently, stories developed from a set of images often lack cohesiveness, relevance, and semantic relationship. In this paper, we propose a novel Vision Transformer Based Model for describing a set of images as a story. The proposed method extracts the distinct features of the input images using a Vision Transformer (ViT). Firstly, input images are divided into 16X16 patches and bundled into a linear projection of flattened patches. The transformation from a single image to multiple image patches captures the visual variety of the input visual patterns. These features are used as input to a Bidirectional-LSTM which is part of the sequence encoder. This captures the past and future image context of all image patches. Then, an attention mechanism is implemented and used to increase the discriminatory capacity of the data fed into the language model, i.e. a Mogrifier-LSTM. The performance of our proposed model is evaluated using the Visual Story-Telling dataset (VIST), and the results show that our model outperforms the current state of the art models.


Differentially Private Stochastic Gradient Descent with Low-Noise

arXiv.org Artificial Intelligence

Modern machine learning algorithms aim to extract fine-grained information from data to provide accurate predictions, which often conflicts with the goal of privacy protection. This paper addresses the practical and theoretical importance of developing privacy-preserving machine learning algorithms that ensure good performance while preserving privacy. In this paper, we focus on the privacy and utility (measured by excess risk bounds) performances of differentially private stochastic gradient descent (SGD) algorithms in the setting of stochastic convex optimization. Specifically, we examine the pointwise problem in the low-noise setting for which we derive sharper excess risk bounds for the differentially private SGD algorithm. In the pairwise learning setting, we propose a simple differentially private SGD algorithm based on gradient perturbation. Furthermore, we develop novel utility bounds for the proposed algorithm, proving that it achieves optimal excess risk rates even for non-smooth losses. Notably, we establish fast learning rates for privacy-preserving pairwise learning under the low-noise condition, which is the first of its kind.


Model-Assisted Probabilistic Safe Adaptive Control With Meta-Bayesian Learning

arXiv.org Artificial Intelligence

Despite the existence of numerous designs, significant research efforts, and successful applications in the field of control systems, the development of a reliable and secure controller that combines robust theoretical foundations with exceptional performance continues to present a formidable challenge. This challenge has captured the attention of researchers from diverse fields, including robotics [1] and healthcare [2], among others. In the context of control systems, safety is evaluated based on the system state. In this study, we focus on probabilistic safe control, wherein a safe controller is expected to prevent the system from entering hazardous states with an acceptable probability [3-5]. Due to the intricate nature of calculating the safe state space for a general dynamics-driven system, ensuring safety by designing or learning a safe controller is rather complex. Existing safe control strategies include model predictive control [6], reachability analysis [7], and control barrier function (CBF) method [8]. In our research, we build upon the CBF method, which ensures that the system state remains within safe regions by defining a forward invariant set. This set is a subset of the safe region and restricts the system state within its boundaries. Furthermore, we take into account the presence of uncertainty, which not only have a more significant impact on the system state than small disturbances [9], and does not have an analytical format as well [10].


Generating Efficient Training Data via LLM-based Attribute Manipulation

arXiv.org Artificial Intelligence

In this paper, we propose a novel method, Chain-of-Thoughts Attribute Manipulation (CoTAM), to guide few-shot learning by carefully crafted data from Large Language Models (LLMs). The main idea is to create data with changes only in the attribute targeted by the task. Inspired by facial attribute manipulation, our approach generates label-switched data by leveraging LLMs to manipulate task-specific attributes and reconstruct new sentences in a controlled manner. Instead of conventional latent representation controlling, we implement chain-of-thoughts decomposition and reconstruction to adapt the procedure to LLMs. Extensive results on text classification and other tasks verify the advantage of CoTAM over other LLM-based text generation methods with the same number of training examples. Analysis visualizes the attribute manipulation effectiveness of CoTAM and presents the potential of LLM-guided learning with even less supervision.