Goto

Collaborating Authors

 Africa


Your device may know you better than you know yourself -- continuous authentication on novel dataset using machine learning

arXiv.org Artificial Intelligence

This research aims to further understanding in the field of continuous authentication using behavioural biometrics. We are contributing a novel dataset that encompasses the gesture data of 15 users playing Minecraft with a Samsung Tablet, each for a duration of 15 minutes. Utilizing this dataset, we employed machine learning (ML) binary classifiers, being Random Forest (RF), K-Nearest Neighbors (KNN), and Support Vector Classifier (SVC), to determine the authenticity of specific user actions. Our most robust model was SVC, which achieved an average accuracy of approximately 90%, demonstrating that touch dynamics can effectively distinguish users. However, further studies are needed to make it viable option for authentication systems. NTRODUCTION The current authentication methods, which are primarily implemented at entry points, can be problematic in numerous scenarios.


FL-GUARD: A Holistic Framework for Run-Time Detection and Recovery of Negative Federated Learning

arXiv.org Artificial Intelligence

Federated learning (FL) is a promising approach for learning a model from data distributed on massive clients without exposing data privacy. It works effectively in the ideal federation where clients share homogeneous data distribution and learning behavior. However, FL may fail to function appropriately when the federation is not ideal, amid an unhealthy state called Negative Federated Learning (NFL), in which most clients gain no benefit from participating in FL. Many studies have tried to address NFL. However, their solutions either (1) predetermine to prevent NFL in the entire learning life-cycle or (2) tackle NFL in the aftermath of numerous learning rounds. Thus, they either (1) indiscriminately incur extra costs even if FL can perform well without such costs or (2) waste numerous learning rounds. Additionally, none of the previous work takes into account the clients who may be unwilling/unable to follow the proposed NFL solutions when using those solutions to upgrade an FL system in use. This paper introduces FL-GUARD, a holistic framework that can be employed on any FL system for tackling NFL in a run-time paradigm. That is, to dynamically detect NFL at the early stage (tens of rounds) of learning and then to activate recovery measures when necessary. Specifically, we devise a cost-effective NFL detection mechanism, which relies on an estimation of performance gain on clients. Only when NFL is detected, we activate the NFL recovery process, in which each client learns in parallel an adapted model when training the global model. Extensive experiment results confirm the effectiveness of FL-GUARD in detecting NFL and recovering from NFL to a healthy learning state. We also show that FL-GUARD is compatible with previous NFL solutions and robust against clients unwilling/unable to take any recovery measures.


Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts

arXiv.org Artificial Intelligence

To meet the shifts in post-pandemic learning needs and the demand of artificial intelligence (AI) advancement on workforce development, the education system seeks new instructional and learning strategies that are personalized, effective, safe, and scalable [8]. Throughout the years, richer and more complex educational data have been generated by the advancement of instructional practices, providing vast potential for analyses but at the same time posing challenges to the approaches that process such data. Conventional quantitative methods are limited by the capacity of calculation and the efficiency of models, hence preventing efforts to improve teaching and learning outcomes. AI/ML approaches are able to effectively process the existing and forthcoming complex data with scalability and precision [5], presenting an unprecedented opportunity to promote the research and instructional practices in education. These characteristics of new data and methods provide timely and actionable insights into the dynamics of the instructional environment. Furthermore, in recent years, this trend has been accelerated by the rapid adoption of generative AI tools, such as ChatGPT and Bard, which synergizes the capabilities of both text analysis and generation. A new field of research has emerged, in which researchers integrate the cutting-edge AI/ML techniques with educational domain knowledge of curriculum, teaching, and learning and to explore crucial questions for instructional improvement.


Bridging Diversity and Uncertainty in Active learning with Self-Supervised Pre-Training

arXiv.org Artificial Intelligence

This study addresses the integration of diversity-based and uncertainty-based sampling strategies in active learning, particularly within the context of self-supervised pre-trained models. We introduce a straightforward heuristic called TCM that mitigates the cold start problem while maintaining strong performance across various data levels. By initially applying TypiClust for diversity sampling and subsequently transitioning to uncertainty sampling with Margin, our approach effectively combines the strengths of both strategies. Our experiments demonstrate that TCM consistently outperforms existing methods across various datasets in both low and high data regimes.


KG-TREAT: Pre-training for Treatment Effect Estimation by Synergizing Patient Data with Knowledge Graphs

arXiv.org Artificial Intelligence

Treatment effect estimation (TEE) is the task of determining the impact of various treatments on patient outcomes. Current TEE methods fall short due to reliance on limited labeled data and challenges posed by sparse and high-dimensional observational patient data. To address the challenges, we introduce a novel pre-training and fine-tuning framework, KG-TREAT, which synergizes large-scale observational patient data with biomedical knowledge graphs (KGs) to enhance TEE. Unlike previous approaches, KG-TREAT constructs dual-focus KGs and integrates a deep bi-level attention synergy method for in-depth information fusion, enabling distinct encoding of treatment-covariate and outcome-covariate relationships. KG-TREAT also incorporates two pre-training tasks to ensure a thorough grounding and contextualization of patient data and KGs. Evaluation on four downstream TEE tasks shows KG-TREAT's superiority over existing methods, with an average improvement of 7% in Area under the ROC Curve (AUC) and 9% in Influence Function-based Precision of Estimating Heterogeneous Effects (IF-PEHE). The effectiveness of our estimated treatment effects is further affirmed by alignment with established randomized clinical trial findings.


A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation

arXiv.org Artificial Intelligence

Knowledge-based, open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. Many types and sources of knowledge have previously been shown to be useful as support knowledge. Even in the era of large language models, response generation grounded in knowledge retrieved from additional up-to-date sources remains a practically important approach. While prior work using single-source knowledge has shown a clear positive correlation between the performances of knowledge selection and response generation, there are no existing multi-source datasets for evaluating support knowledge retrieval. Further, prior work has assumed that the knowledge sources available at test time are the same as during training. This unrealistic assumption unnecessarily handicaps models, as new knowledge sources can become available after a model is trained. In this paper, we present a high-quality benchmark named multi-source Wizard of Wikipedia (Ms.WoW) for evaluating multi-source dialogue knowledge selection and response generation. Unlike existing datasets, it contains clean support knowledge, grounded at the utterance level and partitioned into multiple knowledge sources. We further propose a new challenge, dialogue knowledge plug-and-play, which aims to test an already trained dialogue model on using new support knowledge from previously unseen sources in a zero-shot fashion.


OCD-FL: A Novel Communication-Efficient Peer Selection-based Decentralized Federated Learning

arXiv.org Artificial Intelligence

The conjunction of edge intelligence and the ever-growing Internet-of-Things (IoT) network heralds a new era of collaborative machine learning, with federated learning (FL) emerging as the most prominent paradigm. With the growing interest in these learning schemes, researchers started addressing some of their most fundamental limitations. Indeed, conventional FL with a central aggregator presents a single point of failure and a network bottleneck. To bypass this issue, decentralized FL where nodes collaborate in a peer-to-peer network has been proposed. Despite the latter's efficiency, communication costs and data heterogeneity remain key challenges in decentralized FL. In this context, we propose a novel scheme, called opportunistic communication-efficient decentralized federated learning, a.k.a., OCD-FL, consisting of a systematic FL peer selection for collaboration, aiming to achieve maximum FL knowledge gain while reducing energy consumption. Experimental results demonstrate the capability of OCD-FL to achieve similar or better performances than the fully collaborative FL, while significantly reducing consumed energy by at least 30% and up to 80%.


Spectral Phase Transition and Optimal PCA in Block-Structured Spiked models

arXiv.org Machine Learning

The statistical challenge of inferring a low-dimensional signal from a noisy, high-dimensional observation is ubiquitous across statistics, probability, and machine learning. Spiked random matrix models have recently gained extensive interest, serving as a valuable platform for exploring this issue [30, 51, 42]. A prominent example is the spiked Wigner model, where a rank one matrix is observed through a component-wise homogeneous noise, that has been studied extensively in random matrix theory [10]. Most models, with the spiked Wigner model at the forefront, have focused however on scenarios where the noise is "homogeneous", aiming to understand how the performance of the inference depends on the noise level. Yet in practice, datasets are inherently structured and the exploration of inhomogeneity plays a pivotal role in unraveling their complexities. A prototypical model to study this phenomenon is to improve the aforementioned spiked Wigner model by introducing a block structure in the noise, a model which has been recently introduced in a series of papers [17, 5, 7, 34] and that arises in many different learning contexts such as community detection [17, 34], deep Boltzmann machines [6], or the dense limit of the celebrated degree-corrected stochastic block model [34, 39]. Our goal in this paper is to apply rigorous random matrix theory to such "inhomogenous" spiked models, and to provide an optimal reconstruction method from a spectral algorithm, to generalize the seminal work of [10] (BBP) to inhomogenous matrices.


Reducing the dimensionality and granularity in hierarchical categorical variables

arXiv.org Machine Learning

This may cause overfitting and estimation issues when including such covariates in a predictive model. In current literature, a hierarchical covariate is often incorporated via nested random effects. However, this does not facilitate the assumption of classes having the same effect on the response variable. In this paper, we propose a methodology to obtain a reduced representation of a hierarchical categorical variable. We show how entity embedding can be applied in a hierarchical setting. Subsequently, we propose a top-down clustering algorithm which leverages the information encoded in the embeddings to reduce both the within-level dimensionality as well as the overall granularity of the hierarchical categorical variable. In simulation experiments, we show that our methodology can effectively approximate the true underlying structure of a hierarchical covariate in terms of the effect on a response variable, and find that incorporating the reduced hierarchy improves model fit. We apply our methodology on a real dataset and find that the reduced hierarchy is an improvement over the original hierarchical structure and reduced structures proposed in the literature. MSC classification: 62H30, 68T07 Keywords: hierarchical categorical variable, entity embedding, clustering, predictive modelling, machine learning Data and code availability statement: Data and code are available on https://github.


Efficient Algorithms for Empirical Group Distributional Robust Optimization and Beyond

arXiv.org Machine Learning

We investigate the empirical counterpart of group distributionally robust optimization (GDRO), which aims to minimize the maximal empirical risk across $m$ distinct groups. We formulate empirical GDRO as a $\textit{two-level}$ finite-sum convex-concave minimax optimization problem and develop a stochastic variance reduced mirror prox algorithm. Unlike existing methods, we construct the stochastic gradient by per-group sampling technique and perform variance reduction for all groups, which fully exploits the $\textit{two-level}$ finite-sum structure of empirical GDRO. Furthermore, we compute the snapshot and mirror snapshot point by a one-index-shifted weighted average, which distinguishes us from the naive ergodic average. Our algorithm also supports non-constant learning rates, which is different from existing literature. We establish convergence guarantees both in expectation and with high probability, demonstrating a complexity of $\mathcal{O}\left(\frac{m\sqrt{\bar{n}\ln{m}}}{\varepsilon}\right)$, where $\bar n$ is the average number of samples among $m$ groups. Remarkably, our approach outperforms the state-of-the-art method by a factor of $\sqrt{m}$. Furthermore, we extend our methodology to deal with the empirical minimax excess risk optimization (MERO) problem and manage to give the expectation bound and the high probability bound, accordingly. The complexity of our empirical MERO algorithm matches that of empirical GDRO at $\mathcal{O}\left(\frac{m\sqrt{\bar{n}\ln{m}}}{\varepsilon}\right)$, significantly surpassing the bounds of existing methods.