Education
Review on Fault Diagnosis and Fault-Tolerant Control Scheme for Robotic Manipulators: Recent Advances in AI, Machine Learning, and Digital Twin
Quamar, Md Muzakkir, Nasir, Ali
This comprehensive review article delves into the intricate realm of fault-tolerant control (FTC) schemes tailored for robotic manipulators. Our exploration spans the historical evolution of FTC, tracing its development over time, and meticulously examines the recent breakthroughs fueled by the synergistic integration of cutting-edge technologies such as artificial intelligence (AI), machine learning (ML), and digital twin technologies (DTT). The article places a particular emphasis on the transformative influence these contemporary trends exert on the landscape of robotic manipulator control and fault tolerance. By delving into the historical context, our aim is to provide a comprehensive understanding of the evolution of FTC schemes. This journey encompasses the transition from model-based and signal-based schemes to the role of sensors, setting the stage for an exploration of the present-day paradigm shift enabled by AI, ML, and DTT. The narrative unfolds as we dissect the intricate interplay between these advanced technologies and their applications in enhancing fault tolerance within the domain of robotic manipulators. Our review critically evaluates the impact of these advancements, shedding light on the novel methodologies, techniques, and applications that have emerged in recent times. The overarching goal of this article is to present a comprehensive perspective on the current state of fault diagnosis and fault-tolerant control within the context of robotic manipulators, positioning our exploration within the broader framework of AI, ML, and DTT advancements. Through a meticulous examination of both historical foundations and contemporary innovations, this review significantly contributes to the existing body of knowledge, offering valuable insights for researchers, practitioners, and enthusiasts navigating the dynamic landscape of robotic manipulator control.
Learning solutions of parametric Navier-Stokes with physics-informed neural networks
Naderibeni, M., Reinders, M. J. T., Wu, L., Tax, D. M. J.
We leverage Physics-Informed Neural Networks (PINNs) to learn solution functions of parametric Navier-Stokes Equations (NSE). Our proposed approach results in a feasible optimization problem setup that bypasses PINNs' limitations in converging to solutions of highly nonlinear parametric-PDEs like NSE. We consider the parameter(s) of interest as inputs of PINNs along with spatio-temporal coordinates, and train PINNs on generated numerical solutions of parametric-PDES for instances of the parameters. We perform experiments on the classical 2D flow past cylinder problem aiming to learn velocities and pressure functions over a range of Reynolds numbers as parameter of interest. Provision of training data from generated numerical simulations allows for interpolation of the solution functions for a range of parameters. Therefore, we compare PINNs with unconstrained conventional Neural Networks (NN) on this problem setup to investigate the effectiveness of considering the PDEs regularization in the loss function. We show that our proposed approach results in optimizing PINN models that learn the solution functions while making sure that flow predictions are in line with conservational laws of mass and momentum. Our results show that PINN results in accurate prediction of gradients compared to NN model, this is clearly visible in predicted vorticity fields given that none of these models were trained on vorticity labels.
From Algorithm Worship to the Art of Human Learning: Insights from 50-year journey of AI in Education
Over the past decade, there have been increasing proclama5ons from diverse stakeholders that humanity is at an inflec5on point due to advances in Ar5ficial Intelligence (AI) technologies (e.g., Crawford, 2017). The general public are condi5oned by this messaging to expect both big (though so far largely non-descript) changes to our lives, including to the way that we learn and teach. Warnings have been also ar5culated regarding whether and how AI might fundamentally change the way we perceive reality, how we form our beliefs, or interact with one another (Bostrom, 2017). More recently, ques5ons started to emerge about AI's transforma5ve poten5al (for beLer or worse) for our func5oning at neurocogni5ve, socio-emo5onal, individual and collec5ve levels (UNESCO, 2022; Pedro, et al., 2019, Porayska-Pomsta, 2023), along with concerns regarding the ethical implica5ons of using AI for suppor5ng human decision-making in contexts that are both high-stakes (e.g., for medical diagnoses or for student assessment) and rela5vely low-stakes, e.g., selec5ng movies on streaming sites. Such hope-fear rhetoric is also present in the context of AI applica5ons to suppor5ng human learning in formal and informal contexts. Recent hopes for AI in educa5on (AIED) largely relate to delivering learning at scale across different geographical and cultural contexts, especially in light of growing global teacher shortages and diminishing funding for educa5on in many countries (UNESCO, 2023). These hopes are increasingly used to fuel poli5cally and market mo5vated discourse about the need to'release teachers from tedious tasks' such as standardised assessments to allow them to focus on the'things that maLer' (Gen5le et al., 2023), or to jus5fy the narrowing of the formal educa5on curricula mainly to STEM subjects.
Comparison of Topic Modelling Approaches in the Banking Context
Ogunleye, Bayode, Maswera, Tonderai, Hirsch, Laurence, Gaudoin, Jotham, Brunsdon, Teresa
Topic modelling is a prominent task for automatic topic extraction in many applications such as sentiment analysis and recommendation systems. The approach is vital for service industries to monitor their customer discussions. The use of traditional approaches such as Latent Dirichlet Allocation (LDA) for topic discovery has shown great performances, however, they are not consistent in their results as these approaches suffer from data sparseness and inability to model the word order in a document. Thus, this study presents the use of Kernel Principal Component Analysis (KernelPCA) and K-means Clustering in the BERTopic architecture. We have prepared a new dataset using tweets from customers of Nigerian banks and we use this to compare the topic modelling approaches. Our findings showed KernelPCA and K-means in the BERTopic architecture-produced coherent topics with a coherence score of 0.8463.
O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language Models
Xiao, Yuchen, Sun, Yanchao, Xu, Mengda, Madhushani, Udari, Vann, Jared, Garg, Deepeka, Ganesh, Sumitra
Recent advancements in large language models (LLMs) have exhibited promising performance in solving sequential decision-making problems. By imitating few-shot examples provided in the prompts (i.e., in-context learning), an LLM agent can interact with an external environment and complete given tasks without additional training. However, such few-shot examples are often insufficient to generate high-quality solutions for complex and long-horizon tasks, while the limited context length cannot consume larger-scale demonstrations with long interaction horizons. To this end, we propose an offline learning framework that utilizes offline data at scale (e.g, logs of human interactions) to improve LLM-powered policies without finetuning. The proposed method O3D (Offline Data-driven Discovery and Distillation) automatically discovers reusable skills and distills generalizable knowledge across multiple tasks based on offline interaction data, advancing the capability of solving downstream tasks. Empirical results under two interactive decision-making benchmarks (ALFWorld and WebShop) verify that O3D can notably enhance the decision-making capabilities of LLMs through the offline discovery and distillation process, and consistently outperform baselines across various LLMs.
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
Xu, Yuancheng, Yao, Jiarui, Shu, Manli, Sun, Yanchao, Wu, Zichu, Yu, Ning, Goldstein, Tom, Huang, Furong
Vision-Language Models (VLMs) excel in generating textual responses from visual inputs, yet their versatility raises significant security concerns. This study takes the first step in exposing VLMs' susceptibility to data poisoning attacks that can manipulate responses to innocuous, everyday prompts. We introduce Shadowcast, a stealthy data poisoning attack method where poison samples are visually indistinguishable from benign images with matching texts. Shadowcast demonstrates effectiveness in two attack types. The first is Label Attack, tricking VLMs into misidentifying class labels, such as confusing Donald Trump for Joe Biden. The second is Persuasion Attack, which leverages VLMs' text generation capabilities to craft narratives, such as portraying junk food as health food, through persuasive and seemingly rational descriptions. We show that Shadowcast are highly effective in achieving attacker's intentions using as few as 50 poison samples. Moreover, these poison samples remain effective across various prompts and are transferable across different VLM architectures in the black-box setting. This work reveals how poisoned VLMs can generate convincing yet deceptive misinformation and underscores the importance of data quality for responsible deployments of VLMs. Our code is available at: https://github.com/umd-huang-lab/VLM-Poisoning.
A Survey on Transformer Compression
Tang, Yehui, Wang, Yunhe, Guo, Jianyuan, Tu, Zhijun, Han, Kai, Hu, Hailin, Tao, Dacheng
Abstract--Large models based on the Transformer architecture play increasingly vital roles in artificial intelligence, particularly within the realms of natural language processing (NLP) and computer vision (CV). Model compression methods reduce their memory and computational cost, which is a necessary step to implement the transformer models on practical devices. Given the unique architecture of transformer, featuring alternative attention and Feedforward Neural Network (FFN) modules, specific compression techniques are required. The efficiency of these compression methods is also paramount, as it is usually impractical to retrain large models on the entire training dataset. This survey provides a comprehensive review of recent compression methods, with a specific focus on their application to transformer models. The compression methods are primarily categorized into pruning, quantization, knowledge distillation, and efficient architecture design. In each category, we discuss compression methods for both CV and NLP tasks, highlighting common underlying principles. At last, we delve into the relation between various compression methods, and discuss the further directions in this domain. For example, When quantizing a full-precision model (MLP), convolutional neural network (CNN), recurrent neural (float32) into 8-bit integers, the memory cost can be reduced network (RNN), long short-term memory (LSTM), Transformers, by a factor of four. In recent times, transformer-based models have emerged as the be divided into post-training quantization(PTQ) or quantizationaware prevailing choice across various domains, including both natural training (QAT), in which the former only incurs limited language processing (NLP) and computer vision (CV) domains. Knowledge Considering their strong scaling ability, most of the large models distillation serves as a training strategy, which transfers knowledge with over billions of parameters are based on the transformer from a large model (teacher) to a smaller model (student). The architecture, which are considered as foundational elements for student mimics the behavior of the teacher by emulating the general artificial intelligence (AGI) [1], [2], [3], [4], [5], [6]. Notably, for advanced While large models have demonstrated significant capabilities, models like GPT-4, accessible only through APIs, their generated their exceptionally vast sizes pose challenges for practical instructions and explanations can also guide the learning of the development. For instance, the GPT-3 model has 175 billion student model [7], [8].In addition to obtaining models from predefined parameters and demands approximately about 350GB memory large models, some methods yield efficient architectures model storage (float16). The sheer volume of parameters and by directly reducing the computational complexity of attention the associated computational expenses necessitate devices with modules or FFN modules.
Personalized Language Modeling from Personalized Human Feedback
Li, Xinyu, Lipton, Zachary C., Leqi, Liu
Reinforcement Learning from Human Feedback (RLHF) is the current dominating framework to fine-tune large language models to better align with human preferences. However, the underlying premise of algorithms developed under this framework can be problematic when user preferences encoded in human feedback are diverse. In this work, we aim to address this problem by developing methods for building personalized language models. We first formally introduce the task of learning from personalized human feedback and explain why vanilla RLHF can be problematic in this context. We then propose a general Personalized-RLHF (P-RLHF) framework, which requires one to jointly learn a user model and a language (or reward) model. The user model takes in user information and outputs user representations. Its structure encodes our assumptions about user preferences underlying the feedback data. We develop new learning objectives for personalized reward modeling and personalized Direct Preference Optimization. To demonstrate the efficacy of our method, we test it on real-world text summarization data with annotated preferences and annotator information. We fine-tune GPT-J 6B to obtain personalized language (and reward) models, which outperform non-personalized models in terms of aligning with individual preferences.
Enhancing Textbook Question Answering Task with Large Language Models and Retrieval Augmented Generation
Alawwad, Hessa Abdulrahman, Alhothali, Areej, Naseem, Usman, Alkhathlan, Ali, Jamal, Amani
Textbook question answering (TQA) is a challenging task in artificial intelligence due to the complex nature of context and multimodal data. Although previous research has significantly improved the task, there are still some limitations including the models' weak reasoning and inability to capture contextual information in the lengthy context. The introduction of large language models (LLMs) has revolutionized the field of AI, however, directly applying LLMs often leads to inaccurate answers. This paper proposes a methodology that handle the out-of-domain scenario in TQA where concepts are spread across different lessons by incorporating the retrieval augmented generation (RAG) technique and utilize transfer learning to handle the long context and enhance reasoning abilities. Through supervised fine-tuning of the LLM model Llama-2 and the incorporation of RAG, our architecture outperforms the baseline, achieving a 4.12% accuracy improvement on validation set and 9.84% on test set for non-diagram multiple-choice questions.
PRES: Toward Scalable Memory-Based Dynamic Graph Neural Networks
Su, Junwei, Zou, Difan, Wu, Chuan
Memory-based Dynamic Graph Neural Networks (MDGNNs) are a family of dynamic graph neural networks that leverage a memory module to extract, distill, and memorize long-term temporal dependencies, leading to superior performance compared to memory-less counterparts. However, training MDGNNs faces the challenge of handling entangled temporal and structural dependencies, requiring sequential and chronological processing of data sequences to capture accurate temporal patterns. During the batch training, the temporal data points within the same batch will be processed in parallel, while their temporal dependencies are neglected. This issue is referred to as temporal discontinuity and restricts the effective temporal batch size, limiting data parallelism and reducing MDGNNs' flexibility in industrial applications. This paper studies the efficient training of MDGNNs at scale, focusing on the temporal discontinuity in training MDGNNs with large temporal batch sizes. We first conduct a theoretical study on the impact of temporal batch size on the convergence of MDGNN training. Based on the analysis, we propose PRES, an iterative prediction-correction scheme combined with a memory coherence learning objective to mitigate the effect of temporal discontinuity, enabling MDGNNs to be trained with significantly larger temporal batches without sacrificing generalization performance. Experimental results demonstrate that our approach enables up to a 4x larger temporal batch (3.4x speed-up) during MDGNN training.