Overview
Foundation Models for CPS-IoT: Opportunities and Challenges
Baris, Ozan, Chen, Yizhuo, Dong, Gaofeng, Han, Liying, Kimura, Tomoyoshi, Quan, Pengrui, Wang, Ruijie, Wang, Tianchen, Abdelzaher, Tarek, Bergés, Mario, Liang, Paul Pu, Srivastava, Mani
Methods from machine learning (ML) have transformed the implementation of Perception-Cognition-Communication-Action loops in Cyber-Physical Systems (CPS) and the Internet of Things (IoT), replacing mechanistic and basic statistical models with those derived from data. However, the first generation of ML approaches, which depend on supervised learning with annotated data to create task-specific models, faces significant limitations in scaling to the diverse sensor modalities, deployment configurations, application tasks, and operating dynamics characterizing real-world CPS-IoT systems. The success of task-agnostic foundation models (FMs), including multimodal large language models (LLMs), in addressing similar challenges across natural language, computer vision, and human speech has generated considerable enthusiasm for and exploration of FMs and LLMs as flexible building blocks in CPS-IoT analytics pipelines, promising to reduce the need for costly task-specific engineering. Nonetheless, a significant gap persists between the current capabilities of FMs and LLMs in the CPS-IoT domain and the requirements they must meet to be viable for CPS-IoT applications. In this paper, we analyze and characterize this gap through a thorough examination of the state of the art and our research, which extends beyond it in various dimensions. Based on the results of our analysis and research, we identify essential desiderata that CPS-IoT domain-specific FMs and LLMs must satisfy to bridge this gap. We also propose actions by CPS-IoT researchers to collaborate in developing key community resources necessary for establishing FMs and LLMs as foundational tools for the next generation of CPS-IoT systems.
DOC-Depth: A novel approach for dense depth ground truth generation
de Moreau, Simon, Corsia, Mathias, Bouchiba, Hassan, Almehio, Yasser, Bursuc, Andrei, El-Idrissi, Hafid, Moutarde, Fabien
Accurate depth information is essential for many computer vision applications. Yet, no available dataset recording method allows for fully dense accurate depth estimation in a large scale dynamic environment. In this paper, we introduce DOC-Depth, a novel, efficient and easy-to-deploy approach for dense depth generation from any LiDAR sensor. After reconstructing consistent dense 3D environment using LiDAR odometry, we address dynamic objects occlusions automatically thanks to DOC, our state-of-the art dynamic object classification method. Additionally, DOC-Depth is fast and scalable, allowing for the creation of unbounded datasets in terms of size and time. We demonstrate the effectiveness of our approach on the KITTI dataset, improving its density from 16.1% to 71.2% and release this new fully dense depth annotation, to facilitate future research in the domain. We also showcase results using various LiDAR sensors and in multiple environments. All software components are publicly available for the research community.
A Decade of Action Quality Assessment: Largest Systematic Survey of Trends, Challenges, and Future Directions
Yin, Hao, Parmar, Paritosh, Xu, Daoliang, Zhang, Yang, Zheng, Tianyou, Fu, Weiwei
Action Quality Assessment (AQA) -- the ability to quantify the quality of human motion, actions, or skill levels and provide feedback -- has far-reaching implications in areas such as low-cost physiotherapy, sports training, and workforce development. As such, it has become a critical field in computer vision & video understanding over the past decade. Significant progress has been made in AQA methodologies, datasets, & applications, yet a pressing need remains for a comprehensive synthesis of this rapidly evolving field. In this paper, we present a thorough survey of the AQA landscape, systematically reviewing over 200 research papers using the preferred reporting items for systematic reviews & meta-analyses (PRISMA) framework. We begin by covering foundational concepts & definitions, then move to general frameworks & performance metrics, & finally discuss the latest advances in methodologies & datasets. This survey provides a detailed analysis of research trends, performance comparisons, challenges, & future directions. Through this work, we aim to offer a valuable resource for both newcomers & experienced researchers, promoting further exploration & progress in AQA. Data are available at https://haoyin116.github.io/Survey_of_AQA/
Towards Fast Graph Generation via Autoregressive Noisy Filtration Modeling
Krimmel, Markus, Wiens, Jenna, Borgwardt, Karsten, Chen, Dexiong
Graph generative models often face a critical trade-off between learning complex distributions and achieving fast generation speed. We introduce Autoregressive Noisy Filtration Modeling (ANFM), a novel approach that addresses both challenges. ANFM leverages filtration, a concept from topological data analysis, to transform graphs into short sequences of monotonically increasing subgraphs. This formulation extends the sequence families used in previous autoregressive models. To learn from these sequences, we propose a novel autoregressive graph mixer model. Our experiments suggest that exposure bias might represent a substantial hurdle in autoregressive graph generation and we introduce two mitigation strategies to address it: noise augmentation and a reinforcement learning approach. Incorporating these techniques leads to substantial performance gains, making ANFM competitive with state-of-the-art diffusion models across diverse synthetic and real-world datasets. Notably, ANFM produces remarkably short sequences, achieving a 100-fold speedup in generation time compared to diffusion models. This work marks a significant step toward high-throughput graph generation.
Deep Learning-Based Facial Expression Recognition for the Elderly: A Systematic Review
Gaya-Morey, F. Xavier, Buades-Rubio, Jose M., Palanque, Philippe, Lacuesta, Raquel, Manresa-Yee, Cristina
The rapid aging of the global population has highlighted the need for technologies to support elderly, particularly in healthcare and emotional well-being. Facial expression recognition (FER) systems offer a non-invasive means of monitoring emotional states, with applications in assisted living, mental health support, and personalized care. This study presents a systematic review of deep learning-based FER systems, focusing on their applications for the elderly population. Following a rigorous methodology, we analyzed 31 studies published over the last decade, addressing challenges such as the scarcity of elderly-specific datasets, class imbalances, and the impact of age-related facial expression differences. Our findings show that convolutional neural networks remain dominant in FER, and especially lightweight versions for resource-constrained environments. However, existing datasets often lack diversity in age representation, and real-world deployment remains limited. Additionally, privacy concerns and the need for explainable artificial intelligence emerged as key barriers to adoption. This review underscores the importance of developing age-inclusive datasets, integrating multimodal solutions, and adopting XAI techniques to enhance system usability, reliability, and trustworthiness. We conclude by offering recommendations for future research to bridge the gap between academic progress and real-world implementation in elderly care.
LLMs can be easily Confused by Instructional Distractions
Hwang, Yerin, Kim, Yongil, Koo, Jahyun, Kang, Taegwan, Bae, Hyunkyung, Jung, Kyomin
Despite the fact that large language models (LLMs) show exceptional skill in instruction following tasks, this strength can turn into a vulnerability when the models are required to disregard certain instructions. Instruction-following tasks typically involve a clear task description and input text containing the target data to be processed. However, when the input itself resembles an instruction, confusion may arise, even if there is explicit prompting to distinguish between the task instruction and the input. We refer to this phenomenon as instructional distraction. In this paper, we introduce a novel benchmark, named DIM-Bench, specifically designed to assess LLMs' performance under instructional distraction. The benchmark categorizes real-world instances of instructional distraction and evaluates LLMs across four instruction tasks: rewriting, proofreading, translation, and style transfer -- alongside five input tasks: reasoning, code generation, mathematical reasoning, bias detection, and question answering. Our experimental results reveal that even the most advanced LLMs are susceptible to instructional distraction, often failing to accurately follow user intent in such cases.
CH-MARL: Constrained Hierarchical Multiagent Reinforcement Learning for Sustainable Maritime Logistics
The advent of globalized trade has led to unprecedented growth in the volume and complexity of maritime logistics. As one of the most cost-effective modes of transportation, maritime shipping has become indispensable for connecting economies and supporting international trade. However, this growth comes with substantial environmental and operational challenges. The sector's heavy reliance on fossil fuels contributes significantly to global greenhouse gas (GHG) emissions, accounting for nearly 2.89% of global emissions Smith et al. [2014], [IMO]. Moreover, the International Maritime Organization (IMO) has outlined a strategy to reduce GHG emissions from international shipping by at least 50% by 2050 compared to 2008 levels, aiming for eventual decarbonization [IMO]. These ambitious targets underscore the pressing need for transformative solutions to meet regulatory requirements and societal expectations. Environmental pressures are further compounded by the intricate logistics of coordinating diverse stakeholders, including shipping companies, port authorities, and policymakers, each with unique objectives and constraints.
Adviser-Actor-Critic: Eliminating Steady-State Error in Reinforcement Learning Control
Chen, Donghe, Peng, Yubin, Zheng, Tengjie, Wang, Han, Qu, Chaoran, Cheng, Lin
High-precision control tasks present substantial Dynamic modeling is crucial for understanding robot behavior challenges for reinforcement learning (RL) algorithms, and designing control strategies. However, real-world frequently resulting in suboptimal performance systems often display nonlinear behavior, making it difficult attributed to network approximation inaccuracies to create accurate models. Additionally, the highdimensional and inadequate sample quality.These state space of robots can lead to complex interactions issues are exacerbated when the task requires the between components, further complicating control agent to achieve a precise goal state, as is common (Buşoniu et al., 2018; Zhao et al., 2020a;b; Cao et al., 2023). in robotics and other real-world applications.We To highlight these challenges, we discuss the attributes and introduce Adviser-Actor-Critic (AAC), designed limitations of existing control algorithms.
Multimodal Brain-Computer Interfaces: AI-powered Decoding Methodologies
Li, Siyang, Wang, Hongbin, Chen, Xiaoqing, Wu, Dongrui
Brain-computer interfaces (BCIs) enable direct communication between the brain and external devices. This review highlights the core decoding algorithms that enable multimodal BCIs, including a dissection of the elements, a unified view of diversified approaches, and a comprehensive analysis of the present state of the field. We emphasize algorithmic advancements in cross-modality mapping, sequential modeling, besides classic multi-modality fusion, illustrating how these novel AI approaches enhance decoding of brain data. The current literature of BCI applications on visual, speech, and affective decoding are comprehensively explored. Looking forward, we draw attention on the impact of emerging architectures like multimodal Transformers, and discuss challenges such as brain data heterogeneity and common errors. This review also serves as a bridge in this interdisciplinary field for experts with neuroscience background and experts that study AI, aiming to provide a comprehensive understanding for AI-powered multimodal BCIs.
Rethinking stance detection: A theoretically-informed research agenda for user-level inference using language models
Bhattacharya, Prasanta, Zhang, Hong, Cao, Yiming, Gao, Wei, Loh, Brandon Siyuan, Simons, Joseph J. P., Wong, Liang Ze
Stance detection has emerged as a popular task in natural language processing research, enabled largely by the abundance of target-specific social media data. While there has been considerable research on the development of stance detection models, datasets, and application, we highlight important gaps pertaining to (i) a lack of theoretical conceptualization of stance, and (ii) the treatment of stance at an individual- or user-level, as opposed to message-level. In this paper, we first review the interdisciplinary origins of stance as an individual-level construct to highlight relevant attributes (e.g., psychological features) that might be useful to incorporate in stance detection models. Further, we argue that recent pre-trained and large language models (LLMs) might offer a way to flexibly infer such user-level attributes and/or incorporate them in modelling stance. To better illustrate this, we briefly review and synthesize the emerging corpus of studies on using LLMs for inferring stance, and specifically on incorporating user attributes in such tasks. We conclude by proposing a four-point agenda for pursuing stance detection research that is theoretically informed, inclusive, and practically impactful.