Spatial Reasoning
STARN-GAT: A Multi-Modal Spatio-Temporal Graph Attention Network for Accident Severity Prediction
Nobin, Pritom Ray, Rifat, Imran Ahammad
Accurate prediction of traffic accident severity is critical for improving road safety, optimizing emergency response strategies, and informing the design of safer transportation infrastructure. However, existing approaches often struggle to effectively model the intricate interdependencies among spatial, temporal, and contextual variables that govern accident outcomes. In this study, we introduce STARN-GAT, a Multi-Modal Spatio-Temporal Graph Attention Network, which leverages adaptive graph construction and modality-aware attention mechanisms to capture these complex relationships. Unlike conventional methods, STARN-GAT integrates road network topology, temporal traffic patterns, and environmental context within a unified attention-based framework. The model is evaluated on the Fatality Analysis Reporting System (FARS) dataset, achieving a Macro F1-score of 85 percent, ROC-AUC of 0.91, and recall of 81 percent for severe incidents. To ensure generalizability within the South Asian context, STARN-GAT is further validated on the ARI-BUET traffic accident dataset, where it attains a Macro F1-score of 0.84, recall of 0.78, and ROC-AUC of 0.89. These results demonstrate the model's effectiveness in identifying high-risk cases and its potential for deployment in real-time, safety-critical traffic management systems. Furthermore, the attention-based architecture enhances interpretability, offering insights into contributing factors and supporting trust in AI-assisted decision-making. Overall, STARN-GAT bridges the gap between advanced graph neural network techniques and practical applications in road safety analytics.
Digital Twin Channel-Enabled Online Resource Allocation for 6G: Principle, Architecture and Application
Li, Tongjie, Zhang, Jianhua, Yu, Li, Zhang, Yuxiang, Cai, Yunlong, Xu, Fan, Liu, Guangyi
The emergence of sixth-generation (6G) networks is reshaping wireless communications to support mission-critical applications such as the Industrial Internet of Things (IIoT), autonomous driving, and smart manufacturing. Compared with 5G, 6G imposes significantly more stringent requirements on latency, reliability, adaptability, and end-to-end responsiveness [1, 2]. IIoT scenarios are particularly challenging due to the coexistence of complex radio propagation conditions and diverse service requirements. Dense deployments, metallic scatterers, and dynamic obstacles give rise to severe multipath fading, especially in high-frequency bands such as mmWave and terahertz, where signal stability is highly sensitive to physical structures [3, 4]. In parallel, service demands span multiple categories, such as periodic sensing, closed-loop control, event-triggered communication, and edge computing, each with distinct quality-of-service (QoS) requirements [5]. To address these multifaceted challenges, resource allocation mechanisms must be environment-aware, latency-sensitive, and capable of online adaptation across large-scale, dynamic deployments. Artificial intelligence (AI)-driven resource allocation has attracted growing interest due to its ability to learn underlying correlations from sensing data and historical records.
NAICS-Aware Graph Neural Networks for Large-Scale POI Co-visitation Prediction: A Multi-Modal Dataset and Methodology
Alrubyli, Yazeed, Alomeir, Omar, Wafa, Abrar, Hidvรฉgi, Diรกna, Alrasheed, Hend, Bahrami, Mohsen
Understanding where people go after visiting one business is crucial for urban planning, retail analytics, and location-based services. However, predicting these co-visitation patterns across millions of venues remains challenging due to extreme data sparsity and the complex interplay between spatial proximity and business relationships. Traditional approaches using only geographic distance fail to capture why coffee shops attract different customer flows than fine dining restaurants, even when co-located. We introduce NAICS-aware GraphSAGE, a novel graph neural network that integrates business taxonomy knowledge through learnable embeddings to predict population-scale co-visitation patterns. Our key insight is that business semantics, captured through detailed industry codes, provide crucial signals that pure spatial models cannot explain. The approach scales to massive datasets (4.2 billion potential venue pairs) through efficient state-wise decomposition while combining spatial, temporal, and socioeconomic features in an end-to-end framework. Evaluated on our POI-Graph dataset comprising 94.9 million co-visitation records across 92,486 brands and 48 US states, our method achieves significant improvements over state-of-the-art baselines: the R-squared value increases from 0.243 to 0.625 (a 157 percent improvement), with strong gains in ranking quality (32 percent improvement in NDCG at 10).
Co-Win: Joint Object Detection and Instance Segmentation in LiDAR Point Clouds via Collaborative Window Processing
Li, Haichuan, Westerlund, Tomi
Accurate perception and scene understanding in complex urban environments is a critical challenge for ensuring safe and efficient autonomous navigation. In this paper, we present Co-Win, a novel bird's eye view (BEV) perception framework that integrates point cloud encoding with efficient parallel window-based feature extraction to address the multi-modality inherent in environmental understanding. Our method employs a hierarchical architecture comprising a specialized encoder, a window-based backbone, and a query-based decoder head to effectively capture diverse spatial features and object relationships. Unlike prior approaches that treat perception as a simple regression task, our framework incorporates a variational approach with mask-based instance segmentation, enabling fine-grained scene decomposition and understanding. The Co-Win architecture processes point cloud data through progressive feature extraction stages, ensuring that predicted masks are both data-consistent and contextually relevant. Furthermore, our method produces interpretable and diverse instance predictions, enabling enhanced downstream decision-making and planning in autonomous driving systems.
Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting
Miao, Xingyu, Duan, Haoran, Qian, Quanhao, Wang, Jiuniu, Long, Yang, Shao, Ling, Zhao, Deli, Xu, Ran, Zhang, Gongjie
Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of large-scale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline that converts single-view images into comprehensive, scale- and appearance-realistic 3D representations - including point clouds, camera poses, depth maps, and pseudo-RGBD - via integrated depth estimation, camera calibration, and scale calibration. Our method bridges the gap between the vast repository of imagery and the increasing demand for spatial scene understanding. By automatically generating authentic, scale-aware 3D data from images, we significantly reduce data collection costs and open new avenues for advancing spatial intelligence. We release two generated spatial datasets, i.e., COCO-3D and Objects365-v2-3D, and demonstrate through extensive experiments that our generated data can benefit various 3D tasks, ranging from fundamental perception to MLLM-based reasoning. These results validate our pipeline as an effective solution for developing AI systems capable of perceiving, understanding, and interacting with physical environments.
Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract Representations
Liu, Xiaoan, Jia, Difan, Liu, Xianhao Carton, Gonzalez-Franco, Mar, Zhu-Tian, Chen
Interacting with real-world objects in Mixed Reality (MR) often proves difficult when they are crowded, distant, or partially occluded, hindering straightforward selection and manipulation. We observe that these difficulties stem from performing interaction directly on physical objects, where input is tightly coupled to their physical constraints. Our key insight is to decouple interaction from these constraints by introducing proxies-abstract representations of real-world objects. We embody this concept in Reality Proxy, a system that seamlessly shifts interaction targets from physical objects to their proxies during selection. Beyond facilitating basic selection, Reality Proxy uses AI to enrich proxies with semantic attributes and hierarchical spatial relationships of their corresponding physical objects, enabling novel and previously cumbersome interactions in MR - such as skimming, attribute-based filtering, navigating nested groups, and complex multi object selections - all without requiring new gestures or menu systems. We demonstrate Reality Proxy's versatility across diverse scenarios, including office information retrieval, large-scale spatial navigation, and multi-drone control. An expert evaluation suggests the system's utility and usability, suggesting that proxy-based abstractions offer a powerful and generalizable interaction paradigm for future MR systems.
Sparser2Sparse: Single-shot Sparser-to-Sparse Learning for Spatial Transcriptomics Imputation with Natural Image Co-learning
Fang, Yaoyu, Qian, Jiahe, Wang, Xinkun, Cooper, Lee A., Zhou, Bo
Spatial transcriptomics (ST) has revolutionized biomedical research by enabling high resolution gene expression profiling within tissues. However, the high cost and scarcity of high resolution ST data remain significant challenges. We present Single-shot Sparser-to-Sparse (S2S-ST), a novel framework for accurate ST imputation that requires only a single and low-cost sparsely sampled ST dataset alongside widely available natural images for co-training. Our approach integrates three key innovations: (1) a sparser-to-sparse self-supervised learning strategy that leverages intrinsic spatial patterns in ST data, (2) cross-domain co-learning with natural images to enhance feature representation, and (3) a Cascaded Data Consistent Imputation Network (CDCIN) that iteratively refines predictions while preserving sampled gene data fidelity. Extensive experiments on diverse tissue types, including breast cancer, liver, and lymphoid tissue, demonstrate that our method outperforms state-of-the-art approaches in imputation accuracy. By enabling robust ST reconstruction from sparse inputs, our framework significantly reduces reliance on costly high resolution data, facilitating potential broader adoption in biomedical research and clinical applications. Keywords: Spatial Transcriptomics, Gene Expression Imputation, Single-shot Learning, Natural Image Co-training, Cost Reduction 1. Introduction Spatial transcriptomics (ST) is a cutting-edge technology that enables the investigation of spatially resolved gene expression within tissues (Asp et al., 2020). Traditional transcriptomic approaches, such as single-cell RNA sequencing (scRNA-seq), provide high-throughput, high resolution gene expression profiles but inherently lack spatial context (Aung et al., 2024; Boe et al., 2024; Sankar et al., 2024). However, spatial information is crucial for identifying disease biomarkers, understanding disease progression, and developing personalized treatment strategies.
MSGM: A Multi-Scale Spatiotemporal Graph Mamba for EEG Emotion Recognition
Liu, Hanwen, Gong, Yifeng, Yan, Zuwei, Zhuang, Zeheng, Lu, Jiaxuan
--EEG-based emotion recognition struggles with capturing multi-scale spatiotemporal dynamics and ensuring computational efficiency for real-time applications. T o overcome these challenges, we propose the Multi-Scale Spatiotemporal Graph Mamba (MSGM), a novel framework integrating multi-window temporal segmentation, bimodal spatial graph modeling, and efficient fusion via the Mamba architecture. A multi-depth Graph Convolutional Network (GCN) and token embedding fusion module, paired with Mamba's state-space modeling, enable dynamic spatiotemporal interaction at linear complexity. MOTION recognition has emerged as a critical research frontier with far-reaching implications for human-computer interaction, mental health monitoring, and neurosci-entific exploration [1] [2] [3]. The ability to decode emotional states in real-time promises to revolutionize intelligent systems by enhancing user adaptability and bolstering clinical applications through early detection and management of emotional disorders [4] [5]. As these capabilities become increasingly vital in healthcare and artificial intelligence, there is an urgent need for robust, efficient, and neurophysiologically grounded approaches to overcome both theoretical complexities and practical deployment challenges [6]. Electroencephalography (EEG) stands out as a premier modality for emotion recognition, owing to its unparalleled capacity to non-invasively record brain activity with high temporal resolution, directly capturing the neural signatures of emotional processes [7]. Hanwen Liu and Yifeng Gong are with the School of Electronics and Communication Engineering, Sun Y at-sen University, Shenzhen, 518107, China, e-mail: (liuhw56, gongyf9)@mail2.sysu.edu.cn. Zuwei Y an is with the College of Communication Engineering, Jilin University, Changchun, 130012, China, e-mail: yanzw2422@mails.jlu.edu.cn.
Spatio-Temporal Demand Prediction for Food Delivery Using Attention-Driven Graph Neural Networks
Bhat, Rabia Latief, Gillani, Iqra Altaf
Accurate demand forecasting is critical for enhancing the efficiency and responsiveness of food delivery platforms, where spatial heterogeneity and temporal fluctuations in order volumes directly influence operational decisions. This paper proposes an attention-based Graph Neural Network framework that captures spatial-temporal dependencies by modeling the food delivery environment as a graph. In this graph, nodes represent urban delivery zones, while edges reflect spatial proximity and inter-regional order flow patterns derived from historical data. The attention mechanism dynamically weighs the influence of neighboring zones, enabling the model to focus on the most contextually relevant areas during prediction. Temporal trends are jointly learned alongside spatial interactions, allowing the model to adapt to evolving demand patterns. Extensive experiments on real-world food delivery datasets demonstrate the superiority of the proposed model in forecasting future order volumes with high accuracy. The framework offers a scalable and adaptive solution to support proactive fleet positioning, resource allocation, and dispatch optimization in urban food delivery operations.
Topological Social Choice: Designing a Noise-Robust Polar Distance for Persistence Diagrams
Andrikopoulos, Athanasios, Sampanis, Nikolaos
Topological Data Analysis (TDA) has emerged as a powerful framework for extracting robust and interpretable features from noisy high-dimensional data. In the context of Social Choice Theory, where preference profiles and collective decisions are geometrically rich yet sensitive to perturbations, TDA remains largely unexplored. This work introduces a novel conceptual bridge between these domains by proposing a new metric framework for persistence diagrams tailored to noisy preference data.We define a polar coordinate-based distance that captures both the magnitude and orientation of topological features in a smooth and differentiable manner. Our metric addresses key limitations of classical distances, such as bottleneck and Wasserstein, including instability under perturbation, lack of continuity, and incompatibility with gradient-based learning. The resulting formulation offers improved behavior in both theoretical and applied settings.To the best of our knowledge, this is the first study to systematically apply persistent homology to social choice systems, providing a mathematically grounded method for comparing topological summaries of voting structures and preference dynamics. We demonstrate the superiority of our approach through extensive experiments, including robustness tests and supervised learning tasks, and we propose a modular pipeline for building predictive models from online preference data. This work contributes a conceptually novel and computationally effective tool to the emerging interface of topology and decision theory, opening new directions in interpretable machine learning for political and economic systems.