Spatial Reasoning
Improved distinct bone segmentation in upper-body CT through multi-resolution networks
Schnider, Eva, Wolleb, Julia, Huck, Antal, Toranelli, Mireille, Rauter, Georg, Müller-Gerbl, Magdalena, Cattin, Philippe C.
Purpose: Automated distinct bone segmentation from CT scans is widely used in planning and navigation workflows. U-Net variants are known to provide excellent results in supervised semantic segmentation. However, in distinct bone segmentation from upper body CTs a large field of view and a computationally taxing 3D architecture are required. This leads to low-resolution results lacking detail or localisation errors due to missing spatial context when using high-resolution inputs. Methods: We propose to solve this problem by using end-to-end trainable segmentation networks that combine several 3D U-Nets working at different resolutions. Our approach, which extends and generalizes HookNet and MRN, captures spatial information at a lower resolution and skips the encoded information to the target network, which operates on smaller high-resolution inputs. We evaluated our proposed architecture against single resolution networks and performed an ablation study on information concatenation and the number of context networks. Results: Our proposed best network achieves a median DSC of 0.86 taken over all 125 segmented bone classes and reduces the confusion among similar-looking bones in different locations. These results outperform our previously published 3D U-Net baseline results on the task and distinct-bone segmentation results reported by other groups. Conclusion: The presented multi-resolution 3D U-Nets address current shortcomings in bone segmentation from upper-body CT scans by allowing for capturing a larger field of view while avoiding the cubic growth of the input pixels and intermediate computations that quickly outgrow the computational capacities in 3D. The approach thus improves the accuracy and efficiency of distinct bone segmentation from upper-body CT.
Continual Graph Learning: A Survey
Yuan, Qiao, Guan, Sheng-Uei, Ni, Pin, Luo, Tianlun, Man, Ka Lok, Wong, Prudence, Chang, Victor
Research on continual learning (CL) mainly focuses on data represented in the Euclidean space, while research on graph-structured data is scarce. Furthermore, most graph learning models are tailored for static graphs. However, graphs usually evolve continually in the real world. Catastrophic forgetting also emerges in graph learning models when being trained incrementally. This leads to the need to develop robust, effective and efficient continual graph learning approaches. Continual graph learning (CGL) is an emerging area aiming to realize continual learning on graph-structured data. This survey is written to shed light on this emerging area. It introduces the basic concepts of CGL and highlights two unique challenges brought by graphs. Then it reviews and categorizes recent state-of-the-art approaches, analyzing their strategies to tackle the unique challenges in CGL. Besides, it discusses the main concerns in each family of CGL methods, offering potential solutions. Finally, it explores the open issues and potential applications of CGL.
Rethinking Dimensionality Reduction in Grid-based 3D Object Detection
Huang, Dihe, Chen, Ying, Ding, Yikang, Liao, Jinli, Liu, Jianlin, Wu, Kai, Nie, Qiang, Liu, Yong, Wang, Chengjie, Li, Zhiheng
Bird's eye view (BEV) is widely adopted by most of the current point cloud detectors due to the applicability of well-explored 2D detection techniques. However, existing methods obtain BEV features by simply collapsing voxel or point features along the height dimension, which causes the heavy loss of 3D spatial information. To alleviate the information loss, we propose a novel point cloud detection network based on a Multi-level feature dimensionality reduction strategy, called MDRNet. In MDRNet, the Spatial-aware Dimensionality Reduction (SDR) is designed to dynamically focus on the valuable parts of the object during voxel-to-BEV feature transformation. Furthermore, the Multi-level Spatial Residuals (MSR) is proposed to fuse the multi-level spatial information in the BEV feature maps. Extensive experiments on nuScenes show that the proposed method outperforms the state-of-the-art methods. The code will be available upon publication.
Working with Winding Numbers part2(Algebraic topology)
Abstract: Topological invariance is a powerful concept in different branches of physics as they are particularly robust under perturbations. We generalize the ideas of computing the statistics of winding numbers for a specific parametric model of the chiral Gaussian Unitary Ensemble to other chiral random matrix ensembles. Especially, we address the two chiral symmetry classes, unitary (AIII) and symplectic (CII), and we analytically compute ensemble averages for ratios of determinants with parametric dependence. Abstract: The winding number is a concept in complex analysis which has, in the presence of chiral symmetry, a physics interpretation as the topological index belonging to gapped phases of fermions. We study statistical properties of this topological quantity.
Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel Semantics
Jiang, Jiawei, Pan, Dayan, Ren, Houxing, Jiang, Xiaohan, Li, Chao, Wang, Jingyuan
Trajectory Representation Learning (TRL) is a powerful tool for spatial-temporal data analysis and management. TRL aims to convert complicated raw trajectories into low-dimensional representation vectors, which can be applied to various downstream tasks, such as trajectory classification, clustering, and similarity computation. Existing TRL works usually treat trajectories as ordinary sequence data, while some important spatial-temporal characteristics, such as temporal regularities and travel semantics, are not fully exploited. To fill this gap, we propose a novel Self-supervised trajectory representation learning framework with TemporAl Regularities and Travel semantics, namely START. The proposed method consists of two stages. The first stage is a Trajectory Pattern-Enhanced Graph Attention Network (TPE-GAT), which converts the road network features and travel semantics into representation vectors of road segments. The second stage is a Time-Aware Trajectory Encoder (TAT-Enc), which encodes representation vectors of road segments in the same trajectory as a trajectory representation vector, meanwhile incorporating temporal regularities with the trajectory representation. Moreover, we also design two self-supervised tasks, i.e., span-masked trajectory recovery and trajectory contrastive learning, to introduce spatial-temporal characteristics of trajectories into the training process of our START framework. The effectiveness of the proposed method is verified by extensive experiments on two large-scale real-world datasets for three downstream tasks. The experiments also demonstrate that our method can be transferred across different cities to adapt heterogeneous trajectory datasets.
Multitask Weakly Supervised Learning for Origin Destination Travel Time Estimation
Wang, Hongjun, Zhang, Zhiwen, Fan, Zipei, Chen, Jiyuan, Zhang, Lingyu, Shibasaki, Ryosuke, Song, Xuan
Travel time estimation from GPS trips is of great importance to order duration, ridesharing, taxi dispatching, etc. However, the dense trajectory is not always available due to the limitation of data privacy and acquisition, while the origin destination (OD) type of data, such as NYC taxi data, NYC bike data, and Capital Bikeshare data, is more accessible. To address this issue, this paper starts to estimate the OD trips travel time combined with the road network. Subsequently, a Multitask Weakly Supervised Learning Framework for Travel Time Estimation (MWSL TTE) has been proposed to infer transition probability between roads segments, and the travel time on road segments and intersection simultaneously. Technically, given an OD pair, the transition probability intends to recover the most possible route. And then, the output of travel time is equal to the summation of all segments' and intersections' travel time in this route. A novel route recovery function has been proposed to iteratively maximize the current route's co occurrence probability, and minimize the discrepancy between routes' probability distribution and the inverse distribution of routes' estimation loss. Moreover, the expected log likelihood function based on a weakly supervised framework has been deployed in optimizing the travel time from road segments and intersections concurrently. We conduct experiments on a wide range of real world taxi datasets in Xi'an and Chengdu and demonstrate our method's effectiveness on route recovery and travel time estimation.
Spatial-Temporal Meta-path Guided Explainable Crime Prediction
Sun, Yuting, Chen, Tong, Yin, Hongzhi
Exposure to crime and violence can harm individuals' quality of life and the economic growth of communities. In light of the rapid development in machine learning, there is a rise in the need to explore automated solutions to prevent crimes. With the increasing availability of both fine-grained urban and public service data, there is a recent surge in fusing such cross-domain information to facilitate crime prediction. By capturing the information about social structure, environment, and crime trends, existing machine learning predictive models have explored the dynamic crime patterns from different views. However, these approaches mostly convert such multi-source knowledge into implicit and latent representations (e.g., learned embeddings of districts), making it still a challenge to investigate the impacts of explicit factors for the occurrences of crimes behind the scenes. In this paper, we present a Spatial-Temporal Metapath guided Explainable Crime prediction (STMEC) framework to capture dynamic patterns of crime behaviours and explicitly characterize how the environmental and social factors mutually interact to produce the forecasts. Extensive experiments show the superiority of STMEC compared with other advanced spatiotemporal models, especially in predicting felonies (e.g., robberies and assaults with dangerous weapons).
Discover New Roads with Bing Maps
The Microsoft Maps AI Team has detected 47.8M km of all roads and 1.16M km of missing roads from Open Street Maps (OSM). These new roads were detected using Bing Maps imagery collected between 2020 and 2022 including sources from both Maxar and Airbus. The complete set of roads is now available to Bing Maps Users but also shared on Github with the open data community and is freely available for download and use under the Open Data Commons Open Database License (ODbL). The Microsoft Maps AI team's network was based on UNet and ResNet and the following papers [U-Net] (https://arxiv.org/abs/1505.04597), The model was trained using the Keras toolkit with Bing Maps 512x512 images, it is fully convolutional, which allows images of any size (that are divisible by 64) to be processed by the model (constrained to 1088x1088 with 100 cm/pixel resolution by GPU (graphics processing unit) memory in this instance).
Circular Accessible Depth: A Robust Traversability Representation for UGV Navigation
Xie, Shikuan, Song, Ran, Zhao, Yuenan, Huang, Xueqin, Li, Yibin, Zhang, Wei
In this paper, we present the Circular Accessible Depth (CAD), a robust traversability representation for an unmanned ground vehicle (UGV) to learn traversability in various scenarios containing irregular obstacles. To predict CAD, we propose a neural network, namely CADNet, with an attention-based multi-frame point cloud fusion module, Stability-Attention Module (SAM), to encode the spatial features from point clouds captured by LiDAR. CAD is designed based on the polar coordinate system and focuses on predicting the border of traversable area. Since it encodes the spatial information of the surrounding environment, which enables a semi-supervised learning for the CADNet, and thus desirably avoids annotating a large amount of data. Extensive experiments demonstrate that CAD outperforms baselines in terms of robustness and precision. We also implement our method on a real UGV and show that it performs well in real-world scenarios.