AITopics

Genre:

Research Report > New Finding (1.00)
Research Report > Experimental Study (1.00)

Technology: Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.46)

Neural Information Processing SystemsFeb-12-2026, 01:21:57 GMT

Activating Self-Attention for Multi-Scene Absolute Pose Regression

Multi-scene absolute pose regression addresses the demand for fast and memory-efficient camera pose estimation across various real-world environments.

artificial intelligence, machine learning, natural language, (18 more...)

Country: Europe > United Kingdom > England > Tyne and Wear > Newcastle (0.04)

Genre:

Research Report > Experimental Study (0.93)
Research Report > New Finding (0.67)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Natural Language (1.00)
Information Technology > Sensing and Signal Processing > Image Processing (0.94)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.68)

Neural Information Processing SystemsOct-10-2025, 11:48:50 GMT

Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes Qi Ma

To address these challenges, we introduce "Implicit-Zoo": a

dataset, representation, transformer, (14 more...)

Country:

Europe > Switzerland > Zürich > Zürich (0.04)
Europe > Slovenia > Drava > Municipality of Benedikt > Benedikt (0.04)

Genre: Research Report (0.46)

Technology:

Information Technology > Sensing and Signal Processing > Image Processing (1.00)
Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (1.00)
(2 more...)

Neural Information Processing SystemsOct-10-2025, 00:45:31 GMT

Activating Self-Attention for Multi-Scene Absolute Pose Regression

Multi-scene absolute pose regression addresses the demand for fast and memory-efficient camera pose estimation across various real-world environments.

dataset, experiment, query region, (14 more...)

Country: Europe > United Kingdom > England > Tyne and Wear > Newcastle (0.04)

Genre:

Research Report > Experimental Study (0.93)
Research Report > New Finding (0.67)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Natural Language (1.00)
Information Technology > Sensing and Signal Processing > Image Processing (0.94)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.68)

Ravuri, Srinivas, Xu, Yuan, Zehetner, Martin Ludwig, Motlag, Ketan, Albayrak, Sahin

APR-Transformer: Initial Pose Estimation for Localization in Complex Environments through Absolute Pose Regression

arXiv.org Artificial IntelligenceMay-15-2025

Afterwards, we remove the last propagation layer and classification head and use the remaining components as backbone for our APR-Transformer. We utilize the output of the last remaining propagation layer as feature vectors F x and F q at resolutions of (N, 128, 1024). With each of the 128 vectors corresponding to a reduced point of the original 4096-point data input. The Transformer-compatible input embeddings and associated learned encodings, preserving the spatial information of the backbone outputs, are then computed by first separating the 128 vectors into eight groups based on the absolute z -coordinates, i.e., height, of their corresponding reduced points. Subsequently, the 16 feature vectors per group are sorted in a 4 4 grid based on the x and y coordinates of their reduced points. Afterward, we adapt the procedure used in the image-based APR-Transformer case by computing the learned positional encodings along the three axes to generate the final Transformer inputs. C. Pose Regression and Loss Function L p( x) = D null i =1|x p i x t i| (1) L o(q) = D null i =1|q p i q t i| (2) L pose= L p exp( s x) + s x + L o exp( s q) + s q (3) We train the model variants to minimize the position loss L p, see Equation 1, and orientation loss L o, see Equation 2, for the ground truth pose, where L p and L o are L 1 losses. We combine the position and orientation losses using the formulation by Kendall et al. [28] shown in Equation 3. Where s x and s q are learned parameters that control the balance between the position loss and the orientation loss.

artificial intelligence, deep learning, machine learning, (15 more...)

2505.09356

Country:

Europe > Germany > Berlin (0.04)
Europe > United Kingdom > England > Oxfordshire > Oxford (0.04)
Europe > Switzerland (0.04)
Asia > Middle East > Israel (0.04)

Genre: Research Report (1.00)

Industry: Transportation > Ground > Road (0.93)

Technology:

Information Technology > Artificial Intelligence > Representation & Reasoning (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Statistical Learning (0.68)

arXiv.org Artificial IntelligenceNov-28-2024

Unleashing the Power of Data Synthesis in Visual Localization

Li, Sihang, Tan, Siqi, Chang, Bowen, Zhang, Jing, Feng, Chen, Li, Yiming

Visual localization, which estimates a camera's pose within a known scene, is a long-standing challenge in vision and robotics. Recent end-to-end methods that directly regress camera poses from query images have gained attention for fast inference. However, existing methods often struggle to generalize to unseen views. In this work, we aim to unleash the power of data synthesis to promote the generalizability of pose regression. Specifically, we lift real 2D images into 3D Gaussian Splats with varying appearance and deblurring abilities, which are then used as a data engine to synthesize more posed images. To fully leverage the synthetic data, we build a two-branch joint training pipeline, with an adversarial discriminator to bridge the syn-to-real gap. Experiments on established benchmarks show that our method outperforms state-of-the-art end-to-end approaches, reducing translation and rotation errors by 50% and 21.6% on indoor datasets, and 35.56% and 38.7% on outdoor datasets. We also validate the effectiveness of our method in dynamic driving scenarios under varying weather conditions. Notably, as data synthesis scales up, our method exhibits a growing ability to interpolate and extrapolate training data for localizing unseen views. Project Page: https://ai4ce.github.io/RAP/

artificial intelligence, machine learning, proceedings, (14 more...)

2412.00138

Country:

North America > United States > New York (0.04)
Asia > Japan > Honshū > Chūbu > Ishikawa Prefecture > Kanazawa (0.04)

Genre: Research Report (0.82)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Robots (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.68)

Cheng, Yuzhou, Jiao, Jianhao, Wang, Yue, Kanoulas, Dimitrios

LoGS: Visual Localization via Gaussian Splatting with Fewer Training Images

arXiv.org Artificial IntelligenceOct-15-2024

Visual localization involves estimating a query image's 6-DoF (degrees of freedom) camera pose, which is a fundamental component in various computer vision and robotic tasks. This paper presents LoGS, a vision-based localization pipeline utilizing the 3D Gaussian Splatting (GS) technique as scene representation. This novel representation allows high-quality novel view synthesis. During the mapping phase, structure-from-motion (SfM) is applied first, followed by the generation of a GS map. During localization, the initial position is obtained through image retrieval, local feature matching coupled with a PnP solver, and then a high-precision pose is achieved through the analysis-by-synthesis manner on the GS map. Experimental results on four large-scale datasets demonstrate the proposed approach's SoTA accuracy in estimating camera poses and robustness under challenging few-shot conditions.

artificial intelligence, machine learning, proceedings, (16 more...)

2410.11505

Country:

Europe > United Kingdom > England > Greater London > London (0.04)
Europe > Greece (0.04)
Asia > China > Zhejiang Province > Hangzhou (0.04)

Genre: Research Report (0.64)

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Statistical Learning (0.46)
Information Technology > Artificial Intelligence > Vision > Image Understanding (0.41)

Sidorov, Gennady, Mohrat, Malik, Lebedeva, Ksenia, Rakhimov, Ruslan, Kolyubin, Sergey

GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splatting for Improved Visual Localization

arXiv.org Artificial IntelligenceSep-24-2024

Although various visual localization approaches exist, such as scene coordinate and pose regression, these methods often struggle with high memory consumption or extensive optimization requirements. To address these challenges, we utilize recent advancements in novel view synthesis, particularly 3D Gaussian Splatting (3DGS), to enhance localization. 3DGS allows for the compact encoding of both 3D geometry and scene appearance with its spatial features. Our method leverages the dense description maps produced by XFeat's lightweight keypoint detection and description model. We propose distilling these dense keypoint descriptors into 3DGS to improve the model's spatial understanding, leading to more accurate camera pose predictions through 2D-3D correspondences. After estimating an initial pose, we refine it using a photometric warping loss. Benchmarking on popular indoor and outdoor datasets shows that our approach surpasses state-of-the-art Neural Render Pose (NRP) methods, including NeRFMatch and PNeRFLoc.

localization, pose estimation, query image, (15 more...)

2409.16502

Country:

Asia > Russia (0.04)
Europe > Russia > Central Federal District > Moscow Oblast > Moscow (0.04)

Genre: Research Report (0.65)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Machine Learning (1.00)

Polizzi, Vincenzo, Cannici, Marco, Scaramuzza, Davide, Kelly, Jonathan

FaVoR: Features via Voxel Rendering for Camera Relocalization

arXiv.org Artificial IntelligenceSep-11-2024

Camera relocalization methods range from dense image alignment to direct camera pose regression from a query image. Among these, sparse feature matching stands out as an efficient, versatile, and generally lightweight approach with numerous applications. However, feature-based methods often struggle with significant viewpoint and appearance changes, leading to matching failures and inaccurate pose estimates. To overcome this limitation, we propose a novel approach that leverages a globally sparse yet locally dense 3D representation of 2D features. By tracking and triangulating landmarks over a sequence of frames, we construct a sparse voxel map optimized to render image patch descriptors observed during tracking. Given an initial pose estimate, we first synthesize descriptors from the voxels using volumetric rendering and then perform feature matching to estimate the camera pose. This methodology enables the generation of descriptors for unseen views, enhancing robustness to view changes. We extensively evaluate our method on the 7-Scenes and Cambridge Landmarks datasets. Our results show that our method significantly outperforms existing state-of-the-art feature representation techniques in indoor environments, achieving up to a 39% improvement in median translation error. Additionally, our approach yields comparable results to other methods for outdoor scenarios while maintaining lower memory and computational costs.

descriptor, landmark, representation, (16 more...)

2409.07571

Country:

North America > Canada > Ontario > Toronto (0.14)
North America > United States (0.04)
Europe > United Kingdom > England > Cambridgeshire > Cambridge (0.04)
(2 more...)

Genre: Research Report > New Finding (0.68)

Technology:

Information Technology > Sensing and Signal Processing > Image Processing (1.00)
Information Technology > Artificial Intelligence > Vision (0.96)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.93)