Canonical Policy: Learning Canonical 3D Representation for SE(3)-Equivariant Policy
Zhang, Zhiyuan, Xu, Zhengtong, Lakamsani, Jai Nanda, She, Yu
–arXiv.org Artificial Intelligence
Abstract--Visual Imitation learning has achieved remarkable progress in robotic manipulation, yet generalization to unseen objects, scene layouts, and camera viewpoints remains a key challenge. Recent advances address this by using 3D point clouds, which provide geometry-aware, appearance-invariant representations, and by incorporating equivariance into policy architectures to exploit spatial symmetries. However, existing equivariant approaches often lack interpretability and rigor due to unstructured integration of equivariant components. We introduce canonical policy, a principled framework for 3D equiv-ariant imitation learning that unifies 3D point cloud observations under a canonical representation. We first establish a theory of 3D canonical representations, enabling equivariant observation-to-action mappings by grouping both seen and novel point clouds to a canonical representation. We then propose a flexible policy learning pipeline that leverages geometric symmetries from canonical representation and the expressiveness of modern generative models. We validate canonical policy on 12 diverse simulated tasks and 4 real-world manipulation tasks across 16 configurations, involving variations in object color, shape, camera viewpoint, and robot platform. Compared to state-of-the-art imitation learning policies, canonical policy achieves an average improvement of 18.0% in simulation and 39.7% in real-world experiments, demonstrating superior generalization capability and sample efficiency. Imitation learning has made remarkable progress in robotic manipulation in recent years [1]-[5]. However, generalization and sample efficiency remain key challenges. In particular, visual imitation learning policies [1], [2] often struggle to generalize beyond the fixed dataset of human demonstrations on which they are trained. Unseen variations in object types, scene configurations, geometric layouts, or camera viewpoints can significantly degrade policy performance. To improve generalization in visual imitation, recent research has explored the use of 3D point clouds as input observations [3], [5]-[7]. Point clouds encode the geometric structure of the environment and are invariant to visual dis-tractors such as background textures or object appearances. As a result, policies trained on 3D representations can more effectively learn the mapping between geometry of observations and robot actions, thereby improving both generalization and sample efficiency. Since robotic manipulation takes place in 3D Euclidean space and involves tasks that are often equivariant to rotations and translations, incorporating symmetry priors into the policy architecture can significantly improve both sample efficiency and generalization performance. However, existing approaches that aim to achieve equiv-ariant policy learning from 3D point clouds or 2D images often rely on re-designing the entire network using equivariant neural modules [8]-[10].
arXiv.org Artificial Intelligence
Nov-11-2025