Unsupervised Semantic Segmentation by Distilling Feature Correspondences
Hamilton, Mark, Zhang, Zhoutong, Hariharan, Bharath, Snavely, Noah, Freeman, William T.
Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clusters. Unlike previous works which achieve this with a single end-to-end framework, we propose to separate feature learning from cluster compactification. Empirically, we show that current unsupervised feature learning frameworks already generate dense features whose correlations are semantically consistent. This observation motivates us to design STEGO (Self-supervised Transformer with Energy-based Graph Optimization), a novel framework that distills unsupervised features into highquality discrete semantic labels. At the core of STEGO is a novel contrastive loss function that encourages features to form compact clusters while preserving their relationships across the corpora. STEGO yields a significant improvement over the prior state of the art, on both the CocoStuff (+14 mIoU) and Cityscapes (+9 mIoU) semantic segmentation challenges. Semantic segmentation is the process of classifying each individual pixel of an image into a known ontology. Though semantic segmentation models can detect and delineate objects at a much finer granularity than classification or object detection systems, these systems are hindered by the difficulties of creating labelled training data. In particular, segmenting an image can take over 100 more effort for a human annotator than classifying or drawing bounding boxes (Zlateski et al., 2018). Furthermore, in complex domains such as medicine, biology, or astrophysics, ground-truth segmentation labels may be unknown, ill-defined, or require considerable domain-expertise to provide (Yu et al., 2018). Recently, several works introduced semantic segmentation systems that could learn from weaker forms of labels such as classes, tags, bounding boxes, scribbles, or point annotations (Ren et al., 2020; Pan et al., 2021; Liu et al., 2020; Bilen et al.). However, comparatively few works take up the challenge of semantic segmentation without any form of human supervision or motion cues.
Mar-16-2022
- Country:
- North America > United States
- Massachusetts > Middlesex County
- Cambridge (0.04)
- California > Alameda County
- Oakland (0.04)
- Massachusetts > Middlesex County
- Europe
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- Germany > Brandenburg
- Potsdam (0.04)
- United Kingdom > England
- Asia > Middle East
- Jordan (0.04)
- North America > United States
- Genre:
- Research Report (1.00)
- Industry:
- Technology: