GIN: Generative INvariant Shape Prior for Amodal Instance Segmentation

Zhixuan Li1, Weining Ye1, Tingting Jiang1†, Tiejun Huang1,2
1National Engineering Research Center of Visual Technology, National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University
2Beijing Academy of Artificial Intelligence

Corresponding author.

IEEE Transactions on Multimedia, 2023
Comparison between dictionary-based shape priors and the proposed invariant shape-prior learning.

Dictionary-based methods memorize a finite set of masks and cannot retrieve an unseen shape. GIN instead learns a basic shape independent of translation, rotation, and scaling, then dynamically generates a suitable shape prior for each instance.

Abstract

Amodal instance segmentation predicts the complete shape of an occluded object, including both its visible and occluded regions. Because visual clues are unavailable for the hidden region, shape-prior knowledge is especially valuable. Previous 2D approaches establish a shape dictionary and retrieve the closest stored mask, but cannot provide priors outside that dictionary. We propose the generative invariant shape-prior network (GIN), which learns a basic shape invariant to translation, rotation, and scaling. By decoupling shape-prior learning from transformation, GIN is end-to-end trainable, requires no dictionary construction, and generalizes more effectively. GIN outperforms state-of-the-art methods by large margins on D2SA, COCOA-cls, and KINS.

The Proposed GIN Approach

Overview of the GIN architecture.

GIN first extracts an instance feature with a backbone and ROI-Align. Its invariant amodal learning branch predicts a transformation, normalizes the feature, and learns an invariant shape prior. A complementary vanilla mask branch preserves residual appearance information. The feature refinement network combines both branches to produce the final amodal mask.

Qualitative Results on D2SA

Qualitative comparison of GIN with amodal segmentation baselines on D2SA.

On densely occluded supermarket objects, GIN reconstructs complete object shapes with cleaner boundaries than ORCNN, Mask-RCNN, and the dictionary-based ShapeDict method.

Generalization Across Datasets

Qualitative comparison of GIN on COCOA-cls and KINS.

Results on COCOA-cls and KINS demonstrate that the learned invariant prior transfers from everyday objects to complex street scenes, including challenging people and vehicles with severe occlusion.

BibTeX

@inproceedings{li2023gin,
  author={Li, Zhixuan and Ye, Weining and Jiang, Tingting and Huang, Tiejun},
  title={{GIN}: Generative INvariant Shape Prior for Amodal Instance Segmentation},
  booktitle={IEEE Transactions on Multimedia},
  pages={3924--3936},
  year={2023}
}