BLADE: Box-Level Supervised Amodal Segmentation through Directed Expansion

Zhaochen Liu1,2*, Zhixuan Li3*, Tingting Jiang1,4†
1National Engineering Research Center of Visual Technology, National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University
2AI Innovation Center, School of Computer Science, Peking University
3School of Computer Science and Engineering, Nanyang Technological University
4National Biomedical Imaging Center, Peking University

*Equal contribution. Corresponding author.

AAAI 2024
Illustration of the overlapping region used by BLADE.

BLADE identifies the overlapping region of an object from intersecting amodal bounding boxes. This region contains the possible occluded portion and provides a practical cue for directing expansion from a visible mask to its complete amodal shape.

Abstract

Perceiving the complete shape of occluded objects is essential for human and machine intelligence. While the amodal segmentation task is to predict the complete mask of partially occluded objects, it is time-consuming and labor-intensive to annotate pixel-level ground-truth amodal masks. Box-level supervised amodal segmentation addresses this challenge by relying solely on ground-truth bounding boxes and instance classes as supervision. Nevertheless, current box-level methodologies generate low-resolution masks and imprecise boundaries. We introduce a directed expansion approach from visible masks to corresponding amodal masks. Our hybrid end-to-end network applies distinct segmentation strategies to overlapping and non-overlapping regions. An elaborately designed connectivity loss guides expansion in overlapping regions by leveraging correlations with visible masks. Experiments on several challenging datasets show that BLADE outperforms existing state-of-the-art methods by large margins.

The Proposed BLADE Approach

Schematic overview of the BLADE architecture.

BLADE uses three dynamically generated instance-aware branches. The visible branch predicts the visible mask, the amodal branch predicts a coarse amodal mask, and the region branch estimates the overlapping region. Visible-mask cues and the proposed connectivity loss direct expansion in the amodal branch. The final output uses the coarse amodal prediction inside the estimated overlapping region and the visible prediction elsewhere.

Connectivity Loss

Illustration of the connectivity loss in BLADE.

The connectivity loss combines a neighbor loss and a uniform loss. Neighbor loss encourages local label consistency around predicted visible pixels in the overlapping region, while uniform loss maintains consistency between corresponding visible and amodal predictions. Together with projection and pairwise losses, these terms create a balanced expansion process that avoids both under-expansion and over-expansion.

Qualitative Results

Qualitative comparison of BLADE with box-level supervised baselines.

BLADE predicts higher-resolution amodal masks with more complete occluded regions and smoother boundaries than prior box-level supervised approaches across synthetic and real occlusion examples.

BibTeX

@inproceedings{liu2024blade,
  title={{BLADE}: Box-Level Supervised Amodal Segmentation through Directed Expansion},
  author={Liu, Zhaochen and Li, Zhixuan and Jiang, Tingting},
  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
  pages={3846--3854},
  year={2024}
}