Scinovex
articleTop 1% cited

CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation With Transformers

IEEE Transactions on Intelligent Transportation Systems · 2023 · Vol. 24(12) · pp. 14679–14694
Jiaming ZhangHuayao LiuKailun YangXinxin HuRuiping LiuRainer Stiefelhagen

Abstract

Scene understanding based on image segmentation is a crucial component of autonomous vehicles. Pixel-wise semantic segmentation of RGB images can be advanced by exploiting complementary features from the supplementary modality ( <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">${X}$ </tex-math></inline-formula> -modality). However, covering a wide variety of sensors with a modality-agnostic model remains an unresolved problem due to variations in sensor characteristics among different modalities. Unlike previous modality-specific methods, in this work, we propose a unified fusion framework, CMX, for RGB-X semantic segmentation. To generalize well across different modalities, that often include supplements as well as uncertainties, a unified cross-modal interaction is crucial for modality fusion. Specifically, we design a Cross-Modal Feature Rectification Module (CM-FRM) to calibrate bi-modal features by leveraging the features from one modality to rectify the features of the other modality. With rectified feature pairs, we deploy a Feature Fusion Module (FFM) to perform sufficient exchange of long-range contexts before mixing. To verify CMX, for the first time, we unify five modalities complementary to RGB, i.e., depth, thermal, polarization, event, and LiDAR. Extensive experiments show that CMX generalizes well to diverse multi-modal fusion, achieving state-of-the-art performances on five RGB-Depth benchmarks, as well as RGB-Thermal, RGB-Polarization, and RGB-LiDAR datasets. Besides, to investigate the generalizability to dense-sparse data fusion, we establish an RGB-Event semantic segmentation benchmark based on the EventScape dataset, on which CMX sets the new state-of-the-art. The source code of CMX is publicly available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/huaaaliu/RGBX_Semantic_</uri> Segmentation.

Advanced Neural Network ApplicationsVisual Attention and Saliency DetectionRobotics and Sensor-Based LocalizationRGB color modelArtificial intelligenceSegmentationModality (human–computer interaction)Computer scienceComputer visionLidarFeature (linguistics)Pattern recognition (psychology)Remote sensing

Funding

  • Niedersächsisches Ministerium für Wissenschaft und Kultur
  • Bundesministerium für Arbeit und Soziales
Citations
531
FWCI
63.06
field-weighted impact
References
110
Percentile
100%
vs. same field & year
Citations per year
References
ImageNet Large Scale Visual Recognition Challenge
International Journal of Computer Vision · 2015 · 39,683 citations
DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2017 · 21,645 citations
ERFNet: Efficient Residual Factorized ConvNet for Real-Time Semantic Segmentation
IEEE Transactions on Intelligent Transportation Systems · 2017 · 1,469 citations
Deep High-Resolution Representation Learning for Visual Recognition
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2020 · 4,376 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation With Transformers · Scinovex