Scinovex
article Open AccessTop 1% cited

Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges

IEEE Transactions on Intelligent Transportation Systems · 2020 · Vol. 22(3) · pp. 1341–1360
Di FengChristian SchützLars RosenbaumHeinz HertleinClaudius GläserFabian TimmW. WiesbeckKlaus Dietmayer

Abstract

Recent advancements in perception for autonomous driving are driven by deep learning. In order to achieve robust and accurate scene understanding, autonomous vehicles are usually equipped with different sensors (e.g. cameras, LiDARs, Radars), and multiple sensing modalities can be fused to exploit their complementary properties. In this context, many methods have been proposed for deep multi-modal perception problems. However, there is no general guideline for network architecture design, and questions of “what to fuse”, “when to fuse”, and “how to fuse” remain open. This review paper attempts to systematically summarize methodologies and discuss challenges for deep multi-modal object detection and semantic segmentation in autonomous driving. To this end, we first provide an overview of on-board sensors on test vehicles, open datasets, and background information for object detection and semantic segmentation in autonomous driving research. We then summarize the fusion methodologies and discuss challenges and open questions. In the appendix, we provide tables that summarize topics and methods. We also provide an interactive online platform to navigate each reference: https://boschresearch.github.io/multimodalperception/.

Advanced Neural Network ApplicationsAutonomous Vehicle Technology and SafetyVehicle License Plate RecognitionModalComputer scienceSegmentationArtificial intelligenceObject detectionComputer visionPattern recognition (psychology)

Funding

  • Robert Bosch
Citations
1,297
FWCI
66.04
field-weighted impact
References
365
Percentile
100%
vs. same field & year
Citations per year
Cited by
More Diverse Means Better: Multimodal Deep Learning Meets Remote-Sensing Imagery Classification
IEEE Transactions on Geoscience and Remote Sensing · 2020 · 1,291 citations
Deep Learning for Image and Point Cloud Fusion in Autonomous Driving: A Review
IEEE Transactions on Intelligent Transportation Systems · 2021 · 508 citations
References
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2016 · 52,930 citations
The Pascal Visual Object Classes (VOC) Challenge
International Journal of Computer Vision · 2009 · 19,127 citations
The Pascal Visual Object Classes Challenge: A Retrospective
International Journal of Computer Vision · 2014 · 7,183 citations
Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2015 · 11,231 citations
A Practical Bayesian Framework for Backpropagation Networks
Neural Computation · 1992 · 2,890 citations
ImageNet Large Scale Visual Recognition Challenge
International Journal of Computer Vision · 2015 · 39,683 citations
Adaptive Mixtures of Local Experts
Neural Computation · 1991 · 4,795 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.