Scinovex
articleTop 1% cited

Region-Based Convolutional Networks for Accurate Object Detection and Segmentation

Ross GirshickJeff DonahueTrevor DarrellJitendra Malik

Abstract

Object detection performance, as measured on the canonical PASCAL VOC Challenge datasets, plateaued in the final years of the competition. The best-performing methods were complex ensemble systems that typically combined multiple low-level image features with high-level context. In this paper, we propose a simple and scalable detection algorithm that improves mean average precision (mAP) by more than 50 percent relative to the previous best result on VOC 2012-achieving a mAP of 62.4 percent. Our approach combines two ideas: (1) one can apply high-capacity convolutional networks (CNNs) to bottom-up region proposals in order to localize and segment objects and (2) when labeled training data are scarce, supervised pre-training for an auxiliary task, followed by domain-specific fine-tuning, boosts performance significantly. Since we combine region proposals with CNNs, we call the resulting model an R-CNN or Region-based Convolutional Network. Source code for the complete system is available at http://www.cs.berkeley.edu/~rbg/rcnn.

Advanced Neural Network ApplicationsAdvanced Image and Video Retrieval TechniquesDomain Adaptation and Few-Shot LearningPascal (unit)Computer scienceObject detectionArtificial intelligencePattern recognition (psychology)SegmentationConvolutional neural networkScalabilityImage segmentationContext (archaeology)

Funding

  • Nvidia
  • Defense Advanced Research Projects Agency
Citations
2,864
FWCI
101.85
field-weighted impact
References
114
Percentile
100%
vs. same field & year
Citations per year
Cited by
Deep Learning for Generic Object Detection: A Survey
International Journal of Computer Vision · 2019 · 2,702 citations
Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images
IEEE Transactions on Geoscience and Remote Sensing · 2016 · 1,729 citations
Fully Convolutional Networks for Semantic Segmentation
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2016 · 10,957 citations
Object Detection in 20 Years: A Survey
Proceedings of the IEEE · 2023 · 2,704 citations
Deep learning in video multi-object tracking: A survey
Neurocomputing · 2019 · 638 citations
References
Modeling the Shape of the Scene: A Holistic Representation of the Spatial Envelope
International Journal of Computer Vision · 2001 · 6,378 citations
Learning Hierarchical Features for Scene Labeling
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2012 · 2,704 citations
The Pascal Visual Object Classes (VOC) Challenge
International Journal of Computer Vision · 2009 · 19,127 citations
Selective Search for Object Recognition
International Journal of Computer Vision · 2013 · 6,087 citations
Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2015 · 11,231 citations
Gradient-based learning applied to document recognition
Proceedings of the IEEE · 1998 · 57,014 citations
ImageNet Large Scale Visual Recognition Challenge
International Journal of Computer Vision · 2015 · 39,683 citations
Backpropagation Applied to Handwritten Zip Code Recognition
Neural Computation · 1989 · 11,706 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.