Scinovex
article Open AccessTop 1% cited

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

International Journal of Computer Vision · 2017 · Vol. 123(1) · pp. 32–73
Ranjay KrishnaYuke ZhuOliver GrothJustin JohnsonKenji HataJoshua KravitzStephanie ChenYannis KalantidisLi-Jia LiDavid A. ShammaMichael S. BernsteinLi Fei-Fei

Abstract

Despite progress in perceptual tasks such as image classification, computers still perform poorly on cognitive tasks such as image description and question answering. Cognition is core to tasks that involve not just recognizing, but reasoning about our visual world. However, models used to tackle the rich content in images for cognitive tasks are still being trained using the same datasets designed for perceptual tasks. To achieve success at cognitive tasks, models need to understand the interactions and relationships between objects in an image. When asked “What vehicle is the person riding?”, computers will need to identify the objects in an image as well as the relationships riding(man, carriage) and pulling(horse, carriage) to answer correctly that “the person is riding a horse-drawn carriage.” In this paper, we present the Visual Genome dataset to enable the modeling of such relationships. We collect dense annotations of objects, attributes, and relationships within each image to learn these models. Specifically, our dataset contains over 108K images where each image has an average of $$35$$ objects, $$26$$ attributes, and $$21$$ pairwise relationships between objects. We canonicalize the objects, attributes, relationships, and noun phrases in region descriptions and questions answer pairs to WordNet synsets. Together, these annotations represent the densest and largest dataset of image descriptions, objects, attributes, relationships, and question answer pairs.

Multimodal Machine Learning ApplicationsImage Retrieval and Classification TechniquesAdvanced Image and Video Retrieval TechniquesArtificial intelligenceComputer scienceNatural language processingGenomeImage (mathematics)Pattern recognition (psychology)Computer visionAnnotationCrowdsourcingImage processing

Funding

  • Toyota USA
  • Brown Institute for Media Innovation
  • Multidisciplinary University Research Initiative
  • Office of Naval Research
Citations
5,085
FWCI
143.37
field-weighted impact
References
132
Percentile
100%
vs. same field & year
Citations per year
Cited by
A Comprehensive Survey of Deep Learning for Image Captioning
ACM Computing Surveys · 2019 · 860 citations
A Survey of Deep Active Learning
ACM Computing Surveys · 2021 · 1,001 citations
Places: A 10 Million Image Database for Scene Recognition
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2017 · 3,945 citations
References
Scripts, Plans, Goals, and Understanding: An Inquiry into Human Knowledge Structures
The American Journal of Psychology · 1979 · 2,979 citations
Pedestrian Detection: An Evaluation of the State of the Art
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2011 · 3,229 citations
The Pascal Visual Object Classes (VOC) Challenge
International Journal of Computer Vision · 2009 · 19,127 citations
Long Short-Term Memory
Neural Computation · 1997 · 95,078 citations
WordNet
Communications of the ACM · 1995 · 13,991 citations
A Statistical Approach to Texture Classification from Single Images
International Journal of Computer Vision · 2005 · 1,149 citations
LabelMe: A Database and Web-Based Tool for Image Annotation
International Journal of Computer Vision · 2007 · 4,112 citations
ImageNet Large Scale Visual Recognition Challenge
International Journal of Computer Vision · 2015 · 39,683 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations · Scinovex