Scinovex
reviewTop 1% cited

A Comprehensive Survey of Deep Learning for Image Captioning

ACM Computing Surveys · 2019 · Vol. 51(6) · pp. 1–36
Md Zakir HossainFerdous SohelMohd Fairuz ShiratuddinHamid Laga

Abstract

Generating a description of an image is called image captioning. Image captioning requires recognizing the important objects, their attributes, and their relationships in an image. It also needs to generate syntactically and semantically correct sentences. Deep-learning-based techniques are capable of handling the complexities and challenges of image captioning. In this survey article, we aim to present a comprehensive review of existing deep-learning-based image captioning techniques. We discuss the foundation of the techniques to analyze their performances, strengths, and limitations. We also discuss the datasets and the evaluation metrics popularly used in deep-learning-based automatic image captioning.

Multimodal Machine Learning ApplicationsTopic ModelingAdvanced Image and Video Retrieval TechniquesClosed captioningComputer scienceImage (mathematics)Artificial intelligenceDeep learningNatural language processingInformation retrieval

Funding

  • Australian Research Council
Citations
860
FWCI
44.66
field-weighted impact
References
168
Percentile
100%
vs. same field & year
Citations per year
References
Long Short-Term Memory
Neural Computation · 1997 · 95,078 citations
Gradient-based learning applied to document recognition
Proceedings of the IEEE · 1998 · 57,014 citations
Distinctive Image Features from Scale-Invariant Keypoints
International Journal of Computer Vision · 2004 · 54,768 citations
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
International Journal of Computer Vision · 2017 · 5,085 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.