Physical Sciences → Computer Science → Computer Vision and Pattern Recognition
Multimodal Machine Learning Applications
This cluster of papers focuses on the development and improvement of visual question answering systems, image captioning techniques, and neural networks for understanding and generating descriptions of images and videos. The research involves semantic reasoning, multimodal fusion, scene graph generation, attention mechanisms, and deep learning approaches to bridge the gap between vision and language.
67.4K works worldwide704.7K citations
Visual Question AnsweringImage CaptioningNeural NetworksSemantic ReasoningMultimodal FusionScene Graph GenerationVideo DescriptionAttention MechanismLanguage UnderstandingDeep Learning
Journals publishing in this area
2

IEEE Transactions on Pattern Analysis and Machine Intelligence
ISSN 0162-8828643 articles in this topic
547h-index
11.13Impact
12.1KArticles
1.8MCitations
