Deep video understanding based on language generation
Abstract
The research category touches on "Deep Video Understanding Based on Language Generation" that aims to put the mechanisms their study is developing in the areas of technology on video clips and have such models that can not just keep an eye on what they seek without noticing, and respond dynamically based on the understanding they derive from this content.Research in this area combines computer vision techniques with natural language processing, using Convolutional Neural Networks (CNNs) to detect elements, scenes, and actions in the video, and Long Short-Term Memory (LSTM) networks to catch event temporal dimensions.In this paper, to understand video by generating language, we proposed the proposed hybrid model, which is a deep understanding of video after converting it to frames and then using CNN technology, then processing the datasets and using LSTM technology to generate linguistic descriptions, where two types of data were used, which are videos, and the accuracy rate of the proposed system in converting video to linguistic descriptions and recognizing movements was achieved by 90% for the used dataset, in addition to using deep learning, applied using (CNN).The goal of CNN is to prove the validity of the proposed approach, where a data set will be trained and the results will be tested after using the proposed approach.
How this paper connects to the literature. Drag to explore, click any node to open that paper.
