Scinovex
article Open Access

A hybrid deep learning approach for deepfake detection using spatial and temporal features with attention mechanisms

Abstract

Deepfake attacks threaten the authenticity of digital media, requiring strong detection methodologies to counter them. We therefore propose a deepfake detection system in which EfficientNetB3 acts as a spatial feature extractor, BiLSTM allows for temporal sequence modeling, and the self-attention mechanism creates attention on discriminative frames. The method is tested against the highly challenging Celeb-DF dataset, in which it achieves an accuracy of 85% on the test split. This also suggests that the proposed method successfully captures spatial and temporal discrepancies inside deepfake videos and therefore, is a viable candidate to analyze high-quality synthesized content. Early stop has been applied to prevent the model from overfitting the training data and enhance generalization to unseen data. The future aims of this research are to improve the robustness of the face detector and explore multimodal approaches to improve the inference accuracy further.

Citations
0
FWCI
0.00
field-weighted impact
References
7
Percentile
20%
vs. same field & year
References
Deep learning techniques for deep fake identification: A review
International Journal of Computing and Artificial Intelligence · 2025 · 1 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

A hybrid deep learning approach for deepfake detection using spatial and temporal features with attention mechanisms · Scinovex