Data-Efficient Transformer Architectures for Image-Level Facial Forgery Detection: A Comparative Evaluation of ViT and DeiT

نویسندگان

1 Cloud Architect, Lead of AI initiative Program, Ernst & Young LLP, New York, USA.

2 Professor, ECE Department, Sreenidhi Institute of Science and Technology, India.

3 Assistant Professor, Department of Biotechnology, Vinayaka Mission`s Kirupananda Variyar Engineering College, Salem (Vinayaka Mission`s Research Foundation). India.

4 Assistant Professor, Department of Computer Science& Engineering, B.M.S College of Engineering, Affiliated to Visvesvaraya Technological University, Belagavi, India.

5 Associate Professor, Department of Artificial Intelligence and Machine Learning, Bangalore Institute of Technology, Visvesvaraya Technological University, Bangalore, India.

6 Department of ECE, CMR Technical Campus, Hyderabad, Telangana, India.

doi
10.22059/jitm.2026.106252
چکیده

The rapid development of deepfake technologies has increased the demand for a credible and inter-pretable system for facial forgery detection. This study compares two transformer-based architec-tures—Vision Transformer (ViT) and Distilled Data-Efficient Image Transformer (DeiT)—for de-tecting real and manipulated facial images. The study aims to measure performance in terms of de-tection as well as interpretability and to address the weaknesses of traditional convolutional models. Data augmentation was applied, and a balanced dataset containing 8,000 real and fake images was constructed; both models were then fine-tuned under the same training environment. The explanatory ability of the models was incorporated using LIME. Experimental findings indicate that both models perform well, with DeiT being slightly more accurate at 94.62% than ViT at 93.6%, alongside faster convergence rates and less overfitting. Visualization of the focus on important facial areas confirms that the models reliably register synthetic artifacts. Although promising, generalization across dif-ferent datasets and enhancement of real-time performance remain challenges. Overall, the results validate transformer architectures—especially DeiT—as powerful and explainable deepfake detec-tion algorithms, valuable for ensuring safe and transparent digital media forensics.