Peningkatan Deteksi Deepfake Menggunakan Arsitektur Hibrida Multiskala Vision Transformer dan Xceptionnet (M-Vixnet)

dc.contributor.authorMartri Lina Anggraini
dc.date.accessioned2026-07-29T03:38:08Z
dc.date.issued2026-07-15
dc.description:: Finalisasi file repositori 29 Juli 2026_Kurnadi
dc.description.abstractThe rapid advancement of artificial intelligence, particularly Generative Adversarial Networks (GANs), has significantly improved the Realism of deepfake images, making manual identification increasingly difficult. This development poses serious threats to digital security, including identity fraud, misinformation, and multimedia manipulation. Therefore, an accurate and robust automated deepfake detection system is required to distinguish authentic facial images from manipulated ones. This study proposes M-ViXNet (Multi-Scale ViT and XceptionNet), a hybrid deep learning architecture developed by integrating hierarchical feature extraction from XceptionNet with multi-scale representations learned by ViT. The proposed architecture extracts multi-level features from the Entry Flow, Middle Flow, and Exit Flow of XceptionNet, which are subsequently processed by three ViT branches using patch sizes of 8×8, 16×16, and 32×32, respectively. The extracted features are then fused through an Adaptive Cross-Scale Attention Fusion mechanism to obtain comprehensive local and global feature representations for deepfake detection. Experiments were conducted using the 140k Real and Fake Faces (7030 variant) dataset consisting of 10,000 facial images divided into training, validation, and testing sets. The proposed model was evaluated against three baseline architectures, namely XceptionNet, Vision Transformer (ViT), and ViXNet, using Accuracy, Precision, Recall, F1-Score, and Area Under the Receiver Operating Characteristic Curve (ROC-AUC). Experimental results demonstrate that M-ViXNet achieved the best performance with an Accuracy of 94.00%, Precision of 98.59%, Recall of 99.07%, F1-Score of 92.36%, and ROC-AUC of 98.36%, outperforming all comparison models. Furthermore, Grad-CAM visualization confirms that M-ViXNet focuses more accurately on manipulated facial regions, indicating that the integration of hierarchical CNN features and multi-scale transformer representations effectively improves deepfake detection performance.
dc.description.sponsorshipDPU: Dr. Dwiretno Istiyadi Swasono ST.,M.Kom.
dc.identifier.urihttps://repository.unej.ac.id/handle/123456789/12331
dc.language.isoother
dc.publisherFakultas Ilmu Komputer
dc.subjectDeepfake Detection
dc.subjectViT
dc.subjectXceptionNet
dc.subjectMulti-Scale Learning
dc.subjectAdaptive Cross-Scale Attention Fusion.
dc.titlePeningkatan Deteksi Deepfake Menggunakan Arsitektur Hibrida Multiskala Vision Transformer dan Xceptionnet (M-Vixnet)
dc.title.alternativenone
dc.typeOther

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Repository_Martri Lina Anggraini_222410103074.pdf
Size:
2.03 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed to upon submission
Description: