Peningkatan Deteksi Deepfake Menggunakan Arsitektur Hibrida Multiskala Vision Transformer dan Xceptionnet (M-Vixnet)
| dc.contributor.author | Martri Lina Anggraini | |
| dc.date.accessioned | 2026-07-29T03:38:08Z | |
| dc.date.issued | 2026-07-15 | |
| dc.description | :: Finalisasi file repositori 29 Juli 2026_Kurnadi | |
| dc.description.abstract | The rapid advancement of artificial intelligence, particularly Generative Adversarial Networks (GANs), has significantly improved the Realism of deepfake images, making manual identification increasingly difficult. This development poses serious threats to digital security, including identity fraud, misinformation, and multimedia manipulation. Therefore, an accurate and robust automated deepfake detection system is required to distinguish authentic facial images from manipulated ones. This study proposes M-ViXNet (Multi-Scale ViT and XceptionNet), a hybrid deep learning architecture developed by integrating hierarchical feature extraction from XceptionNet with multi-scale representations learned by ViT. The proposed architecture extracts multi-level features from the Entry Flow, Middle Flow, and Exit Flow of XceptionNet, which are subsequently processed by three ViT branches using patch sizes of 8×8, 16×16, and 32×32, respectively. The extracted features are then fused through an Adaptive Cross-Scale Attention Fusion mechanism to obtain comprehensive local and global feature representations for deepfake detection. Experiments were conducted using the 140k Real and Fake Faces (7030 variant) dataset consisting of 10,000 facial images divided into training, validation, and testing sets. The proposed model was evaluated against three baseline architectures, namely XceptionNet, Vision Transformer (ViT), and ViXNet, using Accuracy, Precision, Recall, F1-Score, and Area Under the Receiver Operating Characteristic Curve (ROC-AUC). Experimental results demonstrate that M-ViXNet achieved the best performance with an Accuracy of 94.00%, Precision of 98.59%, Recall of 99.07%, F1-Score of 92.36%, and ROC-AUC of 98.36%, outperforming all comparison models. Furthermore, Grad-CAM visualization confirms that M-ViXNet focuses more accurately on manipulated facial regions, indicating that the integration of hierarchical CNN features and multi-scale transformer representations effectively improves deepfake detection performance. | |
| dc.description.sponsorship | DPU: Dr. Dwiretno Istiyadi Swasono ST.,M.Kom. | |
| dc.identifier.uri | https://repository.unej.ac.id/handle/123456789/12331 | |
| dc.language.iso | other | |
| dc.publisher | Fakultas Ilmu Komputer | |
| dc.subject | Deepfake Detection | |
| dc.subject | ViT | |
| dc.subject | XceptionNet | |
| dc.subject | Multi-Scale Learning | |
| dc.subject | Adaptive Cross-Scale Attention Fusion. | |
| dc.title | Peningkatan Deteksi Deepfake Menggunakan Arsitektur Hibrida Multiskala Vision Transformer dan Xceptionnet (M-Vixnet) | |
| dc.title.alternative | none | |
| dc.type | Other |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Repository_Martri Lina Anggraini_222410103074.pdf
- Size:
- 2.03 MB
- Format:
- Adobe Portable Document Format
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed to upon submission
- Description:
