A Hybrid Wav2Vec2 and LFCC-CNN Framework for Voice Spoofing Detection
DOI:
https://doi.org/10.24237/djes.2026.19302Keywords:
Automatic Speaker Verification, Deepfake Audio, Feature Fusion, Self-Supervised Learning, ASVspoof2019Abstract
Anti-spoofing systems are important for protecting security-sensitive applications such as financial services, smart assistants, biometric authentication, and access control systems. However, the increasing use of text-to-speech (TTS) and voice conversion (VC) technologies has made it easier to generate spoofed speech that closely resembles genuine human speech. Therefore, developing reliable methods for distinguishing genuine and spoofed speech remains an important challenge. A hybrid voice spoofing detection framework is proposed in this work by fusing Wav2Vec2 representations with Linear Frequency Cepstral Coefficient (LFCC) features. Wav2Vec2 extracts contextual representations directly from raw speech, while LFCC features capture fine-grained spectral characteristics associated with artifacts introduced by speech synthesis and voice conversion. The LFCC features are encoded using a lightweight Convolutional Neural Network (CNN), and the resulting 256-dimensional representation is fused with the 768-dimensional Wav2Vec2 embedding to form a 1024-dimensional feature representation. The fused representation is subsequently processed by a fully connected classification network to distinguish bona fide and spoofed speech. The proposed framework is evaluated on the ASVspoof2019 Logical Access (LA) dataset using standard anti-spoofing evaluation metrics. Experimental results show that the proposed hybrid model achieves an accuracy of 98.95% and an Equal Error Rate (EER) of 1.02%. An ablation study comparing the individual Wav2Vec2 and LFCC-CNN branches with the proposed hybrid model shows the benefit of combining contextual and spectral representations for voice spoofing detection.
Downloads
References
[1]. H. Li, "Spoofing and countermeasures for speaker verification: A survey," Speech Communication, vol. 66, pp. 130–153, 2015, doi: 10.1016/j.specom.2014.10.005.
[2]. M. Todisco, H. Delgado, and N. Evans, "Constant Q cepstral coefficients: A spoofing countermeasure for automatic speaker verification," Computer Speech & Language, vol. 45, pp. 516–535, 2017, doi: 10.1016/j.csl.2017.01.001.
[3]. Z. Wu, J. Yamagishi, T. Kinnunen, C. Hanilçi, M. Sahidullah, A. Sizov, N. Evans, M. Todisco, and H. Delgado, "ASVspoof: The Automatic Speaker Verification Spoofing and Countermeasures Challenge," IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 4, pp. 588–604, Jun. 2017, doi: 10.1109/JSTSP.2017.2671435.
[4]. X. Wang, H. Delgado, H. Tak, J.-w. Jung, H.-j. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. H. Kinnunen, N. Evans, K. A. Lee, and J. Yamagishi, “ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale,” in Proc. Automatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 1–8, doi: 10.21437/ASVspoof.2024-1.
[5]. Y. Eom, Y. Lee, J. S. Um, and H. R. Kim, "Anti-Spoofing Using Transfer Learning with Variational Information Bottleneck," in Proc. Interspeech, 2022, pp. 3568–3572, doi: 10.21437/Interspeech.2022-10200.
[6]. H. Tak, J.-W. Jung, J. Patino, M. Todisco, and N. Evans, "Graph Attention Networks for Anti-Spoofing," in Proc. Interspeech, 2021, pp. 2356–2360, doi: 10.21437/Interspeech.2021-993.
[7]. B. Huang, S. Cui, J. Huang, and X. Kang, “Discriminative Frequency Information Learning for End-to-End Speech Anti-Spoofing,” IEEE Signal Processing Letters, vol. 30, pp. 185–189, 2023, doi: 10.1109/LSP.2023.3251895.
[8]. Y. El Kheir, A. Das, E. E. Erdogan, F. Ritter-Guttierez, T. Polzehl, and S. Möller, “Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection,” in Proc. 2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2025, pp. 1–5, doi: 10.1109/WASPAA66052.2025.11230938.
[9]. A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, "wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, pp. 12449–12460, doi: 10.48550/arXiv.2006.11477.
[10]. M. J. Alam, P. Kenny, G. Bhattacharya, and T. Stafylakis, "Development of CRIM system for the automatic speaker verification spoofing and countermeasures challenge 2015," in Proc. Interspeech, 2015, pp. 2072–2076, doi: 10.21437/Interspeech.2015-469.
[11]. X. Wang et al., “ASVspoof 2019: A Large-Scale Public Database of Synthesized, Converted and Replayed Speech,” Computer Speech & Language, vol. 64, Art. no. 101114, 2020, doi: 10.1016/j.csl.2020.101114.
[12]. J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, N. Evans, and H. Delgado, “ASVspoof 2021: Accelerating Progress in Spoofed and Deepfake Speech Detection,” in Proc. 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge (ASVspoof 2021), 2021, pp. 47–54, doi: 10.21437/ASVSPOOF.2021-8.
[13]. X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, and K. A. Lee, “ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2507–2522, 2023, doi: 10.1109/TASLP.2023.3285283.
[14]. M. Neelima and I. S. Prabha, “Hybrid feature optimization for voice spoof detection using CNN-LSTM,” Traitement du Signal, vol. 41, no. 2, pp. 717–727, Apr. 2024, doi: 10.18280/ts.410214.
[15]. J. Zhou, H. Tao, D. J. Jawawi, D. Wang, E. Ibeke, and C. Biamba, “Voice Spoofing Countermeasure for Voice Replay Attacks Using Deep Learning,” Research Square preprint, Jul. 2022, doi: 10.1186/s13677-022-00306-5.
[16]. T. Arif, A. Javed, M. Alhameed, F. Jeribi, and A. Tahir, “Voice Spoofing Countermeasure for Logical Access Attacks Detection,” IEEE Access, vol. 9, pp. 162857–162868, 2021, doi: 10.1109/ACCESS.2021.3133134.
[17]. K. Shahzad, S. Farhan, Yasin-Ul-Haq, R. Sana, and S. Pathan, “Enhancing Voice Spoofing Detection: A Hybrid Approach with VGGish-LSTM Model for Improved Security in Automatic Speaker Verification Systems,” IEEE Access, vol. 13, pp. 40682–40702, 2025, doi: 10.1109/ACCESS.2025.3544639.
[18]. Y. Ren, H. Peng, L. Li, and Y. Yang, “Lightweight Voice Spoofing Detection Using Improved One-Class Learning and Knowledge Distillation,” IEEE Transactions on Multimedia, vol. 26, pp. 4360–4374, 2024, doi: 10.1109/TMM.2023.3321505.
[19]. Z. M. Almutairi and H. Elgibreen, “Detecting Fake Audio of Arabic Speakers Using Self-Supervised Deep Learning,” IEEE Access, vol. 11, pp. 72134–72147, 2023, doi: 10.1109/ACCESS.2023.3286864.
[20]. Y. Lee, N. Kim, J. Jeong, and I.-Y. Kwak, “Experimental Case Study of Self-Supervised Learning for Voice Spoofing Detection,” IEEE Access, vol. 11, pp. 24216–24226, 2023, doi: 10.1109/ACCESS.2023.3254880.
[21]. J. M. Martín-Doñas and A. Álvarez, “The Vicomtech Audio Deepfake Detection System Based on Wav2Vec2 for the 2022 ADD Challenge,” arXiv preprint arXiv:2203.01573, 2022, doi: 10.48550/arXiv.2203.01573.
[22]. Y. Xiao and R. K. Das, “XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection,” IEEE Signal Processing Letters, vol. 32, pp. 1276–1280, 2025, doi: 10.1109/LSP.2025.3547861.
[23]. F. Javanmardi, S. R. Kadiri, and P. Alku, “Exploring the Impact of Fine-Tuning the Wav2vec2 Model in Database-Independent Detection of Dysarthric Speech,” IEEE Journal of Biomedical and Health Informatics, vol. 28, no. 8, pp. 4951–4962, Aug. 2024, doi: 10.1109/JBHI.2024.3392829.
[24]. C. Yi, J. Wang, N. Cheng, S. Zhou, and B. Xu, “Applying Wav2Vec2.0 to Speech Recognition in Various Low-Resource Languages,” arXiv preprint arXiv:2012.12121, Dec. 2020, doi: 10.48550/arXiv.2012.12121.
[25]. I.-Y. Kwak, S. Kwag, J. Lee, Y. Jeon, J. Hwang, H.-J. Choi, J.-H. Yang, S.-Y. Han, J. H. Huh, C.-H. Lee, and J. W. Yoon, “Voice Spoofing Detection Through Residual Network, Max Feature Map, and Depthwise Separable Convolution,” IEEE Access, vol. 11, pp. 49140–49152, 2023, doi: 10.1109/ACCESS.2023.3275790.
[26]. B. Ustubioglu, G. Tahaoglu, A. Ustubioglu, G. Ulutas, I. Amerini, and M. Kilic, “Multi Pattern Features-Based Spoofing Detection Mechanism Using One Class Learning,” IEEE Access, vol. 12, pp. 117523–117540, 2024, doi: 10.1109/ACCESS.2024.3447572.
[27]. Y. Zhang, Z. Li, J. Lu, W. Wang, and P. Zhang, “Synthetic Speech Detection Based on the Temporal Consistency of Speaker Features,” IEEE Signal Processing Letters, vol. 31, pp. 944–948, 2024, doi: 10.1109/LSP.2024.3381890.
[28]. L. Zhang, X. Wang, E. Cooper, N. Evans, and J. Yamagishi, “The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 813–825, 2023, doi: 10.1109/TASLP.2022.3233236.
[29]. Z. Wang and J. H. L. Hansen, “Toward Improving Synthetic Audio Spoofing Detection Robustness via Meta-Learning and Disentangled Training with Adversarial Examples,” IEEE Access, vol. 12, pp. 99894–99911, 2024, doi: 10.1109/ACCESS.2024.3421281.
[30]. O. A. Shaaban, R. Yildirim, and A. A. Alguttar, “Audio Deepfake Approaches,” IEEE Access, vol. 11, pp. 132652–132682, 2023, doi: 10.1109/ACCESS.2023.3333866.
[31]. J. Lu, Y. Zhang, Z. Li, Z. Shang, W. Wang, and P. Zhang, “Leveraging Distance Information for Generalized Spoofing Speech Detection,” Computer Speech & Language, vol. 94, Art. no. 101804, 2025, doi: 10.1016/j.csl.2025.101804.
[32]. T. Kinnunen, H. Delgado, N. Evans, K. A. Lee, V. Vestman, A. Nautsch, M. Todisco, X. Wang, M. Sahidullah, J. Yamagishi, and D. A. Reynolds, “Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: Fundamentals,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2195–2210, 2020, doi: 10.1109/TASLP.2020.3009494.
[33]. Z. Akhtar, T. L. Pendyala, and V. S. Athmakuri, “Video and Audio Deepfake Datasets and Open Issues in Deepfake Technology: Being Ahead of the Curve,” Forensic Sciences, vol. 4, no. 3, pp. 289–377, 2024, doi: 10.3390/forensicsci4030021.
[34]. M. Sharafudeen et al., “A Blended Framework for Audio Spoof Detection with Sequential Models and Bags of Auditory Bites,” Scientific Reports, vol. 14, Art. no. 20192, 2024, doi: 10.1038/s41598-024-71026-w.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Abdul Ashfaque Basha, K. Venkata Prasad

This work is licensed under a Creative Commons Attribution 4.0 International License.









