Analisis Performa Wav2Vec2.0 pada Transkripsi Nama Barang dalam Kondisi Audio Bersih dan Bising
DOI:
https://doi.org/10.29407/ecv7ey20Abstract
Automatic Speech Recognition (ASR) merupakan teknologi yang memungkinkan suara diubah menjadi teks secara otomatis. Penelitian ini bertujuan menganalisis performa model Wav2Vec2.0 pada transkripsi nama barang berbahasa Indonesia dalam kondisi audio bersih dan bising. Dataset terdiri atas 400 rekaman audio yang dikumpulkan dari empat pembicara dengan kosakata terbatas berupa 20 nama barang. Evaluasi dilakukan menggunakan metrik Word Error Rate (WER) dan Character Error Rate (CER) pada kondisi audio bersih, babble noise, crowd noise, dan vehicle noise. Hasil pengujian menunjukkan bahwa kondisi audio bersih menghasilkan performa terbaik dengan nilai WER sebesar 2,08% dan CER sebesar 0,35%, sedangkan crowd noise menghasilkan tingkat kesalahan tertinggi dengan WER sebesar 22,92% dan CER sebesar 9,44%. Hasil penelitian menunjukkan bahwa peningkatan tingkat kebisingan berpengaruh terhadap penurunan performa transkripsi model Wav2Vec2.0 pada tugas pengenalan kosakata terbatas.
Keywords:
kebisingan, pengenalan suara otomatis, transkripsi nama barang, Wav2Vec2.0##plugins.themes.default.displayStats.downloads##
References
[1] R. Setiawan, N. P. Sugihartanti, and M. I. Ibadurrahman, ―Sistem Manajemen Gudang
Bebasis Web dengan Teknologi Barcode Scanner pada Industri Manufaktur: Sebuah
Kajian Literatur,‖ Integrasi: Jurnal Ilmiah Teknik Industri, vol. 9, no. 2, pp. 124–135,
2024, doi: https://doi.org/10.32502/integrasi.v9i2.181.
[2] P. Q. Izzatiy and S. Sundari, ―IMPLEMENTASI ACCURATE DALAM SISTEM
PENCATATAN KAS KECIL DAN HUTANG PADA PERUSAHAAN
MANUFAKTUR DI SURABAYA,‖ Jurnal Ekonomi Bisnis Manajemen dan
Akuntansi (JEBISMA), vol. 3, no. 3, 2026, doi:
https://doi.org/10.70197/jebisma.v3i3.198.
[3] H. Azis, E. P. Indreswari, and R. Wisudawanto, ―PEMANFAATAN TEKNOLOGI
AUTOMATIC SPEECH RECOGNITION DALAM MENCIPTAKAN
PEMBELAJARAN INKLUSIF: IMPLEMENTASI, EFEKTIFITAS DAN
TANTANGAN,‖ GANESHA: Jurnal Pengabdian Masyarakat, vol. 5, no. 1, pp. 27–33,
2025, doi: https://doi.org/10.36728/ganesha.v5i1.4171.
[4] H. Wijaya, ―Teknologi Pengenalan Suara tentang Metode, Bahasa dan Tantangan:
Systematic Literature Review,‖ bit-Tech, vol. 7, no. 2, pp. 533–544, 2024, doi:
https://doi.org/10.32877/bt.v7i2.1888.
[5] D. Ferdiansyah and C. S. K. Aditya, ―Implementasi Automatic Speech Recognition
Bacaan Al-Qur’an Menggunakan Metode Wav2Vec 2.0 dan OpenAI-Whisper,‖ Jurnal
Teknik Elektro Dan Komputer TRIAC, vol. 11, no. 1, pp. 11–16, 2024, doi:
https://doi.org/10.21107/triac.v11i1.24332.
[6] Abdullah Azzam, Ichsan Taufik, and Aldy Rialdy Atmadja, ―Evaluating End-to-End
ASR for Qur’an Recitation Using Whispers in Low Resource Settings,‖ Bulletin of
Computer Science Research, vol. 5, no. 4, pp. 778–787, Jun. 2025, doi:
10.47065/bulletincsr.v5i4.561.
[7] A. NOERCHOLIS, T. DWIANDINI, and F. S. MUKTI, ―Optimasi Teknologi
WAV2Vec 2.0 menggunakan Spectral Masking untuk meningkatkan Kualitas
Transkripsi Teks Video bagi Tuna Rungu,‖ ELKOMIKA: Jurnal Teknik Energi Elektrik,
Teknik Telekomunikasi, & Teknik Elektronika, vol. 12, no. 4, p. 877, 2024, doi:
https://doi.org/10.26760/elkomika.v12i4.877.
[8] J. Cai, Y. Song, J. Wu, and X. Chen, ―Voice Disorder Classification Using Wav2vec
2.0 Feature Extraction,‖ Journal of Voice, 2024, doi:
https://doi.org/10.1016/j.jvoice.2024.09.002.
[9] J. C. Duarte and S. Colcher, ―Noise-Robust Automatic Speech Recognition: A Case
Study for Communication Interference,‖ Journal on Interactive Systems, vol. 15, no. 1,
pp. 670–681, Jul. 2024, doi: 10.5753/jis.2024.4267.
[10] M. N. Althoff, A. Affandy, A. Luthfiarta, M. W. B. D. Satya, and H. Basiron,
―Leveraging Label Preprocessing for Effective End-to-End Indonesian Automatic
Speech Recognition,‖ Sinkron : jurnal dan penelitian teknik informatika, vol. 9, no. 1,
pp. 55–64, Jan. 2025, doi: 10.33395/sinkron.v9i1.14257.
[11] C. Mena, A. L. Padilla-Ortiz, and F. Orduña-Bustamante, ―Automatic Speech
Recognition in the presence of babble noise and reverberation compared to human
speech intelligibility in Spanish,‖ Comput. Speech Lang., vol. 95, p. 101856, 2026, doi:
10.1016/j.csl.2025.101856.
[12] T. D. K, J. James, D. P. Gopinath, and M. A. K, ―Advocating Character Error Rate for
Multilingual ASR Evaluation,‖ in Findings of the Association for Computational
Linguistics: NAACL 2025, L. Chiruzzo, A. Ritter, and L. Wang, Eds., Albuquerque,
New Mexico: Association for Computational Linguistics, Apr. 2025, pp. 4941–4950.
doi: 10.18653/v1/2025.findings-naacl.277.
[13] A. Baevski, H. Zhou, A. Mohamed, and M. Auli, ―wav2vec 2.0: a framework for selfsupervised learning of speech representations,‖ in Proceedings of the 34th
International Conference on Neural Information Processing Systems, in NIPS ’20. Red
Hook, NY, USA: Curran Associates Inc., 2020, doi: 10.48550/arXiv.2006.11477.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Copyright on any article is retained by the author(s).
- The author grants the journal, right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgment of the work’s authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal’s published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
- The article and any associated published material is distributed under the Creative Commons Attribution-ShareAlike 4.0 International License