Analisis Kinerja CNN+LSTM dalam Klasifikasi Tangisan Bayi Menggunakan Ekstraksi Fitur Audio

Authors

  • David Yoga Wicaksono Universitas Nusantara PGRI Kediri Indonesia
  • Daniel Swanjaya Universitas Nusantara PGRI Kediri Indonesia
  • Resty Wulanningrum Universitas Nusantara PGRI Kediri Indonesia

DOI:

https://doi.org/10.29407/gnrehk65

Abstract

Tangisan adalah alat komunikasi utama bagi bayi untuk menyampaikan kebutuhan, tetapi seringkali sulit dipahami. Studi ini bertujuan untuk mengklasifikasikan lima jenis tangisan bayi menggunakan ekstraksi fitur Mel Spectrogram. Untuk mengatasi ketidakseimbangan data, teknik augmentasi audio dan metode SMOTE diterapkan. Analisis perbandingan arsitektur dievaluasi dalam studi ini. Hasil pengujian menunjukkan bahwa model berbasis konvolusi memberikan kinerja tinggi di atas 95%. Meskipun Convolutional Neural Network memiliki akurasi sedikit lebih tinggi, model hibrida CNN-LSTM terbukti paling optimal karena menghasilkan nilai kesalahan pengujian terendah sebesar 0,1160. Integrasi CNN dan LSTM mampu menangkap fitur spasial-frekuensi serta karakteristik temporal suara secara komprehensif. Kesimpulannya, model hibrida ini memberikan solusi cerdas dan akurat untuk sistem pengenalan suara bayi guna mendukung pengembangan asisten pengasuh di masa depan.

Keywords:

CNN-LSTM, Deep Learning, Mel Spectrogram, SMOTE, Tangisan Bayi

##plugins.themes.default.displayStats.downloads##

##plugins.themes.default.displayStats.noStats##

References

[1] L. Yunita and D. Surayana, “Perkembangan personality sosial usia bayi dan toddler,”

Jurnal Family Education, vol. 1, no. 4, pp. 14–22, 2021.

https://doi.org/10.24036/jfe.v1i4.20

[2] H. Yoo, E. H. Buder, D. D. Bowman, G. M. Bidelman, and D. K. Oller, “Acoustic

correlates and adult perceptions of distress in infant speech-like vocalizations and cries,”

Front. Psychol., vol. 10, p. 1154, 2019. https://doi.org/10.3389/fpsyg.2019.01154

[3] D. E. Schoth and C. Liossi, “A systematic review of experimental paradigms for

exploring biased interpretation of ambiguous information with emotional and neutral

associations,” Front. Psychol., vol. 8, p. 171, 2017.

https://doi.org/10.1371/journal.pone.0318296

[4] C. Li, A. Pourtaherian, L. Van Onzenoort, W. E. T. a Ten, and P. H. N. de With, “Infant

monitoring system for real-time and remote discomfort detection,” IEEE Transactions

on Consumer Electronics, vol. 66, no. 4, pp. 336–345, 2020. DOI:

10.1109/TCE.2020.3031359

[5] A. S. H. George and A. S. George, “Decoding Infant Communication: Understanding

the Meaning Behind Baby’s Cries,” Partners Universal Innovative Research

Publication, vol. 1, no. 1, pp. 1–14, 2023. https://doi.org/10.5281/zenodo.8432700

[6] T. H. Rochadiani, “Pendekatan Transfer Learning Untuk Klasifikasi Tangisan Bayi

Dengan Imbalance Dataset,” The Indonesian Journal of Computer Science, vol. 13, no.

2, 2024. https://doi.org/10.33022/ijcs.v13i2.3834

[7] P. Cariani and J. M. Baker, “Time is of the essence: neural codes, synchronies,

oscillations, architectures,” Front. Comput. Neurosci., vol. 16, p. 898829, 2022.

https://doi.org/10.3389/fncom.2022.898829

[8] M. Tiwari and D. K. Verma, “Enhanced text-independent speaker recognition using

MFCC, Bi-LSTM, and CNN-based noise removal techniques,” Int. J. Speech Technol.,

vol. 27, no. 4, pp. 1013–1026, 2024. https://doi.org/10.1007/s10772-024-10150-4

[9] C. Asuai, A. P. Arinomor, C. T. Atumah, I. F. Kowhoro, and D. E. Ogheneochuko,

“Hybrid CNN-LSTM Architectures for Deepfake Audio Detection Using Mel

Frequency Cepstral Coefficients and Spectogram Analysis,” American Journal of

Mathematical and Computer Modelling, vol. 10, no. 3, pp. 98–109, 2025. DOI:

10.11648/j.ajmcm.20251003.12

[10] B. A. Febryanto and I. Tahyudin, “Perbandingan Algoritma CNN, LSTM, FNN untuk

Diagnosa Fibrosis Hati dengan Citra Medis.,” Techno. com, vol. 24, no. 1, 2025. Doi:

10.62411/tc.v24i1.12020

[11] G. Meriyem, D. Abdelaziz, and E. Mohamed, “From Traditional to Deep Learning

Methods for Baby Cry Audio Signal Classification: A Systematic Literature Review,”

in 2025 International Conference on Circuit, Systems and Communication (ICCSC),

IEEE, 2025, pp. 1–6. DOI: 10.1109/ICCSC66714.2025.11135198

[12] I. P. A. A. Mahendra and K. Kusrini, “Analisis Komparatif Klasifikasi Sentimen

Pengguna Aplikasi Investasi Menggunakan Algoritma Hybrid Cnn-Lstm, Cnn-Gru

Dengan Implementasi Smote,” Teknimedia: Teknologi Informasi dan Multimedia, vol.

6, no. 1, pp. 112–118, 2025. https://doi.org/10.46764/teknimedia.v6i1.258

[13] N. M. Susetio, “Analysis of the Effect of Data Augmentation on the Classification of

Indonesian Pronunciation Accuracy in Pronouncing the English Word’City’Using

Deep Learning Methods with MEL Spectrogram and Convolutional Neural Network

(CNN),” Available at SSRN 5196866, 2025. http://dx.doi.org/10.2139/ssrn.5196866

[14] Haoran Kui, Jiahua Pan, Rong Zong, Hongbo Yang, Weilian Wang, "Heart sound

classification based on log Mel-frequency spectral coefficients features and

convolutional neural networks,” Biomedical Signal Processing and Control, Volume

69, 2021, 102893,ISSN 1746-8094, https://doi.org/10.1016/j.bspc.2021.102893

[15] A. Ustubioglu, B. Ustubioglu, and G. Ulutas, “Mel spectrogram-based audio forgery

detection using CNN,” Signal Image Video Process., vol. 17, no. 5, pp. 2211–2219,

2023. https://doi.org/10.1007/s11760-022-02436-4

[16] C. Aliferis and G. Simon, “Overfitting, underfitting and general model overconfidence

and under-performance pitfalls and best practices in machine learning and AI,”

Artificial intelligence and machine learning in health care and medical sciences: Best

practices and pitfalls, pp. 477–524, 2024. https://doi.org/10.1007/978-3-031-39355-

6_10

[17] D. Fadillah, E. Haerani, F. Wulandari, and F. Syafria, “Klasifikasi Kondisi Janin

Menggunakan Algoritma K-Nearest Neighbors dan Teknik SMOTE Berdasarkan Data

Kardiotogram,” Bulletin of Computer Science Research, vol. 5, no. 4, pp. 482–489,

2025. https://doi.org/10.47065/bulletincsr.v5i4.585

[18] T. Ahmad, J. Wu, H. S. Alwageed, F. Khan, J. Khan, and Y. Lee, “Human activity

recognition based on deep-temporal learning using convolution neural networks

features and bidirectional gated recurrent unit with features selection,” IEEE access,

vol. 11, pp. 33148–33159, 2023. doi: 10.1109/ACCESS.2023.3263155

[19] Lusi Dwi Anggraini, Umi Mahdiyah, Resty Wulanningrum, “Pengujian Simple CNN

Dengan Arsitektur LeNet5 Pada Penyakit Daun Jagung”, inotek, vol. 9, no. 1, pp. 711–

719, Jul. 2025, doi: 10.29407/9qdmem64.

[20] A. Farzad, H. Mashayekhi, and H. Hassanpour, “A comparative performance analysis

of different activation functions in LSTM networks for classification,” Neural Comput.

Appl., vol. 31, no. 7, pp. 2507–2521, 2019. https://doi.org/10.1007/s00521-017-3210-

6

[21] M. C. Rani, R. A. Dewi, F. D. Azkia, M. Wahyudi, and A. S. Budiman, “Perbandingan

Algoritma Random Forest, Naive Bayes, Dan Neural Network Dalam Klasifikasi

Penyakit Jantung,” Jurnal Sains Informatika Terapan, vol. 4, no. 2, pp. 77–84, 2025.

https://doi.org/10.62357/jsit.v4i2.609

Downloads

Published

2026-07-20

How to Cite

Analisis Kinerja CNN+LSTM dalam Klasifikasi Tangisan Bayi Menggunakan Ekstraksi Fitur Audio. (2026). Prosiding SEMNAS INOTEK (Seminar Nasional Inovasi Teknologi), 10(2), 1231-1240. https://doi.org/10.29407/gnrehk65