Analisis Kinerja CNN+LSTM dalam Klasifikasi Tangisan Bayi Menggunakan Ekstraksi Fitur Audio
DOI:
https://doi.org/10.29407/gnrehk65Abstract
Tangisan adalah alat komunikasi utama bagi bayi untuk menyampaikan kebutuhan, tetapi seringkali sulit dipahami. Studi ini bertujuan untuk mengklasifikasikan lima jenis tangisan bayi menggunakan ekstraksi fitur Mel Spectrogram. Untuk mengatasi ketidakseimbangan data, teknik augmentasi audio dan metode SMOTE diterapkan. Analisis perbandingan arsitektur dievaluasi dalam studi ini. Hasil pengujian menunjukkan bahwa model berbasis konvolusi memberikan kinerja tinggi di atas 95%. Meskipun Convolutional Neural Network memiliki akurasi sedikit lebih tinggi, model hibrida CNN-LSTM terbukti paling optimal karena menghasilkan nilai kesalahan pengujian terendah sebesar 0,1160. Integrasi CNN dan LSTM mampu menangkap fitur spasial-frekuensi serta karakteristik temporal suara secara komprehensif. Kesimpulannya, model hibrida ini memberikan solusi cerdas dan akurat untuk sistem pengenalan suara bayi guna mendukung pengembangan asisten pengasuh di masa depan.
Keywords:
CNN-LSTM, Deep Learning, Mel Spectrogram, SMOTE, Tangisan Bayi##plugins.themes.default.displayStats.downloads##
References
[1] L. Yunita and D. Surayana, “Perkembangan personality sosial usia bayi dan toddler,”
Jurnal Family Education, vol. 1, no. 4, pp. 14–22, 2021.
https://doi.org/10.24036/jfe.v1i4.20
[2] H. Yoo, E. H. Buder, D. D. Bowman, G. M. Bidelman, and D. K. Oller, “Acoustic
correlates and adult perceptions of distress in infant speech-like vocalizations and cries,”
Front. Psychol., vol. 10, p. 1154, 2019. https://doi.org/10.3389/fpsyg.2019.01154
[3] D. E. Schoth and C. Liossi, “A systematic review of experimental paradigms for
exploring biased interpretation of ambiguous information with emotional and neutral
associations,” Front. Psychol., vol. 8, p. 171, 2017.
https://doi.org/10.1371/journal.pone.0318296
[4] C. Li, A. Pourtaherian, L. Van Onzenoort, W. E. T. a Ten, and P. H. N. de With, “Infant
monitoring system for real-time and remote discomfort detection,” IEEE Transactions
on Consumer Electronics, vol. 66, no. 4, pp. 336–345, 2020. DOI:
10.1109/TCE.2020.3031359
[5] A. S. H. George and A. S. George, “Decoding Infant Communication: Understanding
the Meaning Behind Baby’s Cries,” Partners Universal Innovative Research
Publication, vol. 1, no. 1, pp. 1–14, 2023. https://doi.org/10.5281/zenodo.8432700
[6] T. H. Rochadiani, “Pendekatan Transfer Learning Untuk Klasifikasi Tangisan Bayi
Dengan Imbalance Dataset,” The Indonesian Journal of Computer Science, vol. 13, no.
2, 2024. https://doi.org/10.33022/ijcs.v13i2.3834
[7] P. Cariani and J. M. Baker, “Time is of the essence: neural codes, synchronies,
oscillations, architectures,” Front. Comput. Neurosci., vol. 16, p. 898829, 2022.
https://doi.org/10.3389/fncom.2022.898829
[8] M. Tiwari and D. K. Verma, “Enhanced text-independent speaker recognition using
MFCC, Bi-LSTM, and CNN-based noise removal techniques,” Int. J. Speech Technol.,
vol. 27, no. 4, pp. 1013–1026, 2024. https://doi.org/10.1007/s10772-024-10150-4
[9] C. Asuai, A. P. Arinomor, C. T. Atumah, I. F. Kowhoro, and D. E. Ogheneochuko,
“Hybrid CNN-LSTM Architectures for Deepfake Audio Detection Using Mel
Frequency Cepstral Coefficients and Spectogram Analysis,” American Journal of
Mathematical and Computer Modelling, vol. 10, no. 3, pp. 98–109, 2025. DOI:
10.11648/j.ajmcm.20251003.12
[10] B. A. Febryanto and I. Tahyudin, “Perbandingan Algoritma CNN, LSTM, FNN untuk
Diagnosa Fibrosis Hati dengan Citra Medis.,” Techno. com, vol. 24, no. 1, 2025. Doi:
10.62411/tc.v24i1.12020
[11] G. Meriyem, D. Abdelaziz, and E. Mohamed, “From Traditional to Deep Learning
Methods for Baby Cry Audio Signal Classification: A Systematic Literature Review,”
in 2025 International Conference on Circuit, Systems and Communication (ICCSC),
IEEE, 2025, pp. 1–6. DOI: 10.1109/ICCSC66714.2025.11135198
[12] I. P. A. A. Mahendra and K. Kusrini, “Analisis Komparatif Klasifikasi Sentimen
Pengguna Aplikasi Investasi Menggunakan Algoritma Hybrid Cnn-Lstm, Cnn-Gru
Dengan Implementasi Smote,” Teknimedia: Teknologi Informasi dan Multimedia, vol.
6, no. 1, pp. 112–118, 2025. https://doi.org/10.46764/teknimedia.v6i1.258
[13] N. M. Susetio, “Analysis of the Effect of Data Augmentation on the Classification of
Indonesian Pronunciation Accuracy in Pronouncing the English Word’City’Using
Deep Learning Methods with MEL Spectrogram and Convolutional Neural Network
(CNN),” Available at SSRN 5196866, 2025. http://dx.doi.org/10.2139/ssrn.5196866
[14] Haoran Kui, Jiahua Pan, Rong Zong, Hongbo Yang, Weilian Wang, "Heart sound
classification based on log Mel-frequency spectral coefficients features and
convolutional neural networks,” Biomedical Signal Processing and Control, Volume
69, 2021, 102893,ISSN 1746-8094, https://doi.org/10.1016/j.bspc.2021.102893
[15] A. Ustubioglu, B. Ustubioglu, and G. Ulutas, “Mel spectrogram-based audio forgery
detection using CNN,” Signal Image Video Process., vol. 17, no. 5, pp. 2211–2219,
2023. https://doi.org/10.1007/s11760-022-02436-4
[16] C. Aliferis and G. Simon, “Overfitting, underfitting and general model overconfidence
and under-performance pitfalls and best practices in machine learning and AI,”
Artificial intelligence and machine learning in health care and medical sciences: Best
practices and pitfalls, pp. 477–524, 2024. https://doi.org/10.1007/978-3-031-39355-
6_10
[17] D. Fadillah, E. Haerani, F. Wulandari, and F. Syafria, “Klasifikasi Kondisi Janin
Menggunakan Algoritma K-Nearest Neighbors dan Teknik SMOTE Berdasarkan Data
Kardiotogram,” Bulletin of Computer Science Research, vol. 5, no. 4, pp. 482–489,
2025. https://doi.org/10.47065/bulletincsr.v5i4.585
[18] T. Ahmad, J. Wu, H. S. Alwageed, F. Khan, J. Khan, and Y. Lee, “Human activity
recognition based on deep-temporal learning using convolution neural networks
features and bidirectional gated recurrent unit with features selection,” IEEE access,
vol. 11, pp. 33148–33159, 2023. doi: 10.1109/ACCESS.2023.3263155
[19] Lusi Dwi Anggraini, Umi Mahdiyah, Resty Wulanningrum, “Pengujian Simple CNN
Dengan Arsitektur LeNet5 Pada Penyakit Daun Jagung”, inotek, vol. 9, no. 1, pp. 711–
719, Jul. 2025, doi: 10.29407/9qdmem64.
[20] A. Farzad, H. Mashayekhi, and H. Hassanpour, “A comparative performance analysis
of different activation functions in LSTM networks for classification,” Neural Comput.
Appl., vol. 31, no. 7, pp. 2507–2521, 2019. https://doi.org/10.1007/s00521-017-3210-
6
[21] M. C. Rani, R. A. Dewi, F. D. Azkia, M. Wahyudi, and A. S. Budiman, “Perbandingan
Algoritma Random Forest, Naive Bayes, Dan Neural Network Dalam Klasifikasi
Penyakit Jantung,” Jurnal Sains Informatika Terapan, vol. 4, no. 2, pp. 77–84, 2025.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Copyright on any article is retained by the author(s).
- The author grants the journal, right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgment of the work’s authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal’s published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
- The article and any associated published material is distributed under the Creative Commons Attribution-ShareAlike 4.0 International License