Klasifikasi Teks Mitos dan Fakta Kesehatan Menggunakan Algoritma Random Forest dengan Optimasi Class Weight
DOI:
https://doi.org/10.29407/49rc6579Abstract
Penyebaran misinformasi kesehatan di ruang digital dapat membangun persepsi ancaman yang keliru pada masyarakat, sehingga diperlukan metode klasifikasi teks. Penelitian ini mengklasifikasikan 200 data teks kesehatan ke dalam kelas mitos 52 data dan fakta 148 data menggunakan algoritma Random Forest. Dataset yang digunakan telah divalidasi oleh pakar medis. Alur penelitian meliputi prapemrosesan teks, ekstraksi fitur TF-IDF, serta pelatihan dan pengujian model dengan berbagai rasio data. Optimasi dilakukan menggunakan parameter class_weight untuk mengatasi ketidakseimbangan data. Hasil pengujian menunjukkan dengan rasio 90:10 mendapat nilai akurasi keseluruhan sebesar 85,00%, dengan f1-score kelas fakta sebesar 0.9032 dan kelas mitos sebesar 0.6667. Kontribusi penelitian ini adalah pembuktian bahwa parameter class_weight mampu menjaga performa klasifikasi kelas minoritas tetap berada di atas ambang batas tebakan acak 0,5000. Hasil ini dapat menjadi referensi pemodelan klasifikasi hoaks kesehatan ke depan.
Keywords:
Klasifikasi Teks, Random Forest, Class Weight, TF-IDF, Informasi Kesehatan##plugins.themes.default.displayStats.downloads##
References
[1] D. Salihu, F. Abuadas, M. Alanaz, M. Dhiabat, A. Alkhadam, and H. Obiedallah, "Digital Health Literacy, Technology Services and Health Equity in Developing Economies Concept, Challenges and Recommendations," Majmaah Journal of Health Sciences, vol. 13, no. 4, p. 210, 2025, doi: 10.5455/mjhs.2025.04.015.
[2] M. O. Gallardo and R. Ebardo, "Online Health Information Seeking in Social Media," 2024, pp. 168-179. doi: 10.1007/978-3-031-53731-8 14.
[3] P. G. G. Wahyu, I. Setiawan, and R. I. Saputri, "Gambaran Perilaku Masyarakat Dalam Mencari Informasi Kesehatan Melalui Internet (Studi pada Kecamatan Pasirjambu, Kabupaten Bandung)," Padjadjaran Journal of Dental Researchers and Students, vol. 7, no. 1, p. 81, Mar. 2023, doi: 10.24198/pjdrs.v7i1.40474.
[4] K. Johnson-Arbor, "I Read Health Information Online Can I Trust It?," JAMA Intern. Med., vol. 184, no. 9, p. 1138, Sep. 2024, doi: 10.1001/jamainternmed.2024.0340.
[5] M. M. Ferreira Caceres et al., "The impact of misinformation on the COVID-19 pandemic," AIMS Public Health, vol. 9, no. 2, pp. 262-277, 2022, doi: 10.3934/publichealth.2022018.
[6] L. S. Martinez, "Health Misinformation and Rumors," in The International Encyclopedia of Health Communication, Wiley, 2022, pp. 1-6. doi: 10.1002/9781119678816.iehc0950.
[7] T. Diwantara, S. Harini, and M. Imamudin, "ANALISIS PERBANDINGAN ALGORITMA MACHINE LEARNING UNTUK KLASIFIKASI TUTUPAN LAHAN DI KABUPATEN BANJAR," Technologia: Jurnal Ilmiah, vol. 16, no. 4, pp. 841-848, Oct. 2025, doi: 10.31602/TJI.V16I4.20683.
[8] I. A. Ropikoh, R. Abdulhakim, U. Enri, and N. Sulistiyowati, "Penerapan Algoritma Support Vector Machine (SVM) untuk Klasifikasi Berita Hoax Covid-19," Journal of Applied Informatics and Computing, vol. 5, no. 1, pp. 64-73, Jul. 2021, doi: 10.30871/JAIC.V5I1.3167.
[9] T. A. Roshinta, E. Kumala, and I. V. Dinata, "Sistem Deteksi Berita Hoax Berbahasa Indonesia Bidang Kesehatan," REMIK: Riset dan E-Jurnal Manajemen Informatika Komputer, vol. 7, no. 2, pp. 1167-1173, Apr. 2023, doi: 10.33395/REMIK.V7I2.12369.
[10] M. D. Hendriyanto and B. N. Sari, "PENERAPAN ALGORITMA K-NEAREST NEIGHBOR DALAM KLASIFIKASI JUDUL BERITA HOAX," JURNAL ILMIAH INFORMATIKA, vol. 10, no. 02, pp. 80-84, Sep. 2022, doi: 10.33884/JIF.V10I02.5477.
[11] M. U. Shalih and T. E. E. Tju, "Pembangunan Fitur dalam Identifikasi Cerdas Hoaks dengan Naive Bayes dan Klasifikasi Decision Tree," Jutisi: Jurnal Ilmiah Teknik Informatika dan Sistem Informasi, vol. 13, no. 1, pp. 142-156, Apr. 2024, doi: 10.35889/JUTISI.V13I1.1731.
[12] R. D. Setyadin, R. H. Winasis, and G. Triyono, "Fake News Detection using the Random Forest Algorithm," SISTEMASI, vol. 14, no. 3, pp. 1142-1153, May 2025, doi: 10.32520/STMSI.V14I3.4995.
[13] M. Salmi, D. Atif, D. Oliva, A. Abraham, and S. Ventura, "Handling imbalanced medical datasets: review of a decade of research," Artificial Intelligence Review 2024 57:10, vol. 57, no. 10, pp. 273-, Sep. 2024, doi: 10.1007/S10462-024-10884-2.
[14] J. U. Kazi, "Natural Language Processing (NLP) Basics," in Python Essentials for Biomedical Data Analysis: An Introductory Textbook, Cham: Springer Nature Switzerland, 2025, pp. 463-505. doi: 10.1007/978-3-031-85600-6_12.
[15] P. K. Reshma, S. Rajagopal, and V. L. Lajish, "A Novel Document and Query Similarity Indexing using VSM for Unstructured Documents," in 2020 6th International Conference on Advanced Computing and Communication Systems (ICACCS), IEEE, Mar. 2020, pp. 676-681. doi: 10.1109/ICACCS48705.2020.9074255.
[16] Z. Sun, G. Wang, P. Li, H. Wang, M. Zhang, and X. Liang, "An improved random forest based on the classification accuracy and correlation measurement of decision trees," Expert Syst. Appl., vol. 237, p. 121549, Mar. 2024, doi: 10.1016/J.ESWA.2023.121549.
[17] H. A. Program et al., "Kajian Performa Metode Class Weight Random Forest pada Klasifikasi Imbalance Data Kelas Curah Hujan," Jurnal Sains, Nalar, dan Aplikasi Teknologi Informasi, vol. 3, no. 1, pp. 42-49, Aug. 2023, doi: 10.20885/SNATI.V3I1.30.
[18] E. Helmud, Fitriyani, and P. Romadiana, "Classification Comparison Performance of Supervised Machine Learning Random Forest and Decision Tree Algorithms Using Confusion Matrix," Jurnal Sisfokom (Sistem Informasi dan Komputer), vol. 13, no. 1, pp. 92-97, Feb. 2024, doi: 10.32736/SISFOKOM.V13I1.1985.
[19] D. M. W. Powers and Ailab, "Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation," Oct. 2020, Accessed: Jun. 12, 2026. [Online]. Available: https://arxiv.org/pdf/2010.16061
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Copyright on any article is retained by the author(s).
- The author grants the journal, right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgment of the work’s authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal’s published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
- The article and any associated published material is distributed under the Creative Commons Attribution-ShareAlike 4.0 International License