EVALUASI KOMPARATIF MANUSIA, CHATGPT, DAN GEMINI DALAM MENGANALISIS POLA KALIMAT PADA TEKS AKADEMIK
DOI:
https://doi.org/10.29407/ka5fbn51Keywords:
Comparison, ChatGPT, Gemini, SentencesAbstract
The increasing interest of high school students in utilizing LLM-based applications (ChatGPT and Gemini) as learning tools needs to be evaluated. Sentence analysis results from ChatGPT and Gemini, which are frequently used by students in their work, require careful review. Therefore, a comparative evaluation of the sentence analysis results from ChatGPT and Gemini is important to determine the reliability of the two LLM applications. This study presents the results of a comparative evaluation of the ChatGPT and Gemini applications in analyzing sentence patterns, based on the analysis of Indonesian linguists as the “gold standard”. A comparative quantitative approach was used in this study. The data used were 139 sentences from academic texts found in the grade XI student textbooks curated by the Ministry of Education, Culture, Research, and Technology. Agreement classification tests and performance tests were conducted between ChatGPT and Gemini. The level of agreement for ChatGPT was classified as almost perfect with a coefficient of 0.843, while Gemini's agreement was classified as moderate with a coefficient of 0.597. The ChatGPT evaluation results showed an accuracy of 89.93%, a macro precision of 91.70%, a macro recall of 87.97%, and a macro F1-score of 89.21%. In ChatGPT, classification errors focused on distinguishing between compound-equivalent and compound-equivalent sentences. Furthermore, the Gemini evaluation results showed an accuracy of 73.38%, a macro precision of 79.77%, a macro recall of 72.42%, and a macro F1-score of 66.99%. In Gemini, classification errors focused on the classification of compound-equivalent sentences. Thus, the analysis by Indonesian linguists, as the gold standard, indicates that ChatGPT's suitability and performance are superior to Gemini's in sentence analysis.
References
Bavaresco, A., Bernardi, R., Bertolazzi, L., Elliott, D., Fernández, R., Gatt, A., Ghaleb, E., Giulianelli, M., Hanna, M., Koller, A., Martins, A. F. T., Mondorf, P., Neplenbroek, V., Pezzelle, S., Plank, B., Schlangen, D., Suglia, A., Surikuchi, A. K., Takmaz, E., & Testoni, A. (2025). LLMs instead of Human Judges ? A Large Scale Empirical Study across 20 NLP Evaluation Tasks. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2, 238–255. https://doi.org/10.18653/v1/2025.acl-short.20
Chen, X., Alexopoulou, T., & Tsimpli, I. (2021). Automatic Extraction of Subordinate Clauses and Its Application in Second Language Acquisition Research. Behavior Research Methods, 53(2), 803–817. https://doi.org/10.3758/s13428-020-01456-7
Damayanti, S., & Zulkarnain. (2026). Analisis Efektivitas Large Language Model (LLM) sebagai Asisten Praktikum pada Mata Kuliah Jaringan Komputer. Klik - Jurnal Ilmu Komputer, 7(1), 30–38. https://doi.org/10.56869/klik.v7i1.778
Dewi Marhamah Syadli, & Dewi Anggraini. (2023). Kalimat Efektif dalam Teks Eksplanasi Siswa Kelas XI SMA Negeri 5 Sijunjung. Jurnal Pendidikan, Bahasa Dan Budaya, 2(2), 01–11. https://doi.org/10.55606/jpbb.v2i2.1350
Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of Metacognitive Laziness: Effects of Generative Artificial Intelligence on Learning Motivation, Processes, and Performance. British Journal of Educational Technology, 56(2), 489–530. https://doi.org/10.1111/bjet.13544
J Richard Landis, & Gary G Koch. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1), 159–174. https://doi.org/https://doi.org/10.2307/2529310
Kiranawati, B. I., Latifah, U., & Khusnah, N. K. (2025). Pemanfaatan Model AI dalam Analisis Sintaksis Bahasa Indonesia: Kajian Linguistik Komputasional. JURNAL PROGRESIF, 3, 7–26. https://journal.univgresik.ac.id/index.php/progresif/article/view/443
Kurniahtunnisa, Manuel, M. Y., Aini, M., & Agustina, T. P. (2025). Persepsi dan Sikap Siswa terhadap Penggunaan Artificial Intelligence. Scholaria: Jurnal Pendidikan Dan Kebudayaan, 15(1), 47–59. https://doi.org/10.24246/j.js.2025.v15.i1.p47-59
Li, M. Y., Liu, A., Wu, Z., & Smith, N. A. (2024). A Taxonomy of Ambiguity Types for NLP. https://doi.org/https://doi.org/10.48550/arXiv.2403.14072
Nainggolan, K. (2025). Sintakmatik dalam Analisis Struktur Kalimat Bahasa Indonesia. JBSI: Jurnal Bahasa Dan Sastra Indonesia, 5(01), 215–222. https://doi.org/10.47709/jbsi.v5i01.6247
Nur Gimas Tiaroh, Lutfi Anifatun Wahidah, Yuwen Nurkhalifah, Atikah Nuha Hibatullah, Isnaini Muyassaroh, Asep Purwo Yudi Utomo, Rossi Galih Kesuma, & Zulfa Fahmy. (2026). Konjungsi Koordinatif dan Konjungsi Subordinatif dalam Teks Berita Kompas Terbitan September 2025. Dinamika Pembelajaran : Jurnal Pendidikan Dan Bahasa, 3(1), 98–120. https://doi.org/10.62383/dilan.v3i1.2909
Owens, R. E., Pavelko, S. L., & Hahs-Vaughn, D. (2024). Growth of Complex Syntax: Coordinate and Subordinate Clause Use in Elementary School–Aged Children. Language, Speech, and Hearing Services in Schools, 55(3), 714–723. https://doi.org/10.1044/2024_LSHSS-23-00102
Rosyada, A., Putri, A. N., Irawati, Y., Sari, R. R., Ghifari, I. Al, Utomo, A. P. Y., & Kesuma, R. G. (2025). Analisis Jenis Konjungsi Koordinatif pada Artikel Pendidikan dan Edukasi dalam Website kelasjuara.id Edisi Oktober 2024 sebagai Bahan Ajar Siswa SMA. PUSTAKA: Jurnal Bahasa Dan Pendidikan, 5(3), 22–46. https://doi.org/https://doi.org/10.56910/pustaka.v5i3.3262
Suganda, A. (2023). Memilih AI yang Tepat untuk Guru: Perbandingan Fitur Gemini, ChatGPT, dan Claude A. Jurnal Inovasi Teknologi Dan Edukasi Teknik, 3(11), 2023. https://doi.org/10.17977/um084.v3.i11.2023.2
Wahyuni, R. N., Suryani, I., & Warni, W. (2024). Pemanfaatan Gemini dalam Pembelajaran Bahasa Indonesia di SMA: Studi Literatur. Diskursus: Jurnal Pendidikan Bahasa Indonesia, 7(3), 446. https://doi.org/10.30998/diskursus.v7i3.26571
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
