Penerapan Metode Weak Supervision dan Open Information Extraction untuk Ekstraksi Triplet Semantik pada Subtitle Youtube Otomotif
DOI:
https://doi.org/10.29407/sxp7wm39Abstract
Pencarian informasi teknis perawatan otomotif melalui transkripsi otomatis YouTube sering terhambat oleh durasi video yang panjang dan derau bahasa lisan . Penelitian ini bertujuan mengembangkan sistem otomatis untuk mengekstrak pengetahuan diagnostik menggunakan kombinasi metode Weak Supervision dan Open Information Extraction. Data primer terdiri dari 1.790 dokumen subtitle YouTube, sedangkan data sekunder berupa kamus acuan istilah otomotif. Hasil pengujian terhadap 10 dokumen uji menunjukkan performa optimal dengan nilai Precision 100 persen, Recall 96,77 persen, dan F1-Score 98,36 persen. Sistem juga berhasil diintegrasikan ke dalam antarmuka web berbasis Streamlit untuk visualisasi data secara interaktif. Penelitian ini berkontribusi pada bidang pengolahan data berupa pengumpulan informasi efisien tanpa pemeriksaan manual. Celah informasi terlewat berupa dua False Negative dipicu oleh tidak terdeteksinya entitas suku cadang pada kamus acuan.
Keywords:
ekstraksi informasi, otomotif, open information extraction, weak supervision, youtube##plugins.themes.default.displayStats.downloads##
References
[1] A. Hotho, A. Nürnberger, and G. Paaß, “A Brief Survey of Text Mining.”
[Online]. Available: http://www.crisp-dm.org/
[2] Ronen. Feldman and James. Sanger, The text mining handbook : advanced
approaches in analyzing unstructured data. Cambridge University Press, 2007.
[3] B. Farhadi, “Enriching Subtitled YouTube Media Fragments via Utilization of
the Web-Based Natural Language Processors and Efficient Semantic Video Annotations,”
2013. [Online]. Available: http://www.gjset.org
[4] D. Jurafsky and J. H. Martin, “Speech and Language Processing An Introduction
to Natural Language Processing, Computational Linguistics, and Speech Recognition with
Language Models Third Edition draft.”
[5] R. P. Kusumawardani and K. N. Kusumawati, “Named entity recognition in the
medical domain for Indonesian language health consultation services using bidirectionallstmcrf algorithm,” in Procedia Computer Science, Elsevier B.V., 2024, pp. 1146–1156.
doi: 10.1016/j.procs.2024.10.344.
[6] H. Darji, J. Mitrović, and M. Granitzer, “German BERT Model for Legal Named
Entity Recognition,” Mar. 2023, doi: 10.5220/0011749400003393.
[7] E. Agichtein and L. Gravano, “Snowball: Extracting Relations from Large PlainText Collections.”
[8] A. Ratner, S. H. Bach, H. Ehrenberg, J. Fries, S. Wu, and C. Ré, “Snorkel: Rapid
training data creation with weak supervision,” in Proceedings of the VLDB Endowment,INOTEK, Vol. 10
ISSN: 2580-3336 (Print) / 2549-7952 (Online)
Url: https://proceeding.unpkediri.ac.id/index.php/inotek/
Prosiding SEMNAS INOTEK (Seminar Nasional Inovasi Teknologi) 2026 1048
Association for Computing Machinery, Nov. 2017, pp. 269–282. doi:
10.14778/3157794.3157797.
[9] S. Pawar, G. K. Palshikar, A. Jain, J. Bhat, and S. Johnson, “Weak Supervision
using Linguistic Knowledge for Information Extraction,” NLPAI.
[10] M. Banko, M. J. Cafarella, S. Soderland, M. Broadhead, and O. Etzioni, “Open
Information Extraction from the Web.”
[11] C. Niklaus, M. Cetto, A. Freitas, and S. Handschuh, “A Survey on Open
Information Extraction,” Jun. 2018, [Online]. Available: http://arxiv.org/abs/1806.05599
[12] L. Cui, F. Wei, and M. Zhou, “Neural Open Information Extraction.” [Online].
Available: https://1drv.ms/u/s!ApPZx_
[13] S. Pawar, G. K. Palshikar, and P. Bhattacharyya, “Relation Extraction : A Survey,”
Dec. 2017, [Online]. Available: http://arxiv.org/abs/1712.05191
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Copyright on any article is retained by the author(s).
- The author grants the journal, right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgment of the work’s authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal’s published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
- The article and any associated published material is distributed under the Creative Commons Attribution-ShareAlike 4.0 International License