IMPLEMENTASI INDOBERT DENGAN AUGMENTASI DATA SINTETIS BERBASIS LARGE LANGUAGE MODEL UNTUK PENILAIAN ESAI OTOMATIS BAHASA INDONESIA
DOI:
https://doi.org/10.61722/jipm.v4i6.3132Keywords:
Automated Essay Scoring, IndoBERT, Augmentasi Data Sintetis, Quadratic Weighted Kappa, Large Language ModelAbstract
Penilaian esai secara manual membutuhkan waktu yang lama dan rentan terhadap inkonsistensi antarpenilai, sementara pengembangan sistem Automated Essay Scoring (AES) untuk Bahasa Indonesia masih terkendala keterbatasan dataset berlabel. Penelitian ini mengimplementasikan model IndoBERT dengan pendekatan regresi untuk menilai jawaban esai siswa dan menerapkan augmentasi data sintetis berbasis Large Language Model (LLM) guna mengatasi ketidakseimbangan kelas pada kategori skor rendah. Dataset primer terdiri atas 556 jawaban esai siswa kelas VII pada empat butir soal mata pelajaran Pendidikan Kewarganegaraan yang dinilai oleh guru, dan menyisakan 551 data setelah pembersihan. Augmentasi menggunakan GPT-4o-mini menghasilkan 160 jawaban sintetis pada kategori skor 0 dan 1 sehingga dataset meningkat menjadi 711 data. Model dievaluasi menggunakan 5-fold cross-validation dengan metrik Quadratic Weighted Kappa (QWK), Mean Absolute Error (MAE), dan korelasi Spearman pada empat kondisi pengujian yang membedakan butir soal terlatih dan tidak terlatih. Hasil menunjukkan model hasil augmentasi memperoleh QWK 0,8942 dan Spearman 0,9063, meningkat dari 0,7884 dan 0,8199 pada model tanpa augmentasi, meskipun MAE sedikit memburuk dari 0,3931 menjadi 0,4113. Namun, ketika model yang sama diuji pada 1.000 jawaban dari 100 butir soal yang belum pernah dilatihkan, QWK turun menjadi 0,4832. Temuan ini menunjukkan bahwa ketersediaan butir soal pada data latih merupakan penentu performa yang lebih dominan dibandingkan sumber maupun jumlah data latih..
References
Aisyah, N., Al Kautsar, M. D., Hidayat, A., Chowdhury, R., & Koto, F. (2025). Evaluating Vision–Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms. arXiv. https://doi.org/10.48550/arXiv.2506.04822
Aliyah, N. E., Sholikah, R. W., Firdausi, H., Ciptaningtyas, H. T., & Sabilla, I. A. (2025). Enhancing Automated Essay Scoring in Bahasa Indonesia with IndoBERT and IndoSBERT. 2025 International Conference on Smart Computing, IoT and Machine Learning (SIML), 1–7. https://doi.org/10.1109/SIML65326.2025.11080721
Amalia, A., Lydia, M. S., Muchtar, M. A., Manik, F. Y., & Gunawan, D. (2025). Mitigating Bias and Assessment Inconsistencies with BERT-Based Automated Short Answer Grading for the Indonesian Language. IAENG International Journal of Computer Science, 52(3).
Badran, N., Le, J., Le, T., & Uchiya, T. (2025). Improving Stress Detection with Synthetic Datasets: GPT-4o-Mini and Transformer Model Evaluation. Dalam N. T. Nguyen et al. (Ed.), Intelligent Information and Database Systems (ACIIDS 2025), Lecture Notes in Computer Science, vol. 15683. Springer. https://doi.org/10.1007/978-981-96-6008-7_9
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. http://arxiv.org/abs/1810.04805
Doewes, A. (2026). Rethinking Automated Essay Scoring: Agreement, Fairness, and Feedback [Disertasi doktoral, Eindhoven University of Technology].
Doewes, A., Kurdhi, N. A., & Saxena, A. (2023). Evaluating Quadratic Weighted Kappa as the Standard Performance Metric for Automated Essay Scoring. Proceedings of the International Conference on Educational Data Mining, 103–113. https://doi.org/10.5281/zenodo.8115784
Geni, L., Yulianti, E., & Sensuse, D. I. (2023). Sentiment Analysis of Tweets Before the 2024 Elections in Indonesia Using Bert Language Models. Jurnal Ilmiah Teknik Elektro Komputer dan Informatika, 9(3), 746–757. https://doi.org/10.26555/jiteki.v9i3.26490
Judijanto, L., Abdullah, G., Abdurahman, A., Lumbu, A., Tumober, R. T., Septikasari, D., Sogalrey, F. A. M., Mahliatussikah, H., & Subhaktiyasa, P. G. (2025). Evaluasi Pembelajaran: Prinsip, Teknik, dan Aplikasi. Sonpedia Publishing.
Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. Proceedings of the 28th International Conference on Computational Linguistics, 757–770.
Kurniawati, F. E., & Mega Pradnya, W. (2020). Implementasi Algoritma Winnowing pada Sistem Penilaian Otomatis Jawaban Esai pada Ujian Online Berbasis Web. Jurnal Teknik Komputer AMIK BSI, 6(2). https://doi.org/10.31294/jtk.v4i2
Li, W., & Liu, H. (2024). Applying large language models for automated essay scoring for non-native Japanese. Humanities and Social Sciences Communications, 11(1). https://doi.org/10.1057/s41599-024-03209-9
Miller, I., Choe, Y., & Emirtekin, E. (2025). Large Language Model-Powered Automated Assessment: A Systematic Review. Applied Sciences, 15(10), 5683. https://doi.org/10.3390/app15105683
Misgna, H., On, B. W., Lee, I., & Choi, G. S. (2025). A survey on deep learning-based automated essay scoring and feedback generation. Artificial Intelligence Review, 58(2). https://doi.org/10.1007/s10462-024-11017-5
Møller, A. G., Pera, A., Dalsgaard, J., & Aiello, L. (2024). The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), 179–192. https://doi.org/10.18653/v1/2024.eacl-short.17
Page, E. B. (1994). Computer Grading of Student Prose, Using Modern Concepts and Software. The Journal of Experimental Education, 62(2), 127–142. https://doi.org/10.1080/00220973.1994.9943835
Pradani, K. A., & Suadaa, L. H. (2023). Automated Essay Scoring Menggunakan Semantic Textual Similarity Berbasis Transformer untuk Penilaian Ujian Esai. Jurnal Teknologi Informasi dan Ilmu Komputer, 10(6), 1177–1184. https://doi.org/10.25126/jtiik.2023107338
Prastyo, P. H., Tungadi, E., & Zuhdi, S. (2025). Indonesian Automated Essay Scoring: A Comparative Study of Pretrained Transformer Models. Information Technology Education Journal, 120–130. https://doi.org/10.59562/intec.v4i2.8069
Putri, H., Susiani, D., Wandani, N. S., & Putri, F. A. (2022). Instrumen Penilaian Hasil Pembelajaran Kognitif pada Tes Uraian dan Tes Objektif. Jurnal Papeda, 4(2).
Rahanra, N., Hossam, A., & Scott, J. (2025). Utilizing Artificial Intelligence (AI) for Automated Feedback on the English Essay Writing Skills of Indonesian University Students. Journal International of Lingua and Technology, 4(3), 260–275. https://doi.org/10.55849/jiltech.v4i3.1121
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems, 30. http://arxiv.org/abs/1706.03762
Winarta, I. M. W. P., Alfarozi, S. A. I., & Hidayah, I. (2024). Atomic Evaluation using Large Language Model for Automated Essay Exam Scoring. 2024 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT), 21–27. https://doi.org/10.1109/COMNETSAT63286.2024.10862459
Zubaidi, A., Munip, A., Widodo, S. A., & Zerrouki, T. (2025). Enhancing Arabic writing skills using ChatGPT-based AI learning models: A tridimensional human-AI collaboration framework. Indonesian Journal of Applied Linguistics, 15(1), 87–101. https://doi.org/10.17509/ijal.v15i1.75378
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 JURNAL ILMIAH PENELITIAN MAHASISWA

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.











