Error Checker Berbasis Neural Network pada Pembentukan Soal Cerita Matematika Dasar

  • Ferri Nurdiansah (Corresponding Author) Universitas Bhinneka PGRI
  • Agung Prasetya Universitas Bhinneka PGRI
  • Taufiq Agung Cahyono Universitas Bhinneka PGRI
Keywords: Error Checker, IndoBERT, Multi-Label Classification, Neural Network, Soal Cerita Matematika

Abstract

Soal cerita matematika (Math Word Problem/MWP) berperan penting dalam melatih kemampuan penalaran logis siswa Sekolah Dasar. Perkembangan Large Language Model (LLM) telah memungkinkan pembuatan soal secara otomatis, namun soal yang dihasilkan kerap mengandung berbagai jenis kesalahan. Penelitian ini bertujuan: (1) mengembangkan dataset MWP berbahasa Indonesia beranotasi 11 kategori kesalahan secara multi-label; (2) merancang model multi-label classification berbasis IndoBERT dan Feed Forward Neural Network (FFNN) untuk mendeteksi kesalahan secara simultan; serta (3) mengevaluasi kinerja model menggunakan metrik Micro F1-Score, Macro F1-Score, dan Hamming Loss. Dataset dikembangkan melalui error-inducing prompting, menghasilkan 2.517 soal cerita yang dianotasi secara manual oleh dua anotator dengan rasio pembagian 80:20 menggunakan stratified random sampling. Model dilatih selama 40 epoch dengan Binary Cross Entropy loss dan optimizer AdamW. Model optimal diperoleh pada epoch ke-13 dengan validation loss 0,2616. Evaluasi pada testing set menghasilkan Micro F1-Score sebesar 0,7657, Macro F1-Score sebesar 0,7631, dan Hamming Loss sebesar 0,1387. Topic Unsafety mencapai F1-Score sempurna (1,000) dan Misspellings mendekati sempurna (0,986), sementara Grade Mismatch menjadi kategori paling sulit dideteksi (F1 = 0,591). Penelitian ini menyimpulkan bahwa pendekatan multi-label classification berbasis IndoBERT efektif untuk deteksi error pada MWP berbahasa Indonesia.

Downloads

Download data is not yet available.

References

A. Prasetya, "Identifying arithmetic operations in math word problems based on recursive neural network and support vector machine," J. Online Inform., vol. 8, no. 2, 2024, doi: 10.29100/joeict.v8i2.7421.

S. Ariyarathne, D. Athapaththu, S. Ekanayake, and S. Ranathunga, "Elementary math word problem generation using large language models," arXiv:2506.05950, 2025, doi: 10.48550/arXiv.2506.05950.

Q. Zhou and D. Huang, "Towards generating math word problems from equations and topics," in Proc. 12th Int. Conf. Nat. Lang. Gener. (INLG 2019), Tokyo, Japan, 2019, pp. 494-503, doi: 10.18653/v1/W19-8661.

X. Kang, Z. Wang, X. Jin, W. Wang, K. Huang, and Q. Wang, "Template-driven LLM-paraphrased framework for tabular math word problem generation," in Proc. AAAI Conf. Artif. Intell., vol. 39, no. 23, 2025, pp. 24303-24311, doi: 10.1609/aaai.v39i23.34607.

T. Christ, P. Bhatt, S. Bhat, and A. Lan, "MATHWELL: Generating educational math word problems at scale," in Findings of EMNLP 2024, Miami, FL, USA, 2024, pp. 1598-1613, doi: 10.48550/arXiv.2402.15861.

S. Mandal and S. K. Naskar, "Classifying and solving arithmetic math word problems: An intelligent math solver," IEEE Trans. Learn. Technol., vol. 14, no. 1, pp. 28-41, Feb. 2021, doi: 10.1109/TLT.2021.3057805.

J. He-Yueya, G. Poesia, R. E. Wang, and N. D. Goodman, "Solving math word problems by combining language models with symbolic solvers," arXiv:2304.09102, 2023, doi: 10.48550/arXiv.2304.09102.

K. Shridhar, J. Macina, M. El-Assady, T. Sinha, M. Kapur, and M. Sachan, "Automatic generation of socratic subquestions for teaching math word problems," in Proc. EMNLP 2022, Abu Dhabi, UAE, 2022, pp. 4136-4149, doi: 10.18653/v1/2022.emnlp-main.277.

J. Kim, Y. Kim, I. Baek, J. Bak, and J. Lee, "It ain't over: A multi-aspect diverse math word problem dataset," in Proc. 2023 Conf. Empirical Methods Nat. Lang. Process. (EMNLP), Singapore, 2023, pp. 14984-15011, doi: 10.18653/v1/2023.emnlp-main.927.

Z. Xie, K. Zhou, S. Shi, and S. Shi, "Adversarial math word problem generation," in Proc. AAAI Conf. Artif. Intell., vol. 38, no. 17, 2024, pp. 19238-19246, doi: 10.1609/aaai.v38i17.29900.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. NAACL-HLT 2019, Minneapolis, MN, USA, 2019, pp. 4171-4186, doi: 10.18653/v1/N19-1423.

F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, "IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP," in Proc. 28th COLING, Barcelona, Spain, 2020, pp. 757-770, doi: 10.18653/v1/2020.coling-main.66.

N. K. Nissa and E. Yulianti, "Multi-label text classification of Indonesian customer reviews using bidirectional encoder representations from transformers language model," Int. J. Electr. Comput. Eng., vol. 13, no. 5, pp. 5641-5652, 2023, doi: 10.11591/ijece.v13i5.pp5641-5652.

G. Z. Nabiilah, I. Nur, E. S. Purwanto, and M. F. Hidayat, "Indonesian multilabel classification using IndoBERT embedding and MBERT classification," Int. J. Electr. Comput. Eng., vol. 14, no. 1, pp. 1071-1078, 2024, doi: 10.11591/ijece.v14i1.pp1071-1078.

N. C. Mei, S. Tiun, and G. Sastria, "Multi-label aspect-sentiment classification on Indonesian cosmetic product reviews with IndoBERT model," Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 11, 2024, doi: 10.14569/IJACSA.2024.0151168.

Y. Sagama and A. Alamsyah, "Multi-label classification of Indonesian online toxicity using BERT and RoBERTa," in Proc. IAICT 2023, Bandung, Indonesia, 2023, pp. 264-269, doi: 10.1109/IAICT59002.2023.10205892.

Y. K. Sari, J. Iskandar, and A. Prasetya, "Penerapan metode Naive Bayes dalam analisis sentimen terhadap cyberbullying," JUSTER: J. Sains dan Terapan, vol. 5, no. 1, pp. 64-73, 2026, doi: 10.57218/juster.v5i1.2556.

Published
2026-09-21
How to Cite
Nurdiansah, F., Prasetya, A., & Cahyono, T. A. (2026). Error Checker Berbasis Neural Network pada Pembentukan Soal Cerita Matematika Dasar. Journal of Artificial Intelligence and Technology Information (JAITI), 4(3), 487-494. https://doi.org/10.58602/jaiti.v4i3.316