Educational data mining for informatics grade prediction using a hybrid K-means framework

Authors

  • Roselilie Simbulan Universitas Teknologi Yogyakarta
  • Arief Hermawan Universitas Teknologi Yogyakarta
  • Donny Avianto Universitas Teknologi Yogyakarta

DOI:

https://doi.org/10.31571/saintek.v15i1.10728

Keywords:

Educational Data Mining, K-Means Clustering, Random Forest, Logistic Regression, Grade Prediction

Abstract

This study proposes a two-phase hybrid Educational Data Mining (EDM) framework that integrates K-Means clustering with supervised classification to predict students' final grade categories in Informatics within the Kurikulum Merdeka competency-based assessment system. The dataset consists of 281 Grade X students, all of whom achieved scores above the Minimum Achievement Criterion (KKTP = 75), with prediction focused on three grade categories: A (≥86), B (80–85), and C (75–79). Four summative assessment scores (S1, S7, S8, and S9) were used as input features. In the first phase, K-Means generated three clusters (Silhouette Score = 0.3198), and the resulting cluster labels were added as an additional feature. In the second phase, Random Forest and Logistic Regression were optimized using Grid Search with 5-Fold Stratified Cross-Validation, while SMOTE was employed to address class imbalance. The results show that Logistic Regression outperformed Random Forest, achieving a test accuracy of 59.65% and a Macro F1-Score of 0.5899, whereas Random Forest achieved 49.12% accuracy and a Macro F1-Score of 0.4799 and exhibited signs of overfitting. Feature importance analysis identified S7, S8, and S9 as the most influential predictors, while the cluster-derived feature contributed more strongly to Random Forest than to Logistic Regression. These findings suggest that well-regularized linear models may generalize better than ensemble methods on small datasets with narrow score distributions. The proposed framework is best positioned as a screening-support tool for early formative intervention in competency-based educational settings.

Downloads

Download data is not yet available.

Author Biographies

Roselilie Simbulan, Universitas Teknologi Yogyakarta

Magister of Technology Information

Arief Hermawan, Universitas Teknologi Yogyakarta

Magister of Technology Information

Donny Avianto, Universitas Teknologi Yogyakarta

Magister of Technology Information

References

Alsariera, Y. A., Baashar, Y., Alkawsi, G., Mustafa, A., Alkahtani, A. A., & Ali, N. (2022). Assessment and evaluation of different machine learning algorithms for predicting student performance. Computational Intelligence and Neuroscience, 2022. https://doi.org/10.1155/2022/4151487

Bulut, O., Tan, B., Mazzullo, E., & Syed, A. (2025). Benchmarking variants of recursive feature elimination: Insights from predictive tasks in education and healthcare. MDPI (Information), 16(476), 1–21. https://doi.org/https://doi.org/10.3390/info16060476

Dimić, G., & Pecić, L. (2024). Proposed approach for overcoming the impact of unbalanced distribution in predicting students ’ performance. TEM Journal, 13(4), 2839–2849. https://doi.org/10.18421/TEM134

Evangelista, E. D. L., & Sy, B. D. (2022). An approach for improved students ’ performance

prediction using homogeneous and heterogeneous ensemble methods. International Journal of Electrical and Computer Engineering (IJECE), 12(5), 5226–5235. https://doi.org/10.11591/ijece.v12i5.pp5226-5235

Khairy, D., Alharbi, N., Amasha, M. A., & Areed, M. F. (2024). Prediction of student exam performance using data mining classification algorithms. Education and Information Technologies, 29(16), 21621–21645. https://doi.org/10.1007/s10639-024-12619-w

Malik, S., Mahanty, C., Mohanty, J., Vaghela, K., Narmadha, T., Sivaranjani, R., Bhutto, J. K., Khan, A., & Zewdie, A. (2025). Enhancing education quality with hybrid clustering and evolutionary neural networks in a multi phase framework. Scientific Reports, 15(21323), 1–18. https://doi.org/https://doi.org/10.1038/s41598-025-04622-z

Men, Y., & Zhao, M. (2025). Hybrid machine learning techniques for improving student management and academic performance. International Journal of Information and Communication Technology, 26(17). https://doi.org/10.1504/ijict.2025.10071317

Namoun, A., & Alshanqiti, A. (2021). Predicting student performance using data mining and learning analytics techniques: A systematic literature review. MDPI (Applied Sciences), 11(237), 1–28. https://doi.org/https://doi.org/10.3390/app11010237

Odoom, S., Osei, E. O., Effah, E. Q., Bakariwie, A., & Rufai, A. A. (2026). Intelligent profiling beyond grades using a stacking ensemble framework for student success prediction. Springer Nature, 5(242), 1–19. https://doi.org/https://doi.org/10.1007/s44217-026-01277-4

Opitz, J. (2024). A closer look at classification evaluation metrics and a critical reflection of common evaluation practice. Transactions of the Association for Computational Linguistics, 12(2018), 820–836. https://doi.org/https://doi.org/10.1162/tacl_a_00675

Pamungkas, L., Dewi, N. A., & Putri, N. A. (2024). Classification of Student Grade Data Using the K-Means Clustering Method. Jurnal Sisfokom (Sistem Informasi Dan Komputer), 13(1), 86–91. https://doi.org/10.32736/sisfokom.v13i1.1983

Raj, S. A. P., & Vidyaathulasiraman. (2021). Determining optimal number of K for e-learning groups clustered using K-medoid. (IJACSA) International Journal of Advanced Computer Science and Applications, 12(6), 400–407. https://doi.org/https://dx.doi.org/10.14569/IJACSA.2021.0120644

Roslan, M. H., & Chen, C. J. (2022). Educational data mining for student performance prediction : A systematic literature review (2015-2021). International Journal of Emerging Technologies in Learning, 17(05), 207–225. https://doi.org/https://doi.org/10.3991/ijet.v17i05.27685

Sayed, H. E. F., & Fouad, M. M. (2025). Hybrid machine learning models for enhanced student performance prediction. SciNexuses, 2, 167–183.

Scornet, E. (2023). Trees, forests, and impurity-based variable importance in regression. Institute of Mathematical Statistics (IMS) Association Des Publications de l’Institut Henri Poincaré., 59(1), 21–52. https://doi.org/https://doi.org/10.1214/21-AIHP1240

Srinivasulu, A., & Palanisamy, V. (2024). Data-Driven revolution in academic support for mathematics underachievers through random forest individual and hybrid model. Journal of Artificial Intelligence and System Modelling, 02(03), 1–21. https://doi.org/https://doi.org/10.22034/jaism.2024.469529.1048

Wongoutong, C. (2024). The impact of neglecting feature scaling in k-means clustering. Public Library of Science (PLOS), 19(12), 1–19. https://doi.org/10.1371/journal.pone.0310839

Wongvorachan, T., He, S., & Bulut, O. (2023). A comparison of undersampling, oversampling, and SMOTE methods for dealing with imbalanced classification in educational data mining. MDPI (Information), 14(54), 1–15. https://doi.org/https://doi.org/10.3390/info14010054

Wu, M. (2026). K-Means clustering-based feature generation for student performance prediction. ICCK Transactions on Educational Data Mining, 2, 14–28. https://doi.org/https://doi.org/10.62762/TEDM.2026.716076

Yağcı, M. (2022). Educational data mining: prediction of students ’ academic performance using machine learning algorithms. Smart Learning Environments, 9(11). https://doi.org/10.1186/s40561-022-00192-z

Yanyu, G., Jizu, L., Cliff, D., & Huayun, D. (2025). Classification and prediction of miners’ emergency response competence under sudden events based on seed k-means and stacking learning algorithm. Engineered Science, 38(1877), 1–26. https://doi.org/https://dx.doi.org/10.30919/es1877

Zhang, X., Zhang, Y., Chen, A. L., Yu, M., & Lihao, Z. (2025). Optimizing multi label student performance prediction with GNN-TINet : A contextual multidimensional deep learning framework. Public Library of Science (PLOS) ONE, 20(1), 1–25. https://doi.org/10.1371/journal.pone.03

Downloads

Published

2026-06-30

How to Cite

Simbulan, R., Hermawan, A., & Avianto, D. (2026). Educational data mining for informatics grade prediction using a hybrid K-means framework. Jurnal Pendidikan Informatika Dan Sains, 15(1), 88–97. https://doi.org/10.31571/saintek.v15i1.10728