Educational data mining for informatics grade prediction using a hybrid K-means framework
DOI:
https://doi.org/10.31571/saintek.v15i1.10728Keywords:
Educational Data Mining, K-Means Clustering, Random Forest, Logistic Regression, Grade PredictionAbstract
This study proposes a two-phase hybrid Educational Data Mining (EDM) framework that integrates K-Means clustering with supervised classification to predict students' final grade categories in Informatics within the Kurikulum Merdeka competency-based assessment system. The dataset consists of 281 Grade X students, all of whom achieved scores above the Minimum Achievement Criterion (KKTP = 75), with prediction focused on three grade categories: A (≥86), B (80–85), and C (75–79). Four summative assessment scores (S1, S7, S8, and S9) were used as input features. In the first phase, K-Means generated three clusters (Silhouette Score = 0.3198), and the resulting cluster labels were added as an additional feature. In the second phase, Random Forest and Logistic Regression were optimized using Grid Search with 5-Fold Stratified Cross-Validation, while SMOTE was employed to address class imbalance. The results show that Logistic Regression outperformed Random Forest, achieving a test accuracy of 59.65% and a Macro F1-Score of 0.5899, whereas Random Forest achieved 49.12% accuracy and a Macro F1-Score of 0.4799 and exhibited signs of overfitting. Feature importance analysis identified S7, S8, and S9 as the most influential predictors, while the cluster-derived feature contributed more strongly to Random Forest than to Logistic Regression. These findings suggest that well-regularized linear models may generalize better than ensemble methods on small datasets with narrow score distributions. The proposed framework is best positioned as a screening-support tool for early formative intervention in competency-based educational settings.
Downloads
References
Alsariera, Y. A., Baashar, Y., Alkawsi, G., Mustafa, A., Alkahtani, A. A., & Ali, N. (2022). Assessment and evaluation of different machine learning algorithms for predicting student performance. Computational Intelligence and Neuroscience, 2022. https://doi.org/10.1155/2022/4151487
Bulut, O., Tan, B., Mazzullo, E., & Syed, A. (2025). Benchmarking variants of recursive feature elimination: Insights from predictive tasks in education and healthcare. MDPI (Information), 16(476), 1–21. https://doi.org/https://doi.org/10.3390/info16060476
Dimić, G., & Pecić, L. (2024). Proposed approach for overcoming the impact of unbalanced distribution in predicting students ’ performance. TEM Journal, 13(4), 2839–2849. https://doi.org/10.18421/TEM134
Evangelista, E. D. L., & Sy, B. D. (2022). An approach for improved students ’ performance
prediction using homogeneous and heterogeneous ensemble methods. International Journal of Electrical and Computer Engineering (IJECE), 12(5), 5226–5235. https://doi.org/10.11591/ijece.v12i5.pp5226-5235
Khairy, D., Alharbi, N., Amasha, M. A., & Areed, M. F. (2024). Prediction of student exam performance using data mining classification algorithms. Education and Information Technologies, 29(16), 21621–21645. https://doi.org/10.1007/s10639-024-12619-w
Malik, S., Mahanty, C., Mohanty, J., Vaghela, K., Narmadha, T., Sivaranjani, R., Bhutto, J. K., Khan, A., & Zewdie, A. (2025). Enhancing education quality with hybrid clustering and evolutionary neural networks in a multi phase framework. Scientific Reports, 15(21323), 1–18. https://doi.org/https://doi.org/10.1038/s41598-025-04622-z
Men, Y., & Zhao, M. (2025). Hybrid machine learning techniques for improving student management and academic performance. International Journal of Information and Communication Technology, 26(17). https://doi.org/10.1504/ijict.2025.10071317
Namoun, A., & Alshanqiti, A. (2021). Predicting student performance using data mining and learning analytics techniques: A systematic literature review. MDPI (Applied Sciences), 11(237), 1–28. https://doi.org/https://doi.org/10.3390/app11010237
Odoom, S., Osei, E. O., Effah, E. Q., Bakariwie, A., & Rufai, A. A. (2026). Intelligent profiling beyond grades using a stacking ensemble framework for student success prediction. Springer Nature, 5(242), 1–19. https://doi.org/https://doi.org/10.1007/s44217-026-01277-4
Opitz, J. (2024). A closer look at classification evaluation metrics and a critical reflection of common evaluation practice. Transactions of the Association for Computational Linguistics, 12(2018), 820–836. https://doi.org/https://doi.org/10.1162/tacl_a_00675
Pamungkas, L., Dewi, N. A., & Putri, N. A. (2024). Classification of Student Grade Data Using the K-Means Clustering Method. Jurnal Sisfokom (Sistem Informasi Dan Komputer), 13(1), 86–91. https://doi.org/10.32736/sisfokom.v13i1.1983
Raj, S. A. P., & Vidyaathulasiraman. (2021). Determining optimal number of K for e-learning groups clustered using K-medoid. (IJACSA) International Journal of Advanced Computer Science and Applications, 12(6), 400–407. https://doi.org/https://dx.doi.org/10.14569/IJACSA.2021.0120644
Roslan, M. H., & Chen, C. J. (2022). Educational data mining for student performance prediction : A systematic literature review (2015-2021). International Journal of Emerging Technologies in Learning, 17(05), 207–225. https://doi.org/https://doi.org/10.3991/ijet.v17i05.27685
Sayed, H. E. F., & Fouad, M. M. (2025). Hybrid machine learning models for enhanced student performance prediction. SciNexuses, 2, 167–183.
Scornet, E. (2023). Trees, forests, and impurity-based variable importance in regression. Institute of Mathematical Statistics (IMS) Association Des Publications de l’Institut Henri Poincaré., 59(1), 21–52. https://doi.org/https://doi.org/10.1214/21-AIHP1240
Srinivasulu, A., & Palanisamy, V. (2024). Data-Driven revolution in academic support for mathematics underachievers through random forest individual and hybrid model. Journal of Artificial Intelligence and System Modelling, 02(03), 1–21. https://doi.org/https://doi.org/10.22034/jaism.2024.469529.1048
Wongoutong, C. (2024). The impact of neglecting feature scaling in k-means clustering. Public Library of Science (PLOS), 19(12), 1–19. https://doi.org/10.1371/journal.pone.0310839
Wongvorachan, T., He, S., & Bulut, O. (2023). A comparison of undersampling, oversampling, and SMOTE methods for dealing with imbalanced classification in educational data mining. MDPI (Information), 14(54), 1–15. https://doi.org/https://doi.org/10.3390/info14010054
Wu, M. (2026). K-Means clustering-based feature generation for student performance prediction. ICCK Transactions on Educational Data Mining, 2, 14–28. https://doi.org/https://doi.org/10.62762/TEDM.2026.716076
Yağcı, M. (2022). Educational data mining: prediction of students ’ academic performance using machine learning algorithms. Smart Learning Environments, 9(11). https://doi.org/10.1186/s40561-022-00192-z
Yanyu, G., Jizu, L., Cliff, D., & Huayun, D. (2025). Classification and prediction of miners’ emergency response competence under sudden events based on seed k-means and stacking learning algorithm. Engineered Science, 38(1877), 1–26. https://doi.org/https://dx.doi.org/10.30919/es1877
Zhang, X., Zhang, Y., Chen, A. L., Yu, M., & Lihao, Z. (2025). Optimizing multi label student performance prediction with GNN-TINet : A contextual multidimensional deep learning framework. Public Library of Science (PLOS) ONE, 20(1), 1–25. https://doi.org/10.1371/journal.pone.03
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Roselilie Simbulan, Arief Hermawan, Donny Avianto

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
In submitting the manuscript to the journal, the authors certify that:
- They are authorized by their co-authors to enter into these arrangements.
- The work described has not been formally published before, except in the form of an abstract or as part of a published lecture, review, thesis, or overlay journal. Please also carefully read Jurnal Pendidikan Informatika dan Sains Posting Your Article Policy at http://journal.ikippgriptk.ac.id/index.php/saintek/about/submissions#onlineSubmissions
- That it is not under consideration for publication elsewhere,
- That its publication has been approved by all the author(s) and by the responsible authorities – tacitly or explicitly – of the institutes where the work has been carried out.
- They secure the right to reproduce any material that has already been published or copyrighted elsewhere.
- They agree to the following license and copyright agreement.
Copyright
Authors who publish with Jurnal Pendidikan Informatika dan Sains agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (CC BY-SA 4.0) that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
Download: 18
