Explainable Machine Learning for Predicting Indonesian Vocational School Accreditation: Geographic Validation, Probability Calibration, and Subgroup Auditing
DOI:
https://doi.org/10.62411/jcta.17721Keywords:
CatBoost, Educational Data Mining, Explainable Artificial Intelligence, Machine Learning, Geographic Validation, Machine learning, Probability Calibration, Subgroup Auditing, Vocational School AccreditationAbstract
Vocational school accreditation is a high-stakes institutional classification, yet the geographic transferability and probability reliability of models based on routinely collected administrative data remain underexplored. This study developed a reproducible, leakage-controlled machine-learning framework for classifying the recorded accreditation classes C, B, and A among Indonesian vocational schools. Its main contribution is an integrated evaluation design that combines province-disjoint validation, probability calibration, cluster-aware uncertainty, and subgroup auditing to test reliability beyond conventional random-validation performance. A cross-sectional dataset containing 149,376 program-level records was deterministically aggregated into 14,391 schools, including 14,134 with valid labels. Eight provinces comprising 1,536 schools were locked before feature selection, model comparison, tuning, and calibration; 12,598 schools from 31 provinces supported development. The selected CatBoost model used 35 structural and vocational features with grouped out-of-fold temperature scaling. On the locked holdout, the calibrated model achieved balanced accuracy of 0.586, Macro F1 of 0.533, quadratic weighted kappa of 0.499, Macro ROC-AUC of 0.757, and expected calibration error of 0.049. School size, teacher resources, vocational staffing, program diversity, and status were most in-fluential. Class B had the lowest recall, and performance varied across school groups and provinces. The framework supports calibrated, auditable preliminary screening, but not automated accreditation or replacement of professional assessors.References
T. Vanpham and L. Dangnguyen, “Management of Self-Assessment and Quality Accreditation Activities at Vocational Colleges in Vietnam: Policy, Practice and Some Solutions,” JETT, vol. 14, no. 3, pp. 326–334, 2023, doi: 10.47750/jett.2023.14.03.04.
B. Susetyo and A. Lie, “Quality indicators of vocational schools: Partial least squares-SEM for small sample,” MethodsX, vol. 15, p. 103620, 2025, doi: 10.1016/j.mex.2025.103620.
OECD, “Teachers and leaders in vocational education and training,” OECD Publishing, 2021. doi: 10.1787/59d4fbb1-en.
S. Wang, F. Wang, Z. Zhu, J. Wang, T. Tran, and Z. Du, “Artificial intelligence in education: A systematic literature review,” Expert Syst. Appl., vol. 252, p. 124167, Oct. 2024, doi: 10.1016/j.eswa.2024.124167.
R. S. Baker and A. Hawn, “Algorithmic bias in education,” Int. J. Artif. Intell. Educ., vol. 32, no. 4, pp. 1052–1092, 2022, doi: 10.1007/s40593-021-00285-9.
A. Mathrani, T. Susnjak, G. Ramaswami, and A. Barczak, “Perspectives on the challenges of generalizability, transparency and ethics in predictive learning analytics,” Comput. Educ. Open, vol. 2, p. 100060, 2021, doi: 10.1016/j.caeo.2021.100060.
D. D. Lee and S. J. Cho, “Predicting the outcomes of the Korean national accreditation system for higher education institutions: A method using disclosure data for outsiders,” Asia Pacific Educ. Rev., vol. 22, no. 4, pp. 715–728, 2021, doi: 10.1007/s12564-021-09710-z.
N. Sghir, A. Adadi, and M. Lahmer, “Recent advances in Predictive Learning Analytics: A decade systematic review (2012–2022),” Educ. Inf. Technol., vol. 28, no. 7, pp. 8299–8333, Jul. 2023, doi: 10.1007/s10639-022-11536-0.
P. W. Koh et al., “WILDS: A benchmark of in-the-wild distribution shifts,” Proceedings of the 38th International Conference on Machine Learning, vol. 139. PMLR, pp. 5637–5664, 2021.
T. Takada et al., “Internal-external cross-validation helped to evaluate the generalizability of prediction models in large clustered datasets,” J. Clin. Epidemiol., vol. 137, pp. 83–91, 2021, doi: 10.1016/j.jclinepi.2021.03.025.
Y. Ovadia et al., “Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift,” Advances in Neural Information Processing Systems 32, vol. 32. Curran Associates, pp. 13969–13980, 2019.
T. Silva Filho, H. Song, M. Perello-Nieto, R. Santos-Rodriguez, M. Kull, and P. Flach, “Classifier calibration: A survey on how to assess and improve predicted class probabilities,” Mach. Learn., vol. 112, no. 9, pp. 3211–3260, 2023, doi: 10.1007/s10994-023-06336-7.
U. Johansson, T. Löfström, and H. Boström, “Calibrating multi-class models,” Proceedings of the Tenth Symposium on Conformal and Probabilistic Prediction and Applications, vol. 152. PMLR, pp. 111–130, 2021.
R. Alfredo et al., “Human-centred learning analytics and AI in education: A systematic literature review,” Comput. Educ. Artif. Intell., vol. 6, p. 100215, 2024, doi: 10.1016/j.caeai.2024.100215.
S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, no. 9, p. 100804, 2023, doi: 10.1016/j.patter.2023.100804.
I. D. Raji et al., “Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing,” Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, pp. 33–44, 2020. doi: 10.1145/3351095.3372873.
R. Ashmore, R. Calinescu, and C. Paterson, “Assuring the machine learning lifecycle: Desiderata, methods, and challenges,” ACM Comput. Surv., vol. 54, no. 5, pp. 111, 1–39, 2021, doi: 10.1145/3453444.
G. S. Collins et al., “TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods,” BMJ, vol. 385, p. e078378, 2024, doi: 10.1136/bmj-2023-078378.
H. Crompton and D. Burke, “Artificial intelligence in higher education: the state of the field,” Int. J. Educ. Technol. High. Educ., vol. 20, no. 1, p. 22, Apr. 2023, doi: 10.1186/s41239-023-00392-8.
G. Ramaswami, T. Susnjak, A. Mathrani, and R. Umer, “Use of predictive analytics within learning analytics dashboards: A review of case studies,” Technol. Knowl. Learn., vol. 28, no. 3, pp. 959–980, 2023, doi: 10.1007/s10758-022-09613-x.
Noripansyah, A. Kadir, D. Kusumaningsih, and Haderiansyah, “Accreditation prediction of early childhood education institutions using machine learning techniques,” J. Tek. Inform., vol. 4, no. 3, pp. 647–655, 2023, doi: 10.52436/1.jutif.2023.4.3.999.
M. B. Musthafa, N. Ngatmari, C. Rahmad, R. A. Asmara, and F. Rahutomo, “Evaluation of university accreditation prediction system,” IOP Conf. Ser. Mater. Sci. Eng., vol. 732, no. 1, p. 12041, 2020, doi: 10.1088/1757-899X/732/1/012041.
A. S. K. Rashid, “The extent of the teacher academic development from the accreditation evaluation system perspective using machine learning,” J. Exp. Theor. Artif. Intell., vol. 35, no. 4, pp. 535–555, 2023, doi: 10.1080/0952813X.2021.1960635.
M. Jannah and P. P. Izati, “Evaluation of classification methods for predicting junior high school accreditation ranks in Indonesia,” Int. J. Eng. Comput. Sci. Appl., vol. 5, no. 1, pp. 43–54, 2026, doi: 10.30812/ijecsa.v5i1.6032.
B. I. Igoche, O. Matthew, P. Bednar, and A. Gegov, “Integrating Structural Causal Model Ontologies with LIME for Fair Machine Learning Explanations in Educational Admissions,” J. Comput. Theor. Appl., vol. 2, no. 1, pp. 65–85, Jun. 2024, doi: 10.62411/jcta.10501.
J. P. Ntayagabiri, Y. Bentaleb, J. Ndikumagenge, and H. El Makhtoum, “A Comparative Analysis of Supervised Machine Learning Algorithms for IoT Attack Detection and Classification,” J. Comput. Theor. Appl., vol. 2, no. 3, pp. 395–409, Feb. 2025, doi: 10.62411/jcta.11901.
H. Suresh and J. Guttag, “A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle,” in Equity and Access in Algorithms, Mechanisms, and Optimization, Oct. 2021, pp. 1–9. doi: 10.1145/3465416.3483305.
N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A Survey on Bias and Fairness in Machine Learning,” ACM Comput. Surv., vol. 54, no. 6, pp. 1–35, Jul. 2022, doi: 10.1145/3457607.
H. Xia and K. Chen, “Bridging algorithmic prediction and teacher judgment: An explainable AI framework with institutional fairness calibration for early warning systems,” Front. Educ., vol. 11, p. 1881571, 2026, doi: 10.3389/feduc.2026.1881571.
W. H. Kruskal and W. A. Wallis, “Use of ranks in one-criterion variance analysis,” J. Am. Stat. Assoc., vol. 47, no. 260, pp. 583–621, 1952, doi: 10.1080/01621459.1952.10483441.
Y. Benjamini and Y. Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,” J. R. Stat. Soc. Ser. B, vol. 57, no. 1, pp. 289–300, 1995, doi: 10.1111/j.2517-6161.1995.tb02031.x.
D. R. Roberts et al., “Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure,” Ecography (Cop.)., vol. 40, no. 8, pp. 913–929, 2017, doi: 10.1111/ecog.02881.
L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin, “CatBoost: unbiased boosting with categorical features,” in Advances in Neural Information Processing Systems 31 (NeurIPS 2018), Jan. 2018. [Online]. Available: http://arxiv.org/abs/1706.09516
M. Sokolova and G. Lapalme, “A systematic analysis of performance measures for classification tasks,” Inf. Process. Manag., vol. 45, no. 4, pp. 427–437, 2009, doi: 10.1016/j.ipm.2009.03.002.
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in Proceedings of the 34th International Conference on Machine Learning, Aug. 2017, pp. 1321–1330. [Online]. Available: http://arxiv.org/abs/1706.04599
B. Efron and R. J. Tibshirani, “An introduction to the bootstrap New York,” NY Chapman Hall, vol. 473, 1993.
S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” in NIPS’17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Nov. 2017, pp. 4768–4777. [Online]. Available: https://dl.acm.org/doi/10.5555/3295222.3295230
G.-J. Hwang, H. Xie, B. W. Wah, and D. Gašević, “Vision, challenges, roles and research issues of artificial intelligence in education,” Comput. Educ. Artif. Intell., vol. 1, p. 100001, 2020, doi: 10.1016/j.caeai.2020.100001.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Muhamad Riyan Maulana, Putu Sudira, Priyanto, Yanuar Agung Fadlullah, Dian Nurdiana, Amir Hamzah bin Sofhi @ Subhi

This work is licensed under a Creative Commons Attribution 4.0 International License.














