Explainable Machine Learning for Predicting Indonesian Vocational School Accreditation: Geographic Validation, Probability Calibration, and Subgroup Auditing

Authors

  • Muhamad Riyan Maulana Universitas Negeri Yogyakarta
  • Putu Sudira Universitas Negeri Yogyakarta
  • Priyanto Universitas Negeri Yogyakarta
  • Yanuar Agung Fadlullah Universitas Negeri Yogyakarta
  • Dian Nurdiana Universitas Terbuka
  • Amir Hamzah bin Sofhi @ Subhi Institut Pendidikan Guru Kampus Pendidikan Teknik

DOI:

https://doi.org/10.62411/jcta.17721

Keywords:

CatBoost, Educational Data Mining, Explainable Artificial Intelligence, Machine Learning, Geographic Validation, Machine learning, Probability Calibration, Subgroup Auditing, Vocational School Accreditation

Abstract

Vocational school accreditation is a high-stakes institutional classification, yet the geographic transferability and probability reliability of models based on routinely collected administrative data remain underexplored. This study developed a reproducible, leakage-controlled machine-learning framework for classifying the recorded accreditation classes C, B, and A among Indonesian vocational schools. Its main contribution is an integrated evaluation design that combines province-disjoint validation, probability calibration, cluster-aware uncertainty, and subgroup auditing to test reliability beyond conventional random-validation performance. A cross-sectional dataset containing 149,376 program-level records was deterministically aggregated into 14,391 schools, including 14,134 with valid labels. Eight provinces comprising 1,536 schools were locked before feature selection, model comparison, tuning, and calibration; 12,598 schools from 31 provinces supported development. The selected CatBoost model used 35 structural and vocational features with grouped out-of-fold temperature scaling. On the locked holdout, the calibrated model achieved balanced accuracy of 0.586, Macro F1 of 0.533, quadratic weighted kappa of 0.499, Macro ROC-AUC of 0.757, and expected calibration error of 0.049. School size, teacher resources, vocational staffing, program diversity, and status were most in-fluential. Class B had the lowest recall, and performance varied across school groups and provinces. The framework supports calibrated, auditable preliminary screening, but not automated accreditation or replacement of professional assessors.

Author Biographies

Muhamad Riyan Maulana, Universitas Negeri Yogyakarta

Department of Technology and Vocational Education, Graduate School, Universitas Negeri Yogyakarta, 55281, Yogyakarta, Indonesia

Putu Sudira, Universitas Negeri Yogyakarta

Department of Technology and Vocational Education, Graduate School, Universitas Negeri Yogyakarta, 55281, Yogyakarta, Indonesia

Priyanto, Universitas Negeri Yogyakarta

Department of Technology and Vocational Education, Graduate School, Universitas Negeri Yogyakarta, 55281, Yogyakarta, Indonesia

Yanuar Agung Fadlullah, Universitas Negeri Yogyakarta

Department of Technology and Vocational Education, Graduate School, Universitas Negeri Yogyakarta, 55281, Yogyakarta, Indonesia

Dian Nurdiana, Universitas Terbuka

Data Science Study Program, Faculty of Science and Technology, Universitas Terbuka, 15437, Tangerang Selatan, Indonesia

Amir Hamzah bin Sofhi @ Subhi, Institut Pendidikan Guru Kampus Pendidikan Teknik

Practicum Unit, Institut Pendidikan Guru Kampus Pendidikan Teknik, 71760, Bandar Baru Enstek, Malaysia

References

T. Vanpham and L. Dangnguyen, “Management of Self-Assessment and Quality Accreditation Activities at Vocational Colleges in Vietnam: Policy, Practice and Some Solutions,” JETT, vol. 14, no. 3, pp. 326–334, 2023, doi: 10.47750/jett.2023.14.03.04.

B. Susetyo and A. Lie, “Quality indicators of vocational schools: Partial least squares-SEM for small sample,” MethodsX, vol. 15, p. 103620, 2025, doi: 10.1016/j.mex.2025.103620.

OECD, “Teachers and leaders in vocational education and training,” OECD Publishing, 2021. doi: 10.1787/59d4fbb1-en.

S. Wang, F. Wang, Z. Zhu, J. Wang, T. Tran, and Z. Du, “Artificial intelligence in education: A systematic literature review,” Expert Syst. Appl., vol. 252, p. 124167, Oct. 2024, doi: 10.1016/j.eswa.2024.124167.

R. S. Baker and A. Hawn, “Algorithmic bias in education,” Int. J. Artif. Intell. Educ., vol. 32, no. 4, pp. 1052–1092, 2022, doi: 10.1007/s40593-021-00285-9.

A. Mathrani, T. Susnjak, G. Ramaswami, and A. Barczak, “Perspectives on the challenges of generalizability, transparency and ethics in predictive learning analytics,” Comput. Educ. Open, vol. 2, p. 100060, 2021, doi: 10.1016/j.caeo.2021.100060.

D. D. Lee and S. J. Cho, “Predicting the outcomes of the Korean national accreditation system for higher education institutions: A method using disclosure data for outsiders,” Asia Pacific Educ. Rev., vol. 22, no. 4, pp. 715–728, 2021, doi: 10.1007/s12564-021-09710-z.

N. Sghir, A. Adadi, and M. Lahmer, “Recent advances in Predictive Learning Analytics: A decade systematic review (2012–2022),” Educ. Inf. Technol., vol. 28, no. 7, pp. 8299–8333, Jul. 2023, doi: 10.1007/s10639-022-11536-0.

P. W. Koh et al., “WILDS: A benchmark of in-the-wild distribution shifts,” Proceedings of the 38th International Conference on Machine Learning, vol. 139. PMLR, pp. 5637–5664, 2021.

T. Takada et al., “Internal-external cross-validation helped to evaluate the generalizability of prediction models in large clustered datasets,” J. Clin. Epidemiol., vol. 137, pp. 83–91, 2021, doi: 10.1016/j.jclinepi.2021.03.025.

Y. Ovadia et al., “Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift,” Advances in Neural Information Processing Systems 32, vol. 32. Curran Associates, pp. 13969–13980, 2019.

T. Silva Filho, H. Song, M. Perello-Nieto, R. Santos-Rodriguez, M. Kull, and P. Flach, “Classifier calibration: A survey on how to assess and improve predicted class probabilities,” Mach. Learn., vol. 112, no. 9, pp. 3211–3260, 2023, doi: 10.1007/s10994-023-06336-7.

U. Johansson, T. Löfström, and H. Boström, “Calibrating multi-class models,” Proceedings of the Tenth Symposium on Conformal and Probabilistic Prediction and Applications, vol. 152. PMLR, pp. 111–130, 2021.

R. Alfredo et al., “Human-centred learning analytics and AI in education: A systematic literature review,” Comput. Educ. Artif. Intell., vol. 6, p. 100215, 2024, doi: 10.1016/j.caeai.2024.100215.

S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, no. 9, p. 100804, 2023, doi: 10.1016/j.patter.2023.100804.

I. D. Raji et al., “Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing,” Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, pp. 33–44, 2020. doi: 10.1145/3351095.3372873.

R. Ashmore, R. Calinescu, and C. Paterson, “Assuring the machine learning lifecycle: Desiderata, methods, and challenges,” ACM Comput. Surv., vol. 54, no. 5, pp. 111, 1–39, 2021, doi: 10.1145/3453444.

G. S. Collins et al., “TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods,” BMJ, vol. 385, p. e078378, 2024, doi: 10.1136/bmj-2023-078378.

H. Crompton and D. Burke, “Artificial intelligence in higher education: the state of the field,” Int. J. Educ. Technol. High. Educ., vol. 20, no. 1, p. 22, Apr. 2023, doi: 10.1186/s41239-023-00392-8.

G. Ramaswami, T. Susnjak, A. Mathrani, and R. Umer, “Use of predictive analytics within learning analytics dashboards: A review of case studies,” Technol. Knowl. Learn., vol. 28, no. 3, pp. 959–980, 2023, doi: 10.1007/s10758-022-09613-x.

Noripansyah, A. Kadir, D. Kusumaningsih, and Haderiansyah, “Accreditation prediction of early childhood education institutions using machine learning techniques,” J. Tek. Inform., vol. 4, no. 3, pp. 647–655, 2023, doi: 10.52436/1.jutif.2023.4.3.999.

M. B. Musthafa, N. Ngatmari, C. Rahmad, R. A. Asmara, and F. Rahutomo, “Evaluation of university accreditation prediction system,” IOP Conf. Ser. Mater. Sci. Eng., vol. 732, no. 1, p. 12041, 2020, doi: 10.1088/1757-899X/732/1/012041.

A. S. K. Rashid, “The extent of the teacher academic development from the accreditation evaluation system perspective using machine learning,” J. Exp. Theor. Artif. Intell., vol. 35, no. 4, pp. 535–555, 2023, doi: 10.1080/0952813X.2021.1960635.

M. Jannah and P. P. Izati, “Evaluation of classification methods for predicting junior high school accreditation ranks in Indonesia,” Int. J. Eng. Comput. Sci. Appl., vol. 5, no. 1, pp. 43–54, 2026, doi: 10.30812/ijecsa.v5i1.6032.

B. I. Igoche, O. Matthew, P. Bednar, and A. Gegov, “Integrating Structural Causal Model Ontologies with LIME for Fair Machine Learning Explanations in Educational Admissions,” J. Comput. Theor. Appl., vol. 2, no. 1, pp. 65–85, Jun. 2024, doi: 10.62411/jcta.10501.

J. P. Ntayagabiri, Y. Bentaleb, J. Ndikumagenge, and H. El Makhtoum, “A Comparative Analysis of Supervised Machine Learning Algorithms for IoT Attack Detection and Classification,” J. Comput. Theor. Appl., vol. 2, no. 3, pp. 395–409, Feb. 2025, doi: 10.62411/jcta.11901.

H. Suresh and J. Guttag, “A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle,” in Equity and Access in Algorithms, Mechanisms, and Optimization, Oct. 2021, pp. 1–9. doi: 10.1145/3465416.3483305.

N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A Survey on Bias and Fairness in Machine Learning,” ACM Comput. Surv., vol. 54, no. 6, pp. 1–35, Jul. 2022, doi: 10.1145/3457607.

H. Xia and K. Chen, “Bridging algorithmic prediction and teacher judgment: An explainable AI framework with institutional fairness calibration for early warning systems,” Front. Educ., vol. 11, p. 1881571, 2026, doi: 10.3389/feduc.2026.1881571.

W. H. Kruskal and W. A. Wallis, “Use of ranks in one-criterion variance analysis,” J. Am. Stat. Assoc., vol. 47, no. 260, pp. 583–621, 1952, doi: 10.1080/01621459.1952.10483441.

Y. Benjamini and Y. Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,” J. R. Stat. Soc. Ser. B, vol. 57, no. 1, pp. 289–300, 1995, doi: 10.1111/j.2517-6161.1995.tb02031.x.

D. R. Roberts et al., “Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure,” Ecography (Cop.)., vol. 40, no. 8, pp. 913–929, 2017, doi: 10.1111/ecog.02881.

L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin, “CatBoost: unbiased boosting with categorical features,” in Advances in Neural Information Processing Systems 31 (NeurIPS 2018), Jan. 2018. [Online]. Available: http://arxiv.org/abs/1706.09516

M. Sokolova and G. Lapalme, “A systematic analysis of performance measures for classification tasks,” Inf. Process. Manag., vol. 45, no. 4, pp. 427–437, 2009, doi: 10.1016/j.ipm.2009.03.002.

C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in Proceedings of the 34th International Conference on Machine Learning, Aug. 2017, pp. 1321–1330. [Online]. Available: http://arxiv.org/abs/1706.04599

B. Efron and R. J. Tibshirani, “An introduction to the bootstrap New York,” NY Chapman Hall, vol. 473, 1993.

S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” in NIPS’17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Nov. 2017, pp. 4768–4777. [Online]. Available: https://dl.acm.org/doi/10.5555/3295222.3295230

G.-J. Hwang, H. Xie, B. W. Wah, and D. Gašević, “Vision, challenges, roles and research issues of artificial intelligence in education,” Comput. Educ. Artif. Intell., vol. 1, p. 100001, 2020, doi: 10.1016/j.caeai.2020.100001.

Downloads

Published

2026-09-12

How to Cite

Maulana, M. R., Sudira, P., Priyanto, P., Fadlullah, Y. A., Nurdiana, D., & bin Sofhi @ Subhi, A. H. (2026). Explainable Machine Learning for Predicting Indonesian Vocational School Accreditation: Geographic Validation, Probability Calibration, and Subgroup Auditing. Journal of Computing Theories and Applications, 4(2), 502–516. https://doi.org/10.62411/jcta.17721