Transformer-Based Support for Content-Validity Pre-Screening in Educational Materials

Authors

  • Safuan Universitas Muhammadiyah Semarang
  • Dhendra Marutho Universitas Muhammadiyah Semarang
  • Ahmad Ilham Universitas Muhammadiyah Semarang
  • Muhammad Munsarif Universitas Muhammadiyah Semarang
  • Wendy Sarasjati Universitas Muhammadiyah Semarang
  • Edy Winarno Universitas Muhammadiyah Semarang
  • Arnold Adimabua Ojugo Federal University of Petroleum Resources
  • De Rosal Ignatius Moses Setiadi Universitas Dian Nuswantoro

DOI:

https://doi.org/10.62411/jcta.16829

Keywords:

Aiken’s V, Content Validity, Educational Materials, Educational NLP, Essentiality Classification, Expert Annotation, Group-Aware Evaluation, Transformer Models

Abstract

Content validity assessment is essential for determining whether educational materials adequately represent intended learning outcomes. However, conventional assessment procedures require substantial expert time and may produce inconsistent decisions across large item collections. This study develops a transformer-based framework to support content-validity pre-screening through two complementary tasks: predicting expert-derived Aiken’s V coefficients and classifying instructional-item essentiality. The final dataset comprised 652 Indonesian-language educational text items independently evaluated by four subject-matter experts. To reduce information leakage, identical and normalized-equivalent texts were grouped before applying a group-aware 70:15:15 training–validation–test split. Classical TF-IDF-based baselines were compared with IndoBERT, multilingual BERT, XLM-RoBERTa, and multilingual DeBERTa-v3. For Aiken’s V regression, multilingual BERT achieved the lowest MAE of 0.0501, the lowest RMSE of 0.0625, and the highest R² of 0.5239, whereas multilingual DeBERTa-v3 achieved the highest Spearman correlation of 0.7532. For essentiality classification, XLM-RoBERTa achieved the highest accuracy of 0.8557 and Macro-F1 of 0.8161, whereas multilingual BERT achieved the highest balanced accuracy of 0.8135 and ROC-AUC of 0.9111. Error analysis showed that the models captured textual patterns associated with expert-derived outcomes but remained limited when judgments depended on broader curricular context, competency hierarchies, prerequisite relationships, or relationships among instructional items. The findings support the use of transformer models as human-in-the-loop decision-support tools for prioritizing uncertain or potentially problematic educational items. However, the framework should be interpreted as a pre-screening mechanism rather than a replacement for expert judgment, and external validation across institutions and disciplines remains necessary.

Author Biographies

Safuan, Universitas Muhammadiyah Semarang

Department of Artificial Intelligence, Universitas Muhammadiyah Semarang, Semarang 50273, Indonesia

Dhendra Marutho, Universitas Muhammadiyah Semarang

Master’s Program in Informatics, Universitas Muhammadiyah Semarang, Semarang 50273, Indonesia

Ahmad Ilham, Universitas Muhammadiyah Semarang

Master’s Program in Informatics, Universitas Muhammadiyah Semarang, Semarang 50273, Indonesia

Muhammad Munsarif, Universitas Muhammadiyah Semarang

Department of Informatics, Universitas Muhammadiyah Semarang, Semarang 50273, Indonesia

Wendy Sarasjati, Universitas Muhammadiyah Semarang

Master’s Program in Informatics, Universitas Muhammadiyah Semarang, Semarang 50273, Indonesia

Edy Winarno, Universitas Muhammadiyah Semarang

Master’s Program in Informatics, Universitas Muhammadiyah Semarang, Semarang 50273, Indonesia

Arnold Adimabua Ojugo, Federal University of Petroleum Resources

Department of Computer Science, Federal University of Petroleum Resources, Effurun330102, Nigeria

De Rosal Ignatius Moses Setiadi, Universitas Dian Nuswantoro

Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang 50131, Indonesia

References

W. G. Spady, Outcome-Based Education: Critical Issues and Answers. Arlington, VA: American Association of School Administrators, 1994. [Online]. Available: https://eric.ed.gov/?id=ED380910

R. M. Harden, “Outcome-Based Education: the future is today,” Med. Teach., vol. 29, no. 7, pp. 625–629, Jan. 2007, doi: 10.1080/01421590701729930.

J. Biggs and C. Tang, Teaching for Quality Learning at University: What the Student Does, 4th ed. Maidenhead: McGraw-Hill/Society for Research into Higher Education/Open University Press, 2011.

M. S. B. Yusoff, “ABC of Content Validation and Content Validity Index Calculation,” Educ. Med. J., vol. 11, no. 2, pp. 49–54, Jun. 2019, doi: 10.21315/eimj2019.11.2.6.

N. Zaki, S. Turaev, K. Shuaib, A. Krishnan, and E. Mohamed, “Automating the mapping of course learning outcomes to program learning outcomes using natural language processing for accurate educational program evaluation,” Educ. Inf. Technol., vol. 28, no. 12, pp. 16723–16742, Dec. 2023, doi: 10.1007/s10639-023-11877-4.

E. Tocto-Cano, S. Paz Collado, and J. L. López-Gonzales, “A Holistic Maturity Model for Quality Assessment and Innovation in Peruvian Universities,” Educ. Sci., vol. 15, no. 2, p. 142, Jan. 2025, doi: 10.3390/educsci15020142.

M. Praseptiawan, A. N. Che Pee, M. H. Zakaria, and A. Noertjahyana, “Advancing the Measurement of MOOCs Software Quality: Validation of Assessment Tools Using the I-CVI Expert Framework,” Int. J. Eng. Sci. Inf. Technol., vol. 5, no. 3, pp. 138–145, May 2025, doi: 10.52088/ijesty.v5i3.911.

L. B. Mokkink, S. Herbelet, P. R. Tuinman, and C. B. Terwee, “Content validity: judging the relevance, comprehensiveness, and comprehensibility of an outcome measurement instrument – a COSMIN perspective,” J. Clin. Epidemiol., vol. 185, p. 111879, Sep. 2025, doi: 10.1016/j.jclinepi.2025.111879.

R. Doneva, S. Gaftandzhıeva, and G. Totkov, “Automated Quality Assurance of Educational Testing,” Turkish Online J. Distance Educ., vol. 19, no. 3, pp. 71–92, Jul. 2018, doi: 10.17718/tojde.444961.

A. M. McCarthy, D. Maor, A. McConney, and C. Cavanaugh, “Digital transformation in education: Critical components for leaders of system change,” Soc. Sci. Humanit. Open, vol. 8, no. 1, p. 100479, 2023, doi: 10.1016/j.ssaho.2023.100479.

M. Kayyali, “The Evolution of Quality Assurance in Higher Education,” in Navigating Quality Assurance and Accreditation in Global Higher Education, 2024, pp. 1–26. doi: 10.4018/979-8-3693-6915-9.ch001.

J. Devlin, M.-W. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, Oct. 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.

A. Conneau et al., “Unsupervised Cross-lingual Representation Learning at Scale,” ArXiv. Apr. 08, 2020. [Online]. Available: http://arxiv.org/abs/1911.02116

S. Devaraju, “Natural Language Processing (NLP) in AI-Driven Recruitment Systems,” Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol., pp. 555–566, Jun. 2022, doi: 10.32628/CSEIT2285241.

G. Tucudean, M. Bucos, B. Dragulescu, and C. D. Caleanu, “Natural language processing with transformers: a review,” PeerJ Comput. Sci., vol. 10, p. e2222, Aug. 2024, doi: 10.7717/peerj-cs.2222.

L. Krstić, V. Aleksić, and M. Krstić, “Artificial Intelligence in Education: A Review,” in Proceedings TIE 2022, 2022, pp. 223–228. doi: 10.46793/TIE22.223K.

P. Premananthan and M. Fahim, “The Role of Machine Learning in Smart Education: Taxonomy, Challenges, and Use Cases,” EAI Endorsed Trans. Tour. Technol. Intell., vol. 1, no. 1, Sep. 2024, doi: 10.4108/eettti.6833.

R. Yang, J. Cao, Z. Wen, Y. Wu, and X. He, “Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Lan-guage Models with Combination of Regression and Ranking,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 1560–1569. doi: 10.18653/v1/2020.findings-emnlp.141.

A. Amalia, D. Gunawan, Y. Fithri, and I. Aulia, “Automated Bahasa Indonesia essay evaluation with latent semantic analysis,” J. Phys. Conf. Ser., vol. 1235, no. 1, p. 012100, Jun. 2019, doi: 10.1088/1742-6596/1235/1/012100.

D. M. Adayilo, I. O. Oyefolahan, J. N. Ndunagu, C. Otuya, E. Malcalm, and K. Twabu, “AI-Powered Tutoring Systems for Per-sonalised Learning Feedback in Developing Secondary Education Contexts,” J. Futur. Artif. Intell. Technol., vol. 2, no. 4, pp. 549–564, Dec. 2025, doi: 10.62411/faith.3048-3719-290.

K. S. Kalyan, A. Rajasekharan, and S. Sangeetha, “AMMU: A survey of transformer-based biomedical pretrained language models,” J. Biomed. Inform., vol. 126, p. 103982, Feb. 2022, doi: 10.1016/j.jbi.2021.103982.

N. Kania, Y. S. Kusumah, J. A. Dahlan, E. Nurlaelah, F. Gürbüz, and E. Bonyah, “Constructing and providing content validity evidence through the Aiken’s V index based on the experts’ judgments of the instrument to measure mathematical problem-solving skills,” REID (Research Eval. Educ., vol. 10, no. 1, pp. 64–79, Jun. 2024, doi: 10.21831/reid.v10i1.71032.

L. R. Aiken, “Three Coefficients for Analyzing the Reliability and Validity of Ratings,” Educ. Psychol. Meas., vol. 45, no. 1, pp. 131–142, Mar. 1985, doi: 10.1177/0013164485451012.

D. M. Rubio, M. Berg-Weger, S. S. Tebb, E. S. Lee, and S. Rauch, “Objectifying content validity: Conducting a content validity study in social work research,” Soc. Work Res., vol. 27, no. 2, pp. 94–104, Jun. 2003, doi: 10.1093/swr/27.2.94.

A. Muzli, R. N. E. Anggraini, and D. Purwitasari, “Cog-CoT: A Cognitive Chain-of-Thought Framework for Bloom’s Taxono-my-Aligned Educational Question Answering,” J. Comput. Theor. Appl., vol. 4, no. 1, pp. 330–353, Aug. 2026, doi: 10.62411/jcta.17084.

D. Marutho, Muljono, S. Rustad, and Purwanto, “Optimizing aspect-based sentiment analysis using sentence embedding trans-former, bayesian search clustering, and sparse attention mechanism,” J. Open Innov. Technol. Mark. Complex., vol. 10, no. 1, p. 100211, Mar. 2024, doi: 10.1016/j.joitmc.2024.100211.

D. Marutho, Muljono, S. Rustad, and Purwanto, “Optimizing Aspect Term Extraction and Sentiment Classification through At-tention Mechanism and Sparse Attention Techniques,” Int. J. Intell. Eng. Syst., vol. 17, no. 5, pp. 1004–1015, Oct. 2024, doi: 10.22266/ijies2024.1031.75.

A. Esteva et al., “A guide to deep learning in healthcare,” Nat. Med., vol. 25, no. 1, pp. 24–29, Jan. 2019, doi: 10.1038/s41591-018-0316-z.

I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, “LEGAL-BERT: The Muppets straight out of Law School,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 2898–2904. doi: 10.18653/v1/2020.findings-emnlp.261.

Downloads

Published

2026-08-31

How to Cite

Safuan, S., Marutho, D., Ilham, A., Munsarif, M., Sarasjati, W., Winarno, E., Ojugo, A. A., & Setiadi, D. R. I. M. (2026). Transformer-Based Support for Content-Validity Pre-Screening in Educational Materials. Journal of Computing Theories and Applications, 4(1), 423–442. https://doi.org/10.62411/jcta.16829