Transformer-Based Support for Content-Validity Pre-Screening in Educational Materials
DOI:
https://doi.org/10.62411/jcta.16829Keywords:
Aiken’s V, Content Validity, Educational Materials, Educational NLP, Essentiality Classification, Expert Annotation, Group-Aware Evaluation, Transformer ModelsAbstract
Content validity assessment is essential for determining whether educational materials adequately represent intended learning outcomes. However, conventional assessment procedures require substantial expert time and may produce inconsistent decisions across large item collections. This study develops a transformer-based framework to support content-validity pre-screening through two complementary tasks: predicting expert-derived Aiken’s V coefficients and classifying instructional-item essentiality. The final dataset comprised 652 Indonesian-language educational text items independently evaluated by four subject-matter experts. To reduce information leakage, identical and normalized-equivalent texts were grouped before applying a group-aware 70:15:15 training–validation–test split. Classical TF-IDF-based baselines were compared with IndoBERT, multilingual BERT, XLM-RoBERTa, and multilingual DeBERTa-v3. For Aiken’s V regression, multilingual BERT achieved the lowest MAE of 0.0501, the lowest RMSE of 0.0625, and the highest R² of 0.5239, whereas multilingual DeBERTa-v3 achieved the highest Spearman correlation of 0.7532. For essentiality classification, XLM-RoBERTa achieved the highest accuracy of 0.8557 and Macro-F1 of 0.8161, whereas multilingual BERT achieved the highest balanced accuracy of 0.8135 and ROC-AUC of 0.9111. Error analysis showed that the models captured textual patterns associated with expert-derived outcomes but remained limited when judgments depended on broader curricular context, competency hierarchies, prerequisite relationships, or relationships among instructional items. The findings support the use of transformer models as human-in-the-loop decision-support tools for prioritizing uncertain or potentially problematic educational items. However, the framework should be interpreted as a pre-screening mechanism rather than a replacement for expert judgment, and external validation across institutions and disciplines remains necessary.References
W. G. Spady, Outcome-Based Education: Critical Issues and Answers. Arlington, VA: American Association of School Administrators, 1994. [Online]. Available: https://eric.ed.gov/?id=ED380910
R. M. Harden, “Outcome-Based Education: the future is today,” Med. Teach., vol. 29, no. 7, pp. 625–629, Jan. 2007, doi: 10.1080/01421590701729930.
J. Biggs and C. Tang, Teaching for Quality Learning at University: What the Student Does, 4th ed. Maidenhead: McGraw-Hill/Society for Research into Higher Education/Open University Press, 2011.
M. S. B. Yusoff, “ABC of Content Validation and Content Validity Index Calculation,” Educ. Med. J., vol. 11, no. 2, pp. 49–54, Jun. 2019, doi: 10.21315/eimj2019.11.2.6.
N. Zaki, S. Turaev, K. Shuaib, A. Krishnan, and E. Mohamed, “Automating the mapping of course learning outcomes to program learning outcomes using natural language processing for accurate educational program evaluation,” Educ. Inf. Technol., vol. 28, no. 12, pp. 16723–16742, Dec. 2023, doi: 10.1007/s10639-023-11877-4.
E. Tocto-Cano, S. Paz Collado, and J. L. López-Gonzales, “A Holistic Maturity Model for Quality Assessment and Innovation in Peruvian Universities,” Educ. Sci., vol. 15, no. 2, p. 142, Jan. 2025, doi: 10.3390/educsci15020142.
M. Praseptiawan, A. N. Che Pee, M. H. Zakaria, and A. Noertjahyana, “Advancing the Measurement of MOOCs Software Quality: Validation of Assessment Tools Using the I-CVI Expert Framework,” Int. J. Eng. Sci. Inf. Technol., vol. 5, no. 3, pp. 138–145, May 2025, doi: 10.52088/ijesty.v5i3.911.
L. B. Mokkink, S. Herbelet, P. R. Tuinman, and C. B. Terwee, “Content validity: judging the relevance, comprehensiveness, and comprehensibility of an outcome measurement instrument – a COSMIN perspective,” J. Clin. Epidemiol., vol. 185, p. 111879, Sep. 2025, doi: 10.1016/j.jclinepi.2025.111879.
R. Doneva, S. Gaftandzhıeva, and G. Totkov, “Automated Quality Assurance of Educational Testing,” Turkish Online J. Distance Educ., vol. 19, no. 3, pp. 71–92, Jul. 2018, doi: 10.17718/tojde.444961.
A. M. McCarthy, D. Maor, A. McConney, and C. Cavanaugh, “Digital transformation in education: Critical components for leaders of system change,” Soc. Sci. Humanit. Open, vol. 8, no. 1, p. 100479, 2023, doi: 10.1016/j.ssaho.2023.100479.
M. Kayyali, “The Evolution of Quality Assurance in Higher Education,” in Navigating Quality Assurance and Accreditation in Global Higher Education, 2024, pp. 1–26. doi: 10.4018/979-8-3693-6915-9.ch001.
J. Devlin, M.-W. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, Oct. 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.
A. Conneau et al., “Unsupervised Cross-lingual Representation Learning at Scale,” ArXiv. Apr. 08, 2020. [Online]. Available: http://arxiv.org/abs/1911.02116
S. Devaraju, “Natural Language Processing (NLP) in AI-Driven Recruitment Systems,” Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol., pp. 555–566, Jun. 2022, doi: 10.32628/CSEIT2285241.
G. Tucudean, M. Bucos, B. Dragulescu, and C. D. Caleanu, “Natural language processing with transformers: a review,” PeerJ Comput. Sci., vol. 10, p. e2222, Aug. 2024, doi: 10.7717/peerj-cs.2222.
L. Krstić, V. Aleksić, and M. Krstić, “Artificial Intelligence in Education: A Review,” in Proceedings TIE 2022, 2022, pp. 223–228. doi: 10.46793/TIE22.223K.
P. Premananthan and M. Fahim, “The Role of Machine Learning in Smart Education: Taxonomy, Challenges, and Use Cases,” EAI Endorsed Trans. Tour. Technol. Intell., vol. 1, no. 1, Sep. 2024, doi: 10.4108/eettti.6833.
R. Yang, J. Cao, Z. Wen, Y. Wu, and X. He, “Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Lan-guage Models with Combination of Regression and Ranking,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 1560–1569. doi: 10.18653/v1/2020.findings-emnlp.141.
A. Amalia, D. Gunawan, Y. Fithri, and I. Aulia, “Automated Bahasa Indonesia essay evaluation with latent semantic analysis,” J. Phys. Conf. Ser., vol. 1235, no. 1, p. 012100, Jun. 2019, doi: 10.1088/1742-6596/1235/1/012100.
D. M. Adayilo, I. O. Oyefolahan, J. N. Ndunagu, C. Otuya, E. Malcalm, and K. Twabu, “AI-Powered Tutoring Systems for Per-sonalised Learning Feedback in Developing Secondary Education Contexts,” J. Futur. Artif. Intell. Technol., vol. 2, no. 4, pp. 549–564, Dec. 2025, doi: 10.62411/faith.3048-3719-290.
K. S. Kalyan, A. Rajasekharan, and S. Sangeetha, “AMMU: A survey of transformer-based biomedical pretrained language models,” J. Biomed. Inform., vol. 126, p. 103982, Feb. 2022, doi: 10.1016/j.jbi.2021.103982.
N. Kania, Y. S. Kusumah, J. A. Dahlan, E. Nurlaelah, F. Gürbüz, and E. Bonyah, “Constructing and providing content validity evidence through the Aiken’s V index based on the experts’ judgments of the instrument to measure mathematical problem-solving skills,” REID (Research Eval. Educ., vol. 10, no. 1, pp. 64–79, Jun. 2024, doi: 10.21831/reid.v10i1.71032.
L. R. Aiken, “Three Coefficients for Analyzing the Reliability and Validity of Ratings,” Educ. Psychol. Meas., vol. 45, no. 1, pp. 131–142, Mar. 1985, doi: 10.1177/0013164485451012.
D. M. Rubio, M. Berg-Weger, S. S. Tebb, E. S. Lee, and S. Rauch, “Objectifying content validity: Conducting a content validity study in social work research,” Soc. Work Res., vol. 27, no. 2, pp. 94–104, Jun. 2003, doi: 10.1093/swr/27.2.94.
A. Muzli, R. N. E. Anggraini, and D. Purwitasari, “Cog-CoT: A Cognitive Chain-of-Thought Framework for Bloom’s Taxono-my-Aligned Educational Question Answering,” J. Comput. Theor. Appl., vol. 4, no. 1, pp. 330–353, Aug. 2026, doi: 10.62411/jcta.17084.
D. Marutho, Muljono, S. Rustad, and Purwanto, “Optimizing aspect-based sentiment analysis using sentence embedding trans-former, bayesian search clustering, and sparse attention mechanism,” J. Open Innov. Technol. Mark. Complex., vol. 10, no. 1, p. 100211, Mar. 2024, doi: 10.1016/j.joitmc.2024.100211.
D. Marutho, Muljono, S. Rustad, and Purwanto, “Optimizing Aspect Term Extraction and Sentiment Classification through At-tention Mechanism and Sparse Attention Techniques,” Int. J. Intell. Eng. Syst., vol. 17, no. 5, pp. 1004–1015, Oct. 2024, doi: 10.22266/ijies2024.1031.75.
A. Esteva et al., “A guide to deep learning in healthcare,” Nat. Med., vol. 25, no. 1, pp. 24–29, Jan. 2019, doi: 10.1038/s41591-018-0316-z.
I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, “LEGAL-BERT: The Muppets straight out of Law School,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 2898–2904. doi: 10.18653/v1/2020.findings-emnlp.261.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Safuan, Dhendra Marutho, Ahmad Ilham, Muhammad Munsarif, Wendy Sarasjati, Edy Winarno, Arnold Adimabua Ojugo, De Rosal Ignatius Moses Setiadi

This work is licensed under a Creative Commons Attribution 4.0 International License.














