Cog-CoT: A Cognitive Chain-of-Thought Framework for Bloom's Taxonomy-Aligned Educational Question Answering
DOI:
https://doi.org/10.62411/jcta.17084Keywords:
Bloom's Taxonomy, Chain-of-Thought Prompting, Cognitive Scaffolding, Educational Question Answering, Large Language Models, Retrieval-Augmented Generation, Prompt Engineering, Cognitive ReasoningAbstract
Large Language Models (LLMs) have shown strong potential in educational question answering, yet they often exhibit cognitive misalignment by generating responses that emphasize linguistic fluency rather than the cognitive depth required by different levels of Bloom's Taxonomy. This limitation arises from three structural gaps: Bloom classifiers are typically disconnected from the generation process, Chain-of-Thought (CoT) reasoning lacks explicit cognitive scaffolding, and Retrieval-Augmented Generation (RAG) verifies factual consistency only at the final output. To address these limitations, this paper proposes Cognitive Chain-of-Thought (Cog-CoT), a four-module framework that integrates Bloom's Taxonomy into the reasoning process of LLMs. The framework consists of (1) a LinearSVC-based Cognitive Classifier (weighted F1 = 0.9915) for Bloom-level prediction, (2) a Hierarchical Cognitive Decomposer that constructs a bottom-up sequence of Bloom-aligned sub-questions, (3) a Cog-CoT Reasoning module employing six level-specific prompt templates with step-wise cosine similarity verification against retrieved context ( = 0.65, MAX_RETRY = 3), and (4) a Bottom-Up Aggregator that synthesizes verified intermediate responses into a coherent final answer. Experiments using three open-weight LLMs (Gemma-3-4B-IT, LLaMA-3.1-8B-Instruct, and Qwen2.5-7B-Instruct) on an Indonesian Social Studies dataset and external English benchmarks (SQuAD 2.0 and ASQA) demonstrate that Cog-CoT consistently outperforms Zero-Shot prompting, achieving improvements of 4.16%–7.36% in GT Cosine Similarity while also producing superior performance across BLEU-4, ROUGE-L, METEOR, and BERTScore. These results demonstrate the effectiveness of integrating cognitive scaffolding with structured reasoning and retrieval verification for educational question answering.References
R. Prakash and R. Litoriya, “Pedagogical Transformation of Bloom Taxonomy’s LOTs into HOTs: An Investigation in Context with IT Education,” Wirel. Pers. Commun., vol. 122, no. 1, pp. 725–736, Jan. 2022, doi: 10.1007/s11277-021-08921-2.
L. Yan et al., “Practical and ethical challenges of large language models in education: A systematic scoping review,” Br. J. Educ. Technol., vol. 55, no. 1, pp. 90–112, Jan. 2024, doi: 10.1111/bjet.13370.
A. E. Evwiekpaefe, D. T. Chinyio, and L. K. Tohomdet, “The Llama–ARCS Adaptive Learning framework: AI–VR Integration System for Real-Time Motivational Feedback in Higher Education,” J. Comput. Theor. Appl., vol. 3, no. 3, pp. 260–273, Feb. 2026, doi: 10.62411/jcta.15031.
A. Herrmann-Werner et al., “Assessing ChatGPT’s Mastery of Bloom’s Taxonomy Using Psychosomatic Medicine Exam Questions: Mixed-Methods Study,” J. Med. Internet Res., vol. 26, p. e52113, Jan. 2024, doi: 10.2196/52113.
E. Kasneci et al., “ChatGPT for good? On opportunities and challenges of large language models for education,” Learn. Individ. Differ., vol. 103, p. 102274, Apr. 2023, doi: 10.1016/j.lindif.2023.102274.
A. Sangodiah, T. Jee San, Y. Tien Fui, L. Ean Heng, R. K. Ayyasamy, and N. A Jalil, “Identifying Optimal Baseline Variant of Unsupervised Term Weighting in Question Classification Based on Bloom Taxonomy,” MENDEL, vol. 28, no. 1, pp. 8–22, Jun. 2022, doi: 10.13164/mendel.2022.1.008.
M. Chindukuri and S. Sivanesan, “Transfer learning for Bloom’s taxonomy-based question classification,” Neural Comput. Appl., vol. 36, no. 31, pp. 19915–19937, Nov. 2024, doi: 10.1007/s00521-024-10241-y.
J. Miao, C. Thongprayoon, S. Suppadungsuk, P. Krisanapan, Y. Radhakrishnan, and W. Cheungpasitporn, “Chain of Thought Utilization in Large Language Models and Application in Nephrology,” Medicina (B. Aires)., vol. 60, no. 1, p. 148, Jan. 2024, doi: 10.3390/medicina60010148.
M. Lu, F. Gao, X. Tang, and L. Chen, “Analysis and prediction in SCR experiments using GPT-4 with an effective chain-of-thought prompting strategy,” iScience, vol. 27, no. 4, p. 109451, Apr. 2024, doi: 10.1016/j.isci.2024.109451.
Y. Liu, “Retrieval-Augmented Generation: Methods, Applications and Challenges,” Appl. Comput. Eng., vol. 142, no. 1, pp. 99–108, Apr. 2025, doi: 10.54254/2755-2721/2025.KL22312.
Z. Chen et al., “MedScaleRE-PF: a prompt-based framework with retrieval-augmented generation, chain-of-thought, and self-verification for scale-specific relation extraction in Chinese medical literature,” Inf. Process. Manag., vol. 62, no. 6, p. 104278, Nov. 2025, doi: 10.1016/j.ipm.2025.104278.
O. Almatrafi and A. Johri, “Leveraging generative AI for course learning outcome categorization using Bloom’s taxonomy,” Comput. Educ. Artif. Intell., vol. 8, p. 100404, Jun. 2025, doi: 10.1016/j.caeai.2025.100404.
Z. Kolagar, F. Zalkow, and A. Zarcone, “Investigating Methods for Mapping Learning Objectives to Bloom’s Revised Taxonomy in Course Descriptions for Higher Education,” in Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), 2025, pp. 415–445. doi: 10.18653/v1/2025.bea-1.32.
N. Duong-Trung, X. Wang, and M. Kravčík, “BloomLLM: Large Language Models Based Question Generation Combining Supervised Fine-Tuning and Bloom’s Taxonomy,” in Lecture Notes in Computer Science, 2024, pp. 93–98. doi: 10.1007/978-3-031-72312-4_11.
G.-G. Lee, E. Latif, X. Wu, N. Liu, and X. Zhai, “Applying large language models and chain-of-thought for automatic scoring,” Comput. Educ. Artif. Intell., vol. 6, p. 100213, Jun. 2024, doi: 10.1016/j.caeai.2024.100213.
X. Zhang et al., “TreeQA: Enhanced LLM-RAG with logic tree reasoning for reliable and interpretable multi-hop question answering,” Knowledge-Based Syst., vol. 330, p. 114526, Nov. 2025, doi: 10.1016/j.knosys.2025.114526.
A. Plaat, A. Wong, S. Verberne, J. Broekens, N. Van Stein, and T. Bäck, “Multi-Step Reasoning with Large Language Models, a Survey,” ACM Comput. Surv., vol. 58, no. 6, pp. 1–35, Apr. 2026, doi: 10.1145/3774896.
Z. Chu et al., “Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1173–1203. doi: 10.18653/v1/2024.acl-long.65.
Z. Li, Z. Wang, W. Wang, K. Hung, H. Xie, and F. L. Wang, “Retrieval-augmented generation for educational application: A systematic survey,” Comput. Educ. Artif. Intell., vol. 8, p. 100417, Jun. 2025, doi: 10.1016/j.caeai.2025.100417.
A. L. Manuel Tampubolon, R. N. E. Anggraini, and S. C. Hidayati, “Retrieval Augmented Generation with Synergizing Reasoning and Acting Prompt Engineering for Indonesian Open-Domain Question Answering,” in 2025 International Conference on Data Science and Its Applications (ICoDSA), Jul. 2025, pp. 333–338. doi: 10.1109/ICoDSA67155.2025.11157319.
S. Jeong, J. Baek, S. Cho, S. J. Hwang, and J. Park, “Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 7036–7050. doi: 10.18653/v1/2024.naacl-long.389.
S. Shaikh, S. M. Daudpotta, and A. S. Imran, “Bloom’s Learning Outcomes’ Automatic Classification Using LSTM and Pretrained Word Embeddings,” IEEE Access, vol. 9, pp. 117887–117909, 2021, doi: 10.1109/ACCESS.2021.3106443.
M. O. Gani et al., “Towards enhanced assessment question classification: a study using machine learning, deep learning, and generative AI,” Conn. Sci., vol. 37, no. 1, Dec. 2025, doi: 10.1080/09540091.2024.2445249.
P. Rajpurkar, R. Jia, and P. Liang, “Know What You Don’t Know: Unanswerable Questions for SQuAD,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2018, pp. 784–789. doi: 10.18653/v1/P18-2124.
I. Stelmakh, Y. Luan, B. Dhingra, and M.-W. Chang, “ASQA: Factoid Questions Meet Long-Form Answers,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 8273–8288. doi: 10.18653/v1/2022.emnlp-main.566.
T. M. Seinen, J. A. Kors, E. M. van Mulligen, and P. R. Rijnbeek, “Using Structured Codes and Free-Text Notes to Measure Information Complementarity in Electronic Health Records: Feasibility and Validation Study,” J. Med. Internet Res., vol. 27, p. e66910, Feb. 2025, doi: 10.2196/66910.
C. Prakash, M. Lind, and E. De La Cruz, “Hybrid Real-time Framework for Detecting Adaptive Prompt Injection Attacks in Large Language Models,” J. Comput. Theor. Appl., vol. 3, no. 3, pp. 286–301, Jan. 2026, doi: 10.62411/jcta.15254.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Alfarabi Muzli, Ratih Nur Esti Anggraini, Diana Purwitasari

This work is licensed under a Creative Commons Attribution 4.0 International License.














