Cross-Domain Faithfulness Evaluation of SHAP and Attention-Based Explanations in Transformer NLP Models
DOI:
https://doi.org/10.62411/jcta.16258Keywords:
Attention-based explanations, Cross-domain evaluation, Explainable artificial intelligence (XAI), Explanation faithfulness, Hate speech detection, Natural language processing, SHAP, Transformer modelsAbstract
Transformer-based models such as BERT, RoBERTa, DistilBERT, and DeBERTa have achieved remarkable performance across a wide range of natural language processing (NLP) tasks. However, their decision-making processes remain difficult to interpret, particularly in high-risk applications such as hate speech detection, where unreliable explanations may undermine model transparency, trust, and accountability. This study investigates whether explainability methods remain faithful and stable under domain shift in transformer-based text classification. Four transformer architectures were fine-tuned and evaluated on two linguistically distinct datasets: IMDb Movie Reviews and Hate Speech Offensive. Model performance and explanation quality were assessed using classification accuracy, macro F1-score, top-k token-removal faithfulness analysis, and cross-domain Spearman rank correlation. Experimental results show that DeBERTa achieved the highest classification performance, reaching accuracies of 95.6% on IMDb and 91.3% on Hate Speech. Across all evaluated models and datasets, SHAP consistently produced higher faithfulness scores than attention-based explanations. Cross-domain analysis further revealed reduced agreement between SHAP and attention-based explanations under domain shift, indicating lower explanation consistency across linguistically distinct domains. Qualitative error analysis further showed that implicit sentiment, sarcasm, and domain-specific slang remain major sources of prediction errors. Overall, the results demonstrate that superior predictive performance does not necessarily correspond to higher explanation faithfulness or stronger cross-domain stability. These findings highlight the importance of jointly evaluating predictive performance, explanation faithfulness, and explanation robustness when developing trustworthy transformer-based NLP systems.References
J. Devlin, M.-W. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, Oct. 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.
R. Yang, J. Cao, Z. Wen, Y. Wu, and X. He, “Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Language Models with Combination of Regression and Ranking,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 1560–1569. doi: 10.18653/v1/2020.findings-emnlp.141.
V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” ArXiv. pp. 2–6, Oct. 02, 2019. [Online]. Available: http://arxiv.org/abs/1910.01108
J. C. Timoneda and S. V. Vera, “BERT, RoBERTa, or DeBERTa? Comparing Performance Across Transformers Models in Political Science Text,” J. Polit., vol. 87, no. 1, pp. 347–364, Jan. 2025, doi: 10.1086/730737.
T. Wolf et al., “Transformers: State-of-the-Art Natural Language Processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2020, pp. 38–45. doi: 10.18653/v1/2020.emnlp-demos.6.
M. Mozafari, R. Farahbakhsh, and N. Crespi, “Hate speech detection and racial bias mitigation in social media based on BERT model,” PLoS One, vol. 15, no. 8, p. e0237861, Aug. 2020, doi: 10.1371/journal.pone.0237861.
A. Madsen, S. Reddy, and S. Chandar, “Post-hoc Interpretability for Neural NLP: A Survey,” ACM Comput. Surv., vol. 55, no. 8, pp. 1–42, Aug. 2023, doi: 10.1145/3546577.
S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” in NIPS’17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Nov. 2017, pp. 4768–4777. [Online]. Available: https://dl.acm.org/doi/10.5555/3295222.3295230
A. Vaswani et al., “Attention Is All You Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017. [Online]. Available: http://arxiv.org/abs/1706.03762
S. Jain and B. C. Wallace, “Attention is not Explanation,” in Proceedings of the 2019 Conference of the North, 2019, pp. 3543–3556. doi: 10.18653/v1/N19-1357.
S. Wiegreffe and Y. Pinter, “Attention is not not Explanation,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 11–20. doi: 10.18653/v1/D19-1002.
M. Danilevsky, K. Qian, R. Aharonov, Y. Katsis, B. Kawas, and P. Sen, “A Survey of the State of Explainable AI for Natural Language Processing,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, 2020, pp. 447–459. doi: 10.18653/v1/2020.aacl-main.46.
T. Davidson, D. Warmsley, M. Macy, and I. Weber, “Automated Hate Speech Detection and the Problem of Offensive Language,” Proc. Int. AAAI Conf. Web Soc. Media, vol. 11, no. 1, pp. 512–515, May 2017, doi: 10.1609/icwsm.v11i1.14955.
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts, “Learning Word Vectors for Sentiment Analysis,” in Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Jun. 2011, pp. 142–150. [Online]. Available: https://aclanthology.org/P11-1015/
I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” ArXiv. Jan. 04, 2019. [Online]. Available: http://arxiv.org/abs/1711.05101
K. L. Tan, C. P. Lee, K. S. M. Anbananthen, and K. M. Lim, “RoBERTa-LSTM: A Hybrid Model for Sentiment Analysis With Transformer and Recurrent Neural Network,” IEEE Access, vol. 10, pp. 21517–21525, 2022, doi: 10.1109/ACCESS.2022.3152828.
N. A. Semary, W. Ahmed, K. Amin, P. Pławiak, and M. Hammad, “Improving sentiment classification using a RoBERTa-based hybrid model,” Front. Hum. Neurosci., vol. 17, Dec. 2023, doi: 10.3389/fnhum.2023.1292010.
C. Sun, L. Huang, and X. Qiu, “Utilizing BERT for Aspect-Based Sentiment Analysis via Constructing Auxiliary Sentence,” in Proceedings of the 2019 Conference of the North, 2019, pp. 380–385. doi: 10.18653/v1/N19-1035.
N. D. A. Saputra, M. Muljono, A. Karim, and D. R. I. M. Setiadi, “End-to-End Fine-Tuning of DeBERTa-Base for Stance Detection,” J. Futur. Artif. Intell. Technol., vol. 2, no. 4, pp. 698–715, Feb. 2026, doi: 10.62411/faith.3048-3719-168.
S. M. Lundberg et al., “From local explanations to global understanding with explainable AI for trees,” Nat. Mach. Intell., vol. 2, no. 1, pp. 56–67, Jan. 2020, doi: 10.1038/s42256-019-0138-9.
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning, “What Does BERT Look at? An Analysis of BERT’s Attention,” in Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 2019, pp. 276–286. doi: 10.18653/v1/W19-4828.
A. Jacovi and Y. Goldberg, “Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 4198–4205. doi: 10.18653/v1/2020.acl-main.386.
M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh, “Beyond Accuracy: Behavioral Testing of NLP Models with CheckList,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 4902–4912. doi: 10.18653/v1/2020.acl-main.442.
A. J. Böck, D. Slijepčević, and M. Zeppelzauer, “Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI,” in 2024 International Conference on Content-Based Multimedia Indexing (CBMI), Sep. 2024, pp. 1–8. doi: 10.1109/CBMI62980.2024.10859247.
P. Atanasova, J. G. Simonsen, C. Lioma, and I. Augenstein, “A Diagnostic Study of Explainability Techniques for Text Classification,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 3256–3274. doi: 10.18653/v1/2020.emnlp-main.263.
P. Hase and M. Bansal, “Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 5540–5552. doi: 10.18653/v1/2020.acl-main.491.
B. Mathew, P. Saha, S. M. Yimam, C. Biemann, P. Goyal, and A. Mukherjee, “HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection,” Proc. AAAI Conf. Artif. Intell., vol. 35, no. 17, pp. 14867–14875, May 2021, doi: 10.1609/aaai.v35i17.17745.
B. Vidgen, T. Thrush, Z. Waseem, and D. Kiela, “Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, pp. 1667–1682. doi: 10.18653/v1/2021.acl-long.132.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Dony Bahtera Firmawan, Brian Rizqi Paradisiaca Darnoto

This work is licensed under a Creative Commons Attribution 4.0 International License.














