AgriMisinfo-ID: A Multi-Platform Dataset and Ensemble Transformer-Based Detection System for Agricultural Misinformation in Indonesian Social Media
DOI:
https://doi.org/10.62411/jcta.17025Keywords:
Agricultural misinformation detection, Ensemble learning, IndoBERT, Indonesian NLP, Knowledge verification, Misinformation detection, Social media, XLM-RoBERTaAbstract
Agricultural misinformation on social media poses risks to Indonesian food security and farmer livelihoods. False claims about fertilizers, pesticides, crop varieties, and farming practices can spread rapidly across social media platforms such as Instagram, YouTube, and Facebook, as well as through publicly available datasets, potentially influencing agricultural decisions and outcomes. This paper introduces AgriMisinfo-ID, a bilingual, multi-platform dataset for agricultural misinformation detection, containing 5,288 labeled samples collected from social media and supplementary public datasets across Indonesian agricultural contexts. A hybrid detection system is proposed that combines an ensemble of fine-tuned Transformer models, IndoBERT and XLM-RoBERTa, with a Knowledge Verification module that cross-references agricultural claims against Wikidata and Wikipedia. Training uses Focal Loss to address class imbalance, together with GPT-4o-mini-based paraphrase augmentation for minority classes. Across three random seeds, the weighted ensemble achieves an F1-Macro of 0.5870 ± 0.0028 and an accuracy of 0.8027 ± 0.0039 on the test set, outperforming the individual models, a TF-IDF/SVM baseline, and an equal-weight ensemble in terms of F1-Macro. The Knowledge Verification module provides evidence-based verdicts that can support the inspection and auditability of model decisions. This work provides a reproducible benchmark for agricultural misinformation research in bilingual, low-resource settings.References
K. Shu, A. Sliva, S. Wang, J. Tang, and H. Liu, “Fake News Detection on Social Media,” ACM SIGKDD Explor. Newsl., vol. 19, no. 1, pp. 22–36, Sep. 2017, doi: 10.1145/3137597.3137600.
E. Aïmeur, S. Amri, and G. Brassard, “Fake news, disinformation and misinformation in social media: a review,” Soc. Netw. Anal. Min., vol. 13, no. 1, p. 30, Feb. 2023, doi: 10.1007/s13278-023-01028-5.
M. R. Islam, S. Liu, X. Wang, and G. Xu, “Deep learning for misinformation detection on online social networks: a survey and new perspectives,” Soc. Netw. Anal. Min., vol. 10, no. 1, p. 82, Dec. 2020, doi: 10.1007/s13278-020-00696-x.
A. Rahmawati, A. Alamsyah, and A. Romadhony, “Hoax News Detection Analysis using IndoBERT Deep Learning Methodology,” in 2022 10th International Conference on Information and Communication Technology (ICoICT), Aug. 2022, pp. 368–373. doi: 10.1109/ICoICT55009.2022.9914902.
B. P. Nayoga, R. Adipradana, R. Suryadi, and D. Suhartono, “Hoax Analyzer for Indonesian News Using Deep Learning Models,” Procedia Comput. Sci., vol. 179, pp. 704–712, 2021, doi: 10.1016/j.procs.2021.01.059.
A. Chowdhury, K. H. Kabir, A.-R. Abdulai, and M. F. Alam, “Systematic Review of Misinformation in Social and Online Media for the Development of an Analytical Framework for Agri-Food Sector,” Sustainability, vol. 15, no. 6, p. 4753, Mar. 2023, doi: 10.3390/su15064753.
F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Nov. 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66.
A. Conneau et al., “Unsupervised Cross-lingual Representation Learning at Scale,” in Proceedings of the 58th Annual Meeting of the As-sociation for Computational Linguistics, 2020, pp. 8440–8451. doi: 10.18653/v1/2020.acl-main.747.
V. U. Gongane, M. V. Munot, and A. D. Anuse, “A survey of explainable AI techniques for detection of fake news and hate speech on social media platforms,” J. Comput. Soc. Sci., vol. 7, no. 1, pp. 587–623, Apr. 2024, doi: 10.1007/s42001-024-00248-9.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal Loss for Dense Object Detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 2, pp. 318–327, Feb. 2020, doi: 10.1109/TPAMI.2018.2858826.
J. C. S. Reis, A. Correia, F. Murai, A. Veloso, and F. Benevenuto, “Supervised Learning for Fake News Detection,” IEEE Intell. Syst., vol. 34, no. 2, pp. 76–81, Mar. 2019, doi: 10.1109/MIS.2019.2899143.
J. Devlin, M.-W. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, Oct. 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.
R. K. Kaliyar, A. Goswami, and P. Narang, “FakeBERT: Fake news detection in social media with a BERT-based deep learning approach,” Multimed. Tools Appl., vol. 80, no. 8, pp. 11765–11788, Mar. 2021, doi: 10.1007/s11042-020-10183-2.
P. Dhiman, A. Kaur, D. Gupta, S. Juneja, A. Nauman, and G. Muhammad, “GBERT: A hybrid deep learning model based on GPT-BERT for fake news detection,” Heliyon, vol. 10, no. 16, p. e35865, Aug. 2024, doi: 10.1016/j.heliyon.2024.e35865.
S. A. Aljawarneh and S. A. Swedat, “Fake News Detection Using Enhanced BERT,” IEEE Trans. Comput. Soc. Syst., vol. 11, no. 4, pp. 4843–4850, Aug. 2024, doi: 10.1109/TCSS.2022.3223786.
E. Essa, K. Omar, and A. Alqahtani, “Fake news detection based on a hybrid BERT and LightGBM models,” Complex Intell. Syst., vol. 9, no. 6, pp. 6581–6592, Dec. 2023, doi: 10.1007/s40747-023-01098-0.
V.-I. Ilie, C.-O. Truica, E.-S. Apostol, and A. Paschke, “Context-Aware Misinformation Detection: A Benchmark of Deep Learning Architectures Using Word Embeddings,” IEEE Access, vol. 9, pp. 162122–162146, 2021, doi: 10.1109/ACCESS.2021.3132502.
S.-Y. Lin, Y.-C. Kung, and F.-Y. Leu, “Predictive intelligence in harmful news identification by BERT-based ensemble learning model with text sentiment analysis,” Inf. Process. Manag., vol. 59, no. 2, p. 102872, Mar. 2022, doi: 10.1016/j.ipm.2022.102872.
M. Al-alshaqi, D. B. Rawat, and C. Liu, “Ensemble Techniques for Robust Fake News Detection: Integrating Transformers, Natural Language Processing, and Machine Learning,” Sensors, vol. 24, no. 18, p. 6062, Sep. 2024, doi: 10.3390/s24186062.
M. E.Almandouh, M. F. Alrahmawy, M. Eisa, M. Elhoseny, and A. S. Tolba, “Ensemble based high performance deep learning models for fake news detection,” Sci. Rep., vol. 14, no. 1, p. 26591, Nov. 2024, doi: 10.1038/s41598-024-76286-0.
M. S. Khan and C. Hongsong, “Hybrid transformer deep neural architectures for enhanced misinformation detection on social media,” Expert Syst. Appl., vol. 300, p. 130470, Mar. 2026, doi: 10.1016/j.eswa.2025.130470.
C. Kumar, M. Bansal, M. A. Khan, V. Kaushik, M. Arquam, and A. Alabdultif, “Graph-augmented transformer ensemble framework for robust and scalable fake news detection in social media ecosystems,” Sci. Rep., vol. 16, no. 1, p. 2001, Dec. 2025, doi: 10.1038/s41598-025-31653-3.
M. F. Mridha, A. J. Keya, M. A. Hamid, M. M. Monowar, and M. S. Rahman, “A Comprehensive Review on Fake News Detection With Deep Learning,” IEEE Access, vol. 9, pp. 156151–156170, 2021, doi: 10.1109/ACCESS.2021.3129329.
D. Plikynas, I. Rizgelienė, and G. Korvel, “Systematic Review of Fake News, Propaganda, and Disinformation: Examining Authors, Content, and Social Impact Through Machine Learning,” IEEE Access, vol. 13, pp. 17583–17629, 2025, doi: 10.1109/ACCESS.2025.3530688.
N. Seddari, A. Derhab, M. Belaoued, W. Halboob, J. Al-Muhtadi, and A. Bouras, “A Hybrid Linguistic and Knowledge-Based Analysis Approach for Fake News Detection on Social Media,” IEEE Access, vol. 10, pp. 62097–62109, 2022, doi: 10.1109/ACCESS.2022.3181184.
B. Xie, X. Ma, J. Wu, J. Yang, and H. Fan, “Knowledge Graph Enhanced Heterogeneous Graph Neural Network for Fake News Detection,” IEEE Trans. Consum. Electron., vol. 70, no. 1, pp. 2826–2837, Feb. 2024, doi: 10.1109/TCE.2023.3324661.
G. Joshi et al., “Explainable Misinformation Detection Across Multiple Social Media Platforms,” IEEE Access, vol. 11, pp. 23634–23646, 2023, doi: 10.1109/ACCESS.2023.3251892.
S. Kuntur, A. Wróblewska, M. Paprzycki, and M. Ganzha, “Under the Influence: A Survey of Large Language Models in Fake News Detection,” IEEE Trans. Artif. Intell., vol. 6, no. 2, pp. 458–476, Feb. 2025, doi: 10.1109/TAI.2024.3471735.
Muhammad Ikram Kaer Sinapoy, Yuliant Sibaroni, and Sri Suryani Prasetyowati, “Comparison of LSTM and IndoBERT Method in Identifying Hoax on Twitter,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 7, no. 3, pp. 657–662, Jun. 2023, doi: 10.29207/resti.v7i3.4830.
T. C. Praha, W. Widodo, and M. Nugraheni, “Indonesian Fake News Classification Using Transfer Learning in CNN and LSTM,” JOIV Int. J. Informatics Vis., vol. 8, no. 3, p. 1213, Sep. 2024, doi: 10.62527/joiv.8.2.2126.
A. D. Safira and E. B. Setiawan, “Hoax Detection in Social Media using Bidirectional Long Short-Term Memory (Bi-LSTM) and 1 Dimensional-Convolutional Neural Network (1D-CNN) Methods,” in 2023 11th International Conference on Information and Communi-cation Technology (ICoICT), Aug. 2023, pp. 355–360. doi: 10.1109/ICoICT58202.2023.10262528.
L. A. Pekandi, R. G. Widjaja, A. Ananta, J. Harefa, and K. Jingga, “Evaluating IndoBERT for Indonesian Hoax News Detection: A Comparative Study with Ensemble and CNN-LSTM Models,” Procedia Comput. Sci., vol. 269, pp. 1625–1633, 2025, doi: 10.1016/j.procs.2025.09.105.
Y. Sagama and A. Alamsyah, “Multi-Label Classification of Indonesian Online Toxicity using BERT and RoBERTa,” in 2023 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology (IAICT), Jul. 2023, pp. 143–149. doi: 10.1109/IAICT59002.2023.10205892.
J. Sirusstara, N. Alexander, A. Alfarisy, S. Achmad, and R. Sutoyo, “Clickbait Headline Detection in Indonesian News Sites using Robustly Optimized BERT Pre-training Approach (RoBERTa),” in 2022 3rd International Conference on Artificial Intelligence and Data Sciences (AiDAS), Sep. 2022, pp. 1–6. doi: 10.1109/AiDAS56890.2022.9918678.
J. Laimeheriwa, I. A. Iswanto, W. Oktriani, and M. F. Hidayat, “A Survey on Indonesian Hoax Analyzer and Fake News Detection Using Deep Learning Techniques,” in 2024 6th International Conference on Cybernetics and Intelligent System (ICORIS), Nov. 2024, pp. 1–6. doi: 10.1109/ICORIS63540.2024.10903927.
D. B. Firmawan and B. R. P. Darnoto, “Cross-Domain Faithfulness Evaluation of SHAP and Attention-Based Explanations in Transformer NLP Models,” J. Comput. Theor. Appl., vol. 4, no. 1, pp. 146–163, Jul. 2026, doi: 10.62411/jcta.16258.
B. R. P. Darnoto and D. B. Firmawan, “Language-Similarity-Guided Transfer Fine-Tuning of Pre-trained Transformer Models for Sentiment Analysis Across 12 Indonesian Regional Languages,” J. Comput. Theor. Appl., vol. 3, no. 4, pp. 547–563, May 2026, doi: 10.62411/jcta.15975.
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal, “FEVER: a Large-scale Dataset for Fact Extraction and VERification,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), 2018, pp. 809–819. doi: 10.18653/v1/N18-1074.
J. Kim, S. Park, Y. Kwon, Y. Jo, J. Thorne, and E. Choi, “FactKG: Fact Verification via Reasoning on Knowledge Graphs,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 16190–16206. doi: 10.18653/v1/2023.acl-long.895.
Z. Chen, Y. Wang, B. Zhao, J. Cheng, X. Zhao, and Z. Duan, “Knowledge Graph Completion: A Review,” IEEE Access, vol. 8, pp. 192435–192456, 2020, doi: 10.1109/ACCESS.2020.3030076.
A. Golda et al., “Privacy and Security Concerns in Generative AI: A Comprehensive Survey,” IEEE Access, vol. 12, pp. 48126–48144, 2024, doi: 10.1109/ACCESS.2024.3381611.
G. G. Şahin, “To Augment or Not to Augment? A Comparative Study on Text Augmentation Techniques for Low-Resource NLP,” Comput. Linguist., vol. 48, no. 1, pp. 5–42, Apr. 2022, doi: 10.1162/coli_a_00425.
F. Muftie and M. Haris, “IndoBERT Based Data Augmentation for Indonesian Text Classification,” in 2023 International Conference on Information Technology Research and Innovation (ICITRI), Aug. 2023, pp. 128–132. doi: 10.1109/ICITRI59340.2023.10250061.
F. Gilardi, M. Alizadeh, and M. Kubli, “ChatGPT outperforms crowd workers for text-annotation tasks,” Proc. Natl. Acad. Sci., vol. 120, no. 30, Jul. 2023, doi: 10.1073/pnas.2305016120.
X. He et al., “AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track), 2024, pp. 165–190. doi: 10.18653/v1/2024.naacl-industry.15.
L. Rokach, “Ensemble-based classifiers,” Artif. Intell. Rev., vol. 33, no. 1–2, pp. 1–39, Feb. 2010, doi: 10.1007/s10462-009-9124-7.
D. Jiang and J. He, “Text Semantic Classification of Long Discourses Based on Neural Networks with Improved Focal Loss,” Comput. Intell. Neurosci., vol. 2021, no. 1, Jan. 2021, doi: 10.1155/2021/8845362.
H. R. Saeidnia, E. Hosseini, B. Lund, M. A. Tehrani, S. Zaker, and S. Molaei, “Artificial intelligence in the battle against disinfor-mation and misinformation: a systematic review of challenges and approaches,” Knowl. Inf. Syst., vol. 67, no. 4, pp. 3139–3158, Apr. 2025, doi: 10.1007/s10115-024-02337-7.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Brian Rizqi Paradisiaca Darnoto, Nelly Oktavia Adiwijaya, Dony Bahtera Firmawan, Fadhel Akhmad Hizham, Nandini Putri Hanifa Jannah, Talitha Puspitasari, Fabyan Yastika Permana

This work is licensed under a Creative Commons Attribution 4.0 International License.














