CAMMRF: A Conflict-Aware Missing-Modality Robust Fusion Framework for Multimodal Sentiment Analysis under Heterogeneous Observability
DOI:
https://doi.org/10.62411/jcta.16538Keywords:
Conflict-aware fusion, CMU-MOSI, Missing modality, Multimodal sentiment analysis, Reliability estimation, Robustness, Sentiment Analysis, Multimodal fusionAbstract
Multimodal sentiment analysis combines linguistic, acoustic, and visual evidence, yet deployment commonly violates the benchmark assumption that every modality is present and mutually consistent. Missing-modality systems generally reconstruct absent channels while conflict-aware systems generally assume full observation, even though absence and disagreement both change how far an observed modality should be trusted. We therefore treat the two jointly and hypothesise that conditioning modality reliability on both availability and cross-modal compatibility, and applying it across aligned and disagreement-preserving representations, yields more consistent performance across observability regimes than matched fusion controls without this coupling. We introduce CAMMRF, a framework that couples a mask- and peer-conditioned reliability estimator, an alignment/conflict subspace decomposition, and stochastic modality dropout with cross-view consistency. We evaluate one frozen protocol on the official CMU-MOSI train/validation/test partitions under eight observability regimes; five independent training seeds (42, 43, 44, 45, and 46) completed, each retaining predictions, checkpoints, and logs. In the full-modality regime CAMMRF is best on all four metrics simultaneously—MAE 1.050 ± 0.028, Pearson correlation 0.593 ± 0.009, weighted F1 75.01 ± 1.72%, and Acc-2 74.91 ± 1.80%—rather than on accuracy alone, and it ranks first by mean Acc-2 in 5 of eight regimes. The advantage persists when audio or vision is missing and under the controlled conflict stress (mean Acc-2 73.81%, a 1.10-point decrease, with matching small changes in MAE, weighted F1, and correlation); however, all four metrics degrade together when text is absent—correlation most steeply—so robustness is regime-dependent rather than universal. Full-budget ablations show that several component removals improve full-observation point estimates, supporting an accuracy–robustness trade-off rather than an independent-component claim. The evidence is limited to one corpus and controlled surrogate baselines and does not establish external generalisation. A single Colab notebook regenerates the protocol and exports a machine-readable result registry.References
A. Zadeh, R. Zellers, E. Pincus, and L.-P. Morency, “MOSI: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,” arXiv Prepr. arXiv1606.06259, 2016.
A. Bagher Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018, pp. 2236–2246. doi: 10.18653/v1/P18-1208.
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L.-P. Morency, “Tensor Fusion Network for Multimodal Sentiment Analysis,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, 2017, pp. 1103–1114. doi: 10.18653/v1/D17-1115.
Z. Liu, Y. Shen, V. B. Lakshminarasimhan, P. P. Liang, A. Bagher Zadeh, and L.-P. Morency, “Efficient Low-rank Multimodal Fusion With Modality-Specific Factors,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018, pp. 2247–2256. doi: 10.18653/v1/P18-1209.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal Transformer for Unaligned Mul-timodal Language Sequences,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 6558–6569. doi: 10.18653/v1/P19-1656.
W. Rahman et al., “Integrating Multimodal Information in Large Pretrained Transformers,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 2359–2369. doi: 10.18653/v1/2020.acl-main.214.
W. Yu, H. Xu, Z. Yuan, and J. Wu, “Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment Analysis,” Proc. AAAI Conf. Artif. Intell., vol. 35, no. 12, pp. 10790–10797, May 2021, doi: 10.1609/aaai.v35i12.17289.
W. Han, H. Chen, and S. Poria, “Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multi-modal Sentiment Analysis,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 9180–9192. doi: 10.18653/v1/2021.emnlp-main.723.
S. Yang, L. Cui, L. Wang, and T. Wang, “Cross-modal contrastive learning for multimodal sentiment recognition,” Appl. Intell., vol. 54, no. 5, pp. 4260–4276, Mar. 2024, doi: 10.1007/s10489-024-05355-8.
Z. Li, B. Xu, C. Zhu, and T. Zhao, “CLMLF:A Contrastive Learning and Multi-Layer Fusion Method for Multimodal Sentiment Detection,” in Findings of the Association for Computational Linguistics: NAACL 2022, 2022, pp. 2282–2294. doi: 10.18653/v1/2022.findings-naacl.175.
Y. Cai, X. Li, Y. Zhang, J. Li, F. Zhu, and L. Rao, “Multimodal sentiment analysis based on multi-layer feature fusion and multi-task learning,” Sci. Rep., vol. 15, no. 1, p. 2126, Jan. 2025, doi: 10.1038/s41598-025-85859-6.
J. Zhao, R. Li, and Q. Jin, “Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, pp. 2608–2618. doi: 10.18653/v1/2021.acl-long.203.
M. Ma, J. Ren, L. Zhao, D. Testuggine, and X. Peng, “Are Multimodal Transformers Robust to Missing Modality?,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 18156–18165. doi: 10.1109/CVPR52688.2022.01764.
J. Zeng, T. Liu, and J. Zhou, “Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing Modalities,” in Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Jul. 2022, pp. 1545–1554. doi: 10.1145/3477495.3532064.
Z. Lian, L. Chen, L. Sun, B. Liu, and J. Tao, “GCNet: Graph Completion Network for Incomplete Multimodal Learning in Con-versation,” IEEE Trans. Pattern Anal. Mach. Intell., pp. 1–14, 2023, doi: 10.1109/TPAMI.2023.3234553.
L. Sun, Z. Lian, B. Liu, and J. Tao, “Efficient Multimodal Transformer With Dual-Level Feature Restoration for Robust Multimodal Sentiment Analysis,” IEEE Trans. Affect. Comput., vol. 15, no. 1, pp. 309–325, Jan. 2024, doi: 10.1109/TAFFC.2023.3274829.
R. Lin and H. Hu, “MissModal: Increasing Robustness to Missing Modality in Multimodal Sentiment Analysis,” Trans. Assoc. Comput. Linguist., vol. 11, pp. 1686–1702, Dec. 2023, doi: 10.1162/tacl_a_00628.
Z. Guo, T. Jin, and Z. Zhao, “Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recog-nition,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1726–1736. doi: 10.18653/v1/2024.acl-long.94.
P. Xiang, C. Lin, K. Wu, and O. Bai, “MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition,” in 2024 14th International Conference on Pattern Recognition Systems (ICPRS), Jul. 2024, pp. 1–7. doi: 10.1109/ICPRS62101.2024.10677820.
M. K. Reza, A. Prater-Bennette, and M. S. Asif, “Robust Multimodal Learning With Missing Modalities via Parameter-Efficient Adaptation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 2, pp. 742–754, Feb. 2025, doi: 10.1109/TPAMI.2024.3476487.
D. Hazarika, R. Zimmermann, and S. Poria, “MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis,” in Proceedings of the 28th ACM International Conference on Multimedia, Oct. 2020, pp. 1122–1131. doi: 10.1145/3394171.3413678.
B. Liang, C. Lou, X. Li, L. Gui, M. Yang, and R. Xu, “Multi-modal sarcasm detection via cross-modal graph convolutional network,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022, pp. 1767–1777. doi: 10.18653/v1/2022.acl-long.124.
H. Liu, R. Wei, G. Tu, J. Lin, C. Liu, and D. Jiang, “Sarcasm driven by sentiment: A sentiment-aware hierarchical fusion network for multimodal sarcasm detection,” Inf. Fusion, vol. 108, p. 102353, Aug. 2024, doi: 10.1016/j.inffus.2024.102353.
S. Pramanick, A. Roy, and V. M. P. Johns, “Multimodal Learning using Optimal Transport for Sarcasm and Humor Detection,” in 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Jan. 2022, pp. 546–556. doi: 10.1109/WACV51458.2022.00062.
Y. Gao, H. Wu, and L. Zhang, “Multi-level Conflict-Aware Network for Multi-modal Sentiment Analysis.” Feb. 14, 2025. doi: 10.20944/preprints202502.1056.v1.
R. Grzeczkowicz et al., “Uncertainty-aware multimodal emotion recognition through dirichlet parameterization,” arXiv Prepr. arXiv2602.09121, 2026.
P. Gong, J. Liu, X. Zhang, X. Li, L. Wei, and H. He, “Towards robust sentiment analysis with multimodal interaction graph and hybrid contrastive learning,” Pattern Recognit., vol. 169, p. 111870, Jan. 2026, doi: 10.1016/j.patcog.2025.111870.
G. Hu and others, “Recent advances in multimodal affective computing: An NLP perspective,” arXiv Prepr. arXiv2409.07388, 2024.
B. Schuller and others, “Affective computing has changed: The foundation model disruption,” arXiv Prepr. arXiv2409.08907, 2024.
Y. Zhao, Y. Peng, L. Zhang, Q. Sun, Z. Zhang, and Y. Zhuang, “Multimodal foundation model-driven user interest modeling and behavior analysis on short video platforms,” arXiv Prepr. arXiv2509.04751, 2025.
Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” ArXiv. Jul. 26, 2019.
G. Degottex, J. Kane, T. Drugman, T. Raitio, and S. Scherer, “COVAREP - a collaborative voice analysis repository for speech technologies,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 960–964. doi: 10.1109/ICASSP.2014.6853739.
S. Stöckli, M. Schulte-Mecklenbeck, S. Borer, and A. C. Samson, “Facial expression analysis with AFFDEX and FACET: A valida-tion study,” Behav. Res. Methods, vol. 50, no. 4, pp. 1446–1460, Aug. 2018, doi: 10.3758/s13428-017-0996-1.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Kajal Chourasiya, Parul Saxena, Praphula Kumar Jain

This work is licensed under a Creative Commons Attribution 4.0 International License.














