Bridging Data Gaps: Multimodal Text-Image Fusion for Improved Medical Diagnosis in Low-Resource Settings

Authors

  • Muhammad Usama Abrar Higher School of Economics University, Nizhny Novgorod, 603155, Russia

DOI:

https://doi.org/10.66108/mna.v5i01.89

Keywords:

Medical Diagnosis, Text-image fusion, Multimodal learning, Healthcare AI, Data Scarcity

Abstract

Medical diagnosis often requires the integration of heterogeneous information sources such as clinical text and medical images. While artificial intelligence (AI)–based multimodal diagnostic systems have shown promising results in data-rich environments, their applicability in low-resource settings remains limited due to data scarcity, weak infrastructure, and limited expert availability. This paper surveys existing multimodal text–image fusion techniques and proposes data-efficient strategies tailored for low-resource healthcare environments. We discuss challenges related to data acquisition, annotation, alignment, and computational constraints, and analyze fusion architectures including early, late, and intermediate fusion. Ethical, practical, and deployment considerations are also examined. By focusing on lightweight, interpretable, and clinically grounded multimodal fusion approaches, this study highlights pathways for improving diagnostic accuracy and accessibility in underserved healthcare systems.

Downloads

Download data is not yet available.

References

Bray, O. H. (1997). Information integration for data fusion. Office of Scientific and Technical Information (OSTI). https://doi.org/10.2172/444047

Domingues, I., Muller, H., Ortiz, A., Dasarathy, B. V., Abreu, P. H., & Calhoun, V. D. (2020). Guest Editorial: Information Fusion for Medical Data: Early, Late, and Deep Fusion Methods for Multimodal Data. In IEEE Journal of Biomedical and Health Informatics (Vol. 24, Issue 1, pp. 14–16). Institute of Electrical and Electronics Engineers (IEEE). https://doi.org/10.1109/jbhi.2019.2958429

Huang, W.-Y., & Davis, J. J. (2011). Multimodality and nanoparticles in medical imaging. In Dalton Transactions (Vol. 40, Issue 23, p. 6087). Royal Society of Chemistry (RSC). https://doi.org/10.1039/c0dt01656j

Mu, S., Cui, M., & Huang, X. (2020). Multimodal Data Fusion in Learning Analytics: A Systematic Review. In Sensors (Vol. 20, Issue 23, p. 6856). MDPI AG. https://doi.org/10.3390/s20236856

Zhou, T., Thung, K., Zhu, X., & Shen, D. (2018). Effective feature learning and fusion of multimodality data using stage‐wise deep neural network for dementia diagnosis. In Human Brain Mapping (Vol. 40, Issue 3, pp. 1001–1016). Wiley. https://doi.org/10.1002/hbm.24428

Gusarova, N., Lobantsev, A., Vatian, A., Kapitonov, A., & Shalyto, A. (2020). Comparative assessment of text-image fusion models for medical diagnostics. In Information and Control Systems (Issue 5, pp. 70–79). State University of Aerospace Instrumentation (SUAI). https://doi.org/10.31799/1684-8853-2020-5-70-79

Akter, M., & Kudapa, S. P. (2024). A comparative analysis of artificial intelligence-integrated bi dashboards for real-time decision support in operations. International Journal of Scientific Interdisciplinary Research, 5(2), 158-191.

Muzammil, S. R., Maqsood, S., Haider, S., & Damaševičius, R. (2020). CSID: A Novel Multimodal Image Fusion Algorithm for Enhanced Clinical Diagnosis. In Diagnostics (Vol. 10, Issue 11, p. 904). MDPI AG. https://doi.org/10.3390/diagnostics10110904

Qi, G., Wang, J., Zhang, Q., Zeng, F., & Zhu, Z. (2017). An Integrated Dictionary-Learning Entropy-Based Medical Image Fusion Framework. In Future Internet (Vol. 9, Issue 4, p. 61). MDPI AG. https://doi.org/10.3390/fi9040061

Tang, L., Qian, J., Li, L., Hu, J., & Wu, X. (2017). Multimodal medical image fusion based on discrete Tchebichef moments and pulse coupled neural network. In International Journal of Imaging Systems and Technology (Vol. 27, Issue 1, pp. 57–65). Wiley. https://doi.org/10.1002/ima.22210

Johnson, A. E. W., Pollard, T. J., Berkowitz, S. J., Greenbaum, N. R., Lungren, M. P., Deng, C., Mark, R. G., & Horng, S. (2019). MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data, 6(1). https://doi.org/10.1038/s41597-019-0322-0

Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., & Ng, A. Y. (2011, June). Multimodal deep learning. In Icml (Vol. 11, pp. 689-696).

Boulahia, S. Y., Amamra, A., Madi, M. R., & Daikh, S. (2021). Early, intermediate and late fusion strategies for robust deep learning-based multimodal action recognition. Machine Vision and Applications, 32(6). https://doi.org/10.1007/s00138-021-01249-8

Guarrasi, V., Aksu, F., Caruso, C. M., Di Feola, F., Rofena, A., Ruffini, F., & Soda, P. (2024). A Systematic Review of Intermediate Fusion in Multimodal Deep Learning for Biomedical Applications. https://doi.org/10.2139/ssrn.4952813

Lu, J., Batra, D., Parikh, D., & Lee, S. (2019). Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32. https://proceedings.neurips.cc/paper_files/paper/2019/file/c74d97b01eae257e44aa9d5bade97baf-Paper.pdf

Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., & Zhang, L. (2018). Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6077–6086. https://doi.org/10.1109/cvpr.2018.00636

Cai, Y., & Rostami, M. (2024). Dynamic Transformer Architecture for Continual Learning of Multimodal Tasks. https://doi.org/10.2139/ssrn.4713352

Huang, B., Yang, F., Yin, M., Mo, X., & Zhong, C. (2020). A Review of Multimodal Medical Image Fusion Techniques. Computational and Mathematical Methods in Medicine, 2020, 1–16. https://doi.org/10.1155/2020/8279342

Kumar, S., Rani, S., Sharma, S., & Min, H. (2024). Multimodality Fusion Aspects of Medical Diagnosis: A Comprehensive Review. Bioengineering, 11(12), 1233. https://doi.org/10.3390/bioengineering11121233

Ullah Khan, S., Ahmad Khan, M., Azhar, M., Khan, F., Lee, Y., & Javed, M. (2023). Multimodal medical image fusion towards future research: A review. Journal of King Saud University - Computer and Information Sciences, 35(8), 101733. https://doi.org/10.1016/j.jksuci.2023.101733

Gu, X., Xia, Y., & Zhang, J. (2024). Multimodal medical image fusion based on interval gradients and convolutional neural networks. BMC Medical Imaging, 24(1). https://doi.org/10.1186/s12880-024-01418-x

Baltrusaitis, T., Ahuja, C., & Morency, L.-P. (2019). Multimodal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423–443. https://doi.org/10.1109/tpami.2018.2798607

Additional Files

Published

2026-03-16

How to Cite

Abrar, M. U. (2026). Bridging Data Gaps: Multimodal Text-Image Fusion for Improved Medical Diagnosis in Low-Resource Settings. Machines and Algorithms, 5(01), 44–56. https://doi.org/10.66108/mna.v5i01.89

Issue

Section

Reviews

Categories