Automated Detection of Diabetic Retinopathy Using Vision Transformers: A Deep Learning Approach

Authors

  • Tauseef Ahmad IAG Software House, Multan, 66000, Pakistan

DOI:

https://doi.org/10.66108/mna.v5i01.113

Keywords:

Deep Learning, Vision Transformer, Diabetic Retinopathy Diagnosis, Fundus Imaging, Automated Screening, Retinal Image Classification

Abstract

Diabetic Retinopathy (DR) is a prevalently occurring impediment of diabetes mellitus. It is also considered to be one of the leading causes of visual impairment in the world. The importance of timely intervention and treatment is vital in finding the stages of DR and classifying them accordingly in the initial stages of detection. This study enhances a deep learning resolution on Vision Transformers (ViT) in image automated detection and grading of DR on retinal fundi images. The proposed model processes high-resolution retinal images by dividing them into patches before applying the self-attention mechanisms to identify more intricate features that indicate the extent of the DR severity by relying on the APTOS 2019 Blindness Detection dataset. The measures of the model performance are accuracy, precision, recall, and F1-score which demonstrate the efficiency of the model in classifying the different stages of DR. The results indicate the potential of ViT-based architectures in enhancing the procedure of screening offered by DR, which is a promising method in ophthalmic diagnosis.

Downloads

Download data is not yet available.

References

Gettinger, K., Lee, D., Tomita, Y., Negishi, K., & Kurihara, T. (2025). Diabetic Retinopathy, a Comprehensive Overview on Pathophysiology and Relevant Experimental Models. International Journal of Molecular Sciences, 26(20), 9882. https://doi.org/10.3390/ijms26209882

Fong, D. S., Aiello, L. P., Ferris, F. L., & Klein, R. (2004). Diabetic Retinopathy. Diabetes Care, 27(10), 2540–2553. https://doi.org/10.2337/diacare.27.10.2540

Wang, A. (2025). Vision Transformers (ViTs): A New Era in Computer Vision – A Review. Journal of Computing and Electronic Information Management, 18(3), 23–30. https://doi.org/10.54097/41t0vx90

Awasthi, V., Awasthi, N., Kumar, H., Singh, S., Singh, P. P., Dixit, P., & Agarwal, R. (2024). ViT-HHO: Optimized vision transformer for diabetic retinopathy detection using Harris Hawk optimization. MethodsX, 13, 103018. https://doi.org/10.1016/j.mex.2024.103018

Mhasawade, A., Rawal, G., Roje, P., Raut, R., & Devkar, A. (2023). Comparative study of SVM, KNN and Decision Tree for Diabetic Retinopathy Detection. 2023 International Conference on Computational Intelligence and Sustainable Engineering Solutions (CISES), 166–170. https://doi.org/10.1109/cises58720.2023.10183456

Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., Venugopalan, S., Widner, K., Madams, T., Cuadros, J., Kim, R., Raman, R., Nelson, P. C., Mega, J. L., & Webster, D. R. (2016). Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs. JAMA, 316(22), 2402. https://doi.org/10.1001/jama.2016.17216

Pratt, H., Coenen, F., Broadbent, D. M., Harding, S. P., & Zheng, Y. (2016). Convolutional Neural Networks for Diabetic Retinopathy. Procedia Computer Science, 90, 200–205. https://doi.org/10.1016/j.procs.2016.07.014

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., N.Gomez, A., Kaiser, L., & Polosukhin, I. (2025). Attention Is All You Need. https://doi.org/10.65215/r5bs2d54

Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. https://doi.org/10.48550/arXiv.2010.11929

Chen, J., Mei, J., Li, X., Lu, Y., Yu, Q., Wei, Q., Luo, X., Xie, Y., Adeli, E., Wang, Y., Lungren, M. P., Zhang, S., Xing, L., Lu, L., Yuille, A., & Zhou, Y. (2024). TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers. Medical Image Analysis, 97, 103280. https://doi.org/10.1016/j.media.2024.103280

Chen, J., Mei, J., Li, X., Lu, Y., Yu, Q., Wei, Q., Luo, X., Xie, Y., Adeli, E., Wang, Y., Lungren, M. P., Zhang, S., Xing, L., Lu, L., Yuille, A., & Zhou, Y. (2024). TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers. Medical Image Analysis, 97, 103280. https://doi.org/10.1016/j.media.2024.103280

Mallampalli, P., & R, M. (2025). Diabetic Retinopathy Classification Using Swin Transformer: A Vision Transformer-Based Approach. 2025 3rd International Conference on Inventive Computing and Informatics (ICICI), 1181–1188. https://doi.org/10.1109/icici65870.2025.11069800

Wu, J., Hu, R., Xiao, Z., Chen, J., & Liu, J. (2021). Vision Transformer‐based recognition of diabetic retinopathy grade. Medical Physics, 48(12), 7850–7863. Portico. https://doi.org/10.1002/mp.15312

Yao, Z., Yuan, Y., Shi, Z., Mao, W., Zhu, G., Zhang, G., & Wang, Z. (2022). FunSwin: A deep learning method to analysis diabetic retinopathy grade and macular edema risk based on fundus images. Frontiers in Physiology, 13. https://doi.org/10.3389/fphys.2022.961386

Liu, Y., Yao, D., Ma, Y., Wang, H., Wang, J., Bai, X., Zeng, G., & Liu, Y. (2025). STMF-DRNet: A multi-branch fine-grained classification model for diabetic retinopathy using Swin-TransformerV2. Biomedical Signal Processing and Control, 103, 107352. https://doi.org/10.1016/j.bspc.2024.107352

Karthik, Maggie, & Dane, S. (2019). APTOS 2019 Blindness Detection [Data set]. Kaggle. https://www.kaggle.com/competitions/aptos2019-blindness-detection

Halder, A., Gharami, S., Sadhu, P., Singh, P. K., Woźniak, M., & Ijaz, M. F. (2024). Implementing vision transformer for classifying 2D biomedical images. Scientific Reports, 14(1). https://doi.org/10.1038/s41598-024-63094-9

Mohaideen Abdul Kadhar, K., & Anand, G. (2024). Python Libraries for Image Processing. Industrial Vision Systems with Raspberry Pi, 33–72. https://doi.org/10.1007/979-8-8688-0097-9_3

Dosovitskiy, A., & Brox, T. (2016). Inverting Visual Representations with Convolutional Networks. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4829–4837. https://doi.org/10.1109/cvpr.2016.522

Mishra, P. (2019). CNN and RNN Using PyTorch. PyTorch Recipes, 49–109. https://doi.org/10.1007/978-1-4842-4258-2_3

Additional Files

Published

2026-03-16

How to Cite

Tauseef Ahmad. (2026). Automated Detection of Diabetic Retinopathy Using Vision Transformers: A Deep Learning Approach. Machines and Algorithms, 5(01), 33–43. https://doi.org/10.66108/mna.v5i01.113

Issue

Section

Articles

Categories