Efficient Glaucoma Detection from Retinal Fundus Images Using a Lightweight DeiT-Small Vision Transformer and Test-Time Augmentation
Main Article Content
Abstract
Glaucoma is one of the major causes of permanent blindness in the world, and since the progressive nature of the disease leads to considerable damage of the optic nerve before diagnosis, there is a need to develop precise and effective automated screening tools. This paper suggests the implementation of a lightweight vision transformer network for glaucoma screening based on the DeiT-Small architecture. The images were downsized to 224 × 224 resolution and transformed into patches. Four transformer blocks were frozen while the other blocks were fine-tuned via layer-wise learning-rate decay (LLRD) along with data augmentation and Test-Time Augmentation (TTA). Testing was performed on a balanced dataset with 9,540 images: 8,000 for training, 770 for validation and 770 for testing. In the independent test dataset, the model achieved 91.65% accuracy, 91.55% F1-score, 91.56% sensitivity, 95.06% specificity, and 95.84% AUC-ROC. Test-time augmentation was applied to enhance the reliability of predictions without re-training the architecture. According to the Grad-CAM visualization of predictions, the model made decisions based on clinically important features of the retinal structure, such as optic disc and optic cup. These results show the possibility of the effectiveness of a transformer-based technique in terms of automated glaucoma diagnosis, whereas further external validation using the datasets that represent the glaucoma distribution in the population is required.
Article Details
Section

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
References
References
[1] R. R. A. Bourne et al., “Global estimates on the number of people blind or visually impaired by glaucoma: A meta-analysis from 2000 to 2020,” Eye, vol. 38, no. 11, pp. 2036–2046, 2024, doi: 10.1038/s41433-024-02995-5.
[2] Y. C. Tham, X. Li, T. Y. Wong, H. A. Quigley, T. Aung, and C. Y. Cheng, “Global prevalence of glaucoma and projections of glaucoma burden through 2040: A systematic review and meta-analysis,” Ophthalmology, vol. 121, no. 11, pp. 2081–2090, Nov. 2014, doi: 10.1016/j.ophtha.2014.05.013.
[3] H. A. Quigley and A. T. Broman, “The number of people with glaucoma worldwide in 2010 and 2020,” British Journal of Ophthalmology, vol. 90, no. 3, p. 262, Mar. 2006, doi: 10.1136/bjo.2005.081224.
[4] American Academy of Ophthalmology, Primary Open-Angle Glaucoma Preferred Practice Pattern Guidelines, San Francisco, CA, USA, 2020.
[5] European Glaucoma Society, Terminology and Guidelines for Glaucoma, 5th ed., British Journal of Ophthalmology, vol. 105, Suppl. 1, pp. 1–169, 2021.
[6] D. R. Anderson, “Normal-tension glaucoma (low-tension glaucoma),” Survey of Ophthalmology, vol. 48, no. 2, pp. S52–S57, 2003.
[7] R. S. Harwerth, E. L. Carter-Dawson, J. K. Shen, J. M. Smith, and R. L. Crawford, “Ganglion cell losses underlying visual field defects from experimental glaucoma,” Progress in Retinal and Eye Research, vol. 18, no. 4, pp. 443–464, 1999.
[8] Z. Li et al., “Development and validation of a deep learning system for detecting glaucomatous optic neuropathy using fundus photographs,” JAMA Ophthalmology, vol. 136, no. 12, pp. 1363–1370, 2018.
[9] M. Raghu et al., “Transfusion: Understanding transfer learning for medical imaging,” Nature Medicine, vol. 25, pp. 764–769, 2019.
[10] W. Brendel and M. Bethge, “Approximating CNNs with bag-of-local-features models reveals their strong shape bias,” arXiv preprint arXiv:1904.00760, 2019.
[11] A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” ICLR, 2021.
[12] H. Touvron et al., “Training data-efficient image transformers & distillation through attention,” ICML, 2021.
[13] S. Pathan, P. Kumar, R. M. Pai, and S. V. Bhandary, “An automated classification framework for glaucoma detection in fundus images using ensemble of dynamic selection methods,” Progress in Artificial Intelligence, vol. 12, no. 3, pp. 287–301, 2023, doi: 10.1007/s13748-023-00304-x.
[14] L. Pascal, O. J. Perdomo, X. Bost, B. Huet, S. Otálora, and M. A. Zuluaga, “Multi-task deep learning for glaucoma detection from color fundus images,” Scientific Reports, vol. 12, no. 1, Dec. 2022, doi: 10.1038/s41598-022-16262-8.
[15] Y. Li, Y. Han, Z. Li, Y. Zhong, and Z. Guo, “A transfer learning-based multimodal neural network combining metadata and multiple medical images for glaucoma type diagnosis,” Scientific Reports, vol. 13, no. 1, p. 12076, 2023, doi: 10.1038/s41598-022-27045-6.
[16] V. K. Velpula and L. D. Sharma, “Multi-stage glaucoma classification using pre-trained convolutional neural networks and voting-based classifier fusion,” Frontiers in Physiology, vol. 14, 2023, https://doi: 10.3389/fphys.2023.1175881 .
[17] S. Saha, J. Vignarajan, and S. Frost, “A fast and fully automated system for glaucoma detection using color fundus photographs,” Scientific Reports, vol. 13, no. 1, Dec. 2023, doi: 10.1038/s41598-023-44473-0.
[18] R. Fan et al., “Detecting glaucoma from fundus photographs using deep learning without convolutions: Transformer for improved generalization,” Ophthalmology Science, vol. 3, no. 1, p. 100233, 2023, doi: 10.1016/j.xops.2022.100233.
[19] S. Chakraborty, A. Roy, P. Pramanik, D. Valenkova, and R. Sarkar, “A dual attention-aided DenseNet-121 for classification of glaucoma from fundus images,” arXiv preprint arXiv:2406.15113, Jun. 2024.
[20] K. J. Tina Brivitha and S. Sophia, “Integrated Deep Learning Framework for Automated Glaucoma Detection, Optic Disc/ Cup Segmentation, and CDR Calculation,” in Proceedings of International Conference on Visual Analytics and Data Visualization, ICVADV 2025, Institute of Electrical and Electronics Engineers Inc., 2025, pp. 887–894. doi: 10.1109/ICVADV63329.2025.10961500.
[21] R. Kiefer, “Glaucoma dataset: EyePACS-AIROGS-light-V2,” Kaggle Dataset, 2024. Available: https://www.kaggle.com/datasets/deathtrooper/glaucoma-dataset-eyepacs-airogs-light-v2
[22] H. Touvron, M. Cord, and H. Jégou, “DeiT III: Revenge of the ViT,” in Computer Vision – ECCV 2022, S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner, Eds., Cham: Springer Nature Switzerland, 2022, pp. 516–533.
[23] Y. Li and Q. Zhou, “Estimation of Doppler velocity from incoherent scatter spectra using context-aware transformers,” Atmos Meas Tech, vol. 19, no. 11, pp. 3865–3874, Jun. 2026, doi: 10.5194/amt-19-3865-2026.
[24] I. Loshchilov and F. Hutter, “DECOUPLED WEIGHT DECAY REGULARIZATION.” [Online]. Available: https://github.com/loshchil/AdamW-and-SGDW