Faithfulness and Stability of Grad-CAM Explanations in Lung and Colon Cancer Image Classification: A Comparative Evaluation

Authors

  • Alina Varga Department of Health Informatics and Clinical Data Science, University of Barcelona Author
  • Adeel Raza Department of Health Informatics and Clinical Data Science, Heidelberg University Author
  • Elias Fischer Department of Health Informatics and Clinical Data Science, University of São Paulo Author

Abstract

Grad-CAM is widely used to explain convolutional neural-network predictions in medical imaging, although visually convincing heatmaps may not be stable or faithful to model behaviour. We compared Grad-CAM explanations from four lung and colon histopathology classifiers across 6,400 images using model randomisation, input perturbation, repeated training, and region-overlap tests. Classification performance was high for all models, but explanation stability varied substantially between architectures. Heatmap overlap across retrained models ranged from 0.41 to 0.68 despite similar predictive accuracy. Explanations were more stable for large, spatially coherent lesions than for diffuse tissue patterns. In model-randomisation tests, some heatmaps retained visually similar structures after upper network layers were disrupted, indicating limited faithfulness in a subset of cases. Smoothing improved visual consistency but did not reliably improve causal sensitivity. The results show that classification accuracy cannot be used as a proxy for explanation quality. Grad-CAM outputs in cancer imaging should be accompanied by robustness and faithfulness checks before they are interpreted as evidence that a model relies on clinically meaningful regions.

References

1. Begum MH, Talukder SI, Ahmed MJ, Khan MFI, Azam MA. XAI-driven multimodal deep learning for early sepsis prediction in ICU. Journal of Science, Technology and Social Transformation. 2025;1(1):30-41. doi:10.64235/j62xmk30.

2. Azam MA, et al. Machine learning meets XAI: Grad-CAM visualization for enhanced lung and colon cancer detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546191/

3. Azam MA, Ansari I, Haque GMM, Jahid A. Leveraging health information systems and predictive analytics to improve patient outcomes: a data-driven approach. The American Journal of Medical Sciences and Pharmaceutical Research. 2026;8(3):45-70. doi:10.37547/tajmspr/Volume08Issue03-06.

4. He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016. p. 770-778. doi:10.1109/CVPR.2016.90.

5. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE International Conference on Computer Vision. 2017. p. 618-626. doi:10.1109/ICCV.2017.74.

6. Adebayo J, Gilmer J, Muelly M, Goodfellow I, Hardt M, Kim B. Sanity checks for saliency maps. Adv Neural Inf Process Syst. 2018;31. Available from: https://arxiv.org/abs/1810.03292

7. Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018;15(11):e1002683. doi:10.1371/journal.pmed.1002683.

8. Roberts M, Driggs D, Thorpe M, Gilbey J, Yeung M, Ursprung S, et al. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nat Mach Intell. 2021;3:199-217. doi:10.1038/s42256-021-00307-0.

9. Borkowski AA, Bui MM, Thomas LB, Wilson CP, DeLand LA, Mastorides SM. Lung and colon cancer histopathological image dataset (LC25000). arXiv [Preprint]. 2019. doi:10.48550/arXiv.1912.12142.

10. Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330. Available from: https://proceedings.mlr.press/v70/guo17a.html

11. Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi:10.1371/journal.pone.0118432.

12. Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. doi:10.1016/j.patter.2023.100804.

13. Gebru T, Morgenstern J, Vecchione B, Vaughan JW, Wallach H, Daumé H III, et al. Datasheets for datasets. Commun ACM. 2021;64(12):86-92. doi:10.1145/3458723.

14. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378.

15. Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58. doi:10.7326/M18-1376.

16. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

17. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.

18. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.

Published

2026-06-01