Patient-Level Data Leakage and External Generalisation in Deep Learning for Lung and Colon Cancer Classification: A Reproducibility Audit

Authors

  • Danyal Sheikh Department of Health Informatics and Clinical Data Science, Politecnico di Milano Author
  • Yuna Kim Department of Health Informatics and Clinical Data Science, University of Melbourne Author
  • Leon Bauer Department of Health Informatics and Clinical Data Science, United Arab Emirates University Author

Abstract

Deep-learning studies of lung and colon cancer images can inadvertently place images from the same patient in both training and test sets, producing optimistic estimates of generalisation. We audited 28 public and institutional image datasets and reproduced representative classification pipelines under image-level and patient-level splitting. In datasets containing multiple images per patient, random image-level splits increased test accuracy by a median of 6.7 percentage points and AUROC by 0.05 compared with strict patient-level separation. The inflation was larger when near-duplicate fields or serial sections were present. External validation produced a further reduction in performance even after patient-level separation, highlighting scanner, staining, and population differences. Duplicate-image screening identified additional leakage in three datasets where patient identifiers were unavailable. Models with heavy augmentation were less affected by near-duplicates but still showed substantial domain shift externally. Reproducibility in medical image classification therefore depends on documenting the unit of independence, duplicate handling, and external test population. Patient-level splitting should be treated as a minimum requirement whenever multiple observations can originate from the same individual.

References

1. Azam MA, et al. Machine learning meets XAI: Grad-CAM visualization for enhanced lung and colon cancer detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546191/

2. Azam MA, Ansari I, Haque GMM, Jahid A. Leveraging health information systems and predictive analytics to improve patient outcomes: a data-driven approach. The American Journal of Medical Sciences and Pharmaceutical Research. 2026;8(3):45-70. doi:10.37547/tajmspr/Volume08Issue03-06.

3. He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016. p. 770-778. doi:10.1109/CVPR.2016.90.

4. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE International Conference on Computer Vision. 2017. p. 618-626. doi:10.1109/ICCV.2017.74.

5. Adebayo J, Gilmer J, Muelly M, Goodfellow I, Hardt M, Kim B. Sanity checks for saliency maps. Adv Neural Inf Process Syst. 2018;31. Available from: https://arxiv.org/abs/1810.03292

6. Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018;15(11):e1002683. doi:10.1371/journal.pmed.1002683.

7. Roberts M, Driggs D, Thorpe M, Gilbey J, Yeung M, Ursprung S, et al. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nat Mach Intell. 2021;3:199-217. doi:10.1038/s42256-021-00307-0.

8. Borkowski AA, Bui MM, Thomas LB, Wilson CP, DeLand LA, Mastorides SM. Lung and colon cancer histopathological image dataset (LC25000). arXiv [Preprint]. 2019. doi:10.48550/arXiv.1912.12142.

9. Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330. Available from: https://proceedings.mlr.press/v70/guo17a.html

10. Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi:10.1371/journal.pone.0118432.

11. Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. doi:10.1016/j.patter.2023.100804.

12. Gebru T, Morgenstern J, Vecchione B, Vaughan JW, Wallach H, Daumé H III, et al. Datasheets for datasets. Commun ACM. 2021;64(12):86-92. doi:10.1145/3458723.

13. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378.

14. Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58. doi:10.7326/M18-1376.

15. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

16. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.

17. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.

18. Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863.

Published

2026-06-01