Patient-Level Data Leakage and External Generalisation in Deep Learning for Lung and Colon Cancer Classification: A Reproducibility Audit
Abstract
Deep-learning studies of lung and colon cancer images can inadvertently place images from the same patient in both training and test sets, producing optimistic estimates of generalisation. We audited 28 public and institutional image datasets and reproduced representative classification pipelines under image-level and patient-level splitting. In datasets containing multiple images per patient, random image-level splits increased test accuracy by a median of 6.7 percentage points and AUROC by 0.05 compared with strict patient-level separation. The inflation was larger when near-duplicate fields or serial sections were present. External validation produced a further reduction in performance even after patient-level separation, highlighting scanner, staining, and population differences. Duplicate-image screening identified additional leakage in three datasets where patient identifiers were unavailable. Models with heavy augmentation were less affected by near-duplicates but still showed substantial domain shift externally. Reproducibility in medical image classification therefore depends on documenting the unit of independence, duplicate handling, and external test population. Patient-level splitting should be treated as a minimum requirement whenever multiple observations can originate from the same individual.
References
1. Azam MA, et al. Machine learning meets XAI: Grad-CAM visualization for enhanced lung and colon cancer detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546191/
2. Azam MA, Ansari I, Haque GMM, Jahid A. Leveraging health information systems and predictive analytics to improve patient outcomes: a data-driven approach. The American Journal of Medical Sciences and Pharmaceutical Research. 2026;8(3):45-70. doi:10.37547/tajmspr/Volume08Issue03-06.
3. He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016. p. 770-778. doi:10.1109/CVPR.2016.90.
4. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE International Conference on Computer Vision. 2017. p. 618-626. doi:10.1109/ICCV.2017.74.
5. Adebayo J, Gilmer J, Muelly M, Goodfellow I, Hardt M, Kim B. Sanity checks for saliency maps. Adv Neural Inf Process Syst. 2018;31. Available from: https://arxiv.org/abs/1810.03292
6. Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018;15(11):e1002683. doi:10.1371/journal.pmed.1002683.
7. Roberts M, Driggs D, Thorpe M, Gilbey J, Yeung M, Ursprung S, et al. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nat Mach Intell. 2021;3:199-217. doi:10.1038/s42256-021-00307-0.
8. Borkowski AA, Bui MM, Thomas LB, Wilson CP, DeLand LA, Mastorides SM. Lung and colon cancer histopathological image dataset (LC25000). arXiv [Preprint]. 2019. doi:10.48550/arXiv.1912.12142.
9. Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330. Available from: https://proceedings.mlr.press/v70/guo17a.html
10. Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi:10.1371/journal.pone.0118432.
11. Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. doi:10.1016/j.patter.2023.100804.
12. Gebru T, Morgenstern J, Vecchione B, Vaughan JW, Wallach H, Daumé H III, et al. Datasheets for datasets. Commun ACM. 2021;64(12):86-92. doi:10.1145/3458723.
13. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378.
14. Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58. doi:10.7326/M18-1376.
15. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.
16. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.
17. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.
18. Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863.
Published
Issue
Section
License
Authors retain copyright. Articles published under the Creative Commons Attribution 4.0 International licence (CC BY 4.0) may be shared and adapted for any purpose, including commercially, provided appropriate credit is given, a link to the licence is supplied, and changes are indicated. No additional legal or technological restrictions may be imposed. Third-party material is included only where its credit line permits. Licence: https://creativecommons.org/licenses/by/4.0/. Earlier publications remain subject to their stated licence and author agreements unless the rights holder authorizes a change.