Data Lineage and Temporal Validation in Hospital Readmission Prediction: A Reproducibility Study of Integrated Health Information Systems

Authors

  • Elias Fischer Department of Health Informatics and Clinical Data Science, University of São Paulo Author
  • Hassan Saleh Department of Health Informatics and Clinical Data Science, Stanford University Author
  • Adam Bennett Department of Health Informatics and Clinical Data Science, American University of Beirut Author

Abstract

Hospital readmission models are vulnerable to data leakage when predictor timestamps, data refresh schedules, and clinical events are not aligned with the intended prediction time. We audited an integrated health-information pipeline containing 64,732 adult admissions and reproduced a 30-day readmission model under three feature-extraction strategies. A conventional retrospective extract produced an AUROC of 0.79, but performance fell to 0.72 when all variables unavailable at discharge were removed and to 0.70 under prospective temporal validation. Several high-ranking predictors were generated or finalised after the nominal prediction time, including coding and reconciliation variables. Reconstructing data lineage allowed 11 leakage-prone features to be identified and removed. Although discrimination decreased, calibration improved in the later validation period after model updating. The study demonstrates that apparent readmission performance can be materially inflated by hidden timing assumptions in integrated hospital systems. Reproducible clinical prediction requires explicit lineage, feature-availability timestamps, and temporal validation that mirrors the intended deployment workflow.

References

1. Begum MH, Talukder SI, Ahmed MJ, Khan MFI, Azam MA. XAI-driven multimodal deep learning for early sepsis prediction in ICU. Journal of Science, Technology and Social Transformation. 2025;1(1):30-41. doi:10.64235/j62xmk30.

2. Azam MA, et al. Machine learning meets XAI: Grad-CAM visualization for enhanced lung and colon cancer detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546191/

3. Azam MA, et al. TransFuse: an interpretable transformer ensemble framework for Bangla smishing detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546480/

4. Azam MA, Ansari I, Haque GMM, Jahid A. Leveraging health information systems and predictive analytics to improve patient outcomes: a data-driven approach. The American Journal of Medical Sciences and Pharmaceutical Research. 2026;8(3):45-70. doi:10.37547/tajmspr/Volume08Issue03-06.

5. Johnson AEW, Bulgarelli L, Shen L, Gayles A, Shammout A, Horng S, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023;10:1. doi:10.1038/s41597-022-01899-x.

6. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006;26(6):565-574. doi:10.1177/0272989X06295361.

7. Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019;17:230. doi:10.1186/s12916-019-1466-7.

8. Kelly CJ, Karthikesalingam A, Suleyman M, Corrado G, King D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019;17:195. doi:10.1186/s12916-019-1426-2.

9. Breiman L. Random forests. Mach Learn. 2001;45:5-32. doi:10.1023/A:1010933404324.

10. Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016. p. 785-794. doi:10.1145/2939672.2939785.

11. Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330. Available from: https://proceedings.mlr.press/v70/guo17a.html

12. Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi:10.1371/journal.pone.0118432.

13. Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. doi:10.1016/j.patter.2023.100804.

14. Demšar J. Statistical comparisons of classifiers over multiple data sets. J Mach Learn Res. 2006;7:1-30. Available from: https://www.jmlr.org/papers/v7/demsar06a.html

15. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378.

16. Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58. doi:10.7326/M18-1376.

17. Riley RD, Snell KIE, Ensor J, Burke DL, Harrell FE Jr, Moons KGM, et al. Minimum sample size for developing a multivariable prediction model: PART II - binary and time-to-event outcomes. Stat Med. 2019;38(7):1276-1296. doi:10.1002/sim.7992.

18. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

Published

2026-06-01