Missing-Modality Robustness in Explainable Sepsis Prediction: An External Validation Study

Authors

  • Inaya Haque Department of Health Informatics and Clinical Data Science, University of Chile Author
  • Iris Keller Department of Health Informatics and Clinical Data Science, University College Dublin Author

Abstract

Multimodal sepsis models may perform well when laboratory, vital-sign, and clinical-note inputs are complete but fail when one source is delayed or unavailable. We externally validated an explainable multimodal sepsis model in 21,406 admissions from a hospital not used for development and systematically removed individual modalities to measure robustness. With complete data, the model achieved an AUROC of 0.86 and acceptable calibration. Removing laboratory data produced the largest decline in discrimination, while removal of clinical text had a smaller effect on AUROC but reduced early-warning sensitivity for a subset of patients. A modality-dropout training strategy improved performance under missing inputs and reduced variation in predicted risk. Explanation consistency also deteriorated when modalities were absent; attributions sometimes shifted toward weak proxy variables. Recalibration on the external site improved probability accuracy without restoring all lost discrimination. External validation under realistic missing-modality conditions provides a more informative test of deployment readiness than complete-case evaluation alone. Robust sepsis prediction should include explicit missingness handling, site-level recalibration, and checks that explanations remain clinically coherent when input sources fail.

References

1. Azam MA, et al. Machine learning meets XAI: Grad-CAM visualization for enhanced lung and colon cancer detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546191/

2. Azam MA, et al. TransFuse: an interpretable transformer ensemble framework for Bangla smishing detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546480/

3. Azam MA, Ansari I, Haque GMM, Jahid A. Leveraging health information systems and predictive analytics to improve patient outcomes: a data-driven approach. The American Journal of Medical Sciences and Pharmaceutical Research. 2026;8(3):45-70. doi:10.37547/tajmspr/Volume08Issue03-06.

4. Khan MFI, Begum MH, Rahman MA, Limon GQ, Azam MA, Masum AKM. A comprehensive review of advances in transformer, GAN, and attention mechanisms: their role in multimodal learning and applications across NLP. International Journal of Science and Research Archive. 2025;15(1):454-459. doi:10.30574/ijsra.2025.15.1.0980.

5. Begum MH, Talukder SI, Ahmed MJ, Khan MFI, Azam MA. XAI-driven multimodal deep learning for early sepsis prediction in ICU. Journal of Science, Technology and Social Transformation. 2025;1(1):30-41. doi:10.64235/j62xmk30.

6. Johnson AEW, Bulgarelli L, Shen L, Gayles A, Shammout A, Horng S, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023;10:1. doi:10.1038/s41597-022-01899-x.

7. Singer M, Deutschman CS, Seymour CW, Shankar-Hari M, Annane D, Bauer M, et al. The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA. 2016;315(8):801-810. doi:10.1001/jama.2016.0287.

8. Wong A, Otles E, Donnelly JP, Krumm A, McCullough J, DeTroyer-Cooley O, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med. 2021;181(8):1065-1070. doi:10.1001/jamainternmed.2021.2626.

9. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006;26(6):565-574. doi:10.1177/0272989X06295361.

10. Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019;17:230. doi:10.1186/s12916-019-1466-7.

11. Kelly CJ, Karthikesalingam A, Suleyman M, Corrado G, King D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019;17:195. doi:10.1186/s12916-019-1426-2.

12. Breiman L. Random forests. Mach Learn. 2001;45:5-32. doi:10.1023/A:1010933404324.

13. Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016. p. 785-794. doi:10.1145/2939672.2939785.

14. Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330. Available from: https://proceedings.mlr.press/v70/guo17a.html

15. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378.

16. Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58. doi:10.7326/M18-1376.

17. Riley RD, Snell KIE, Ensor J, Burke DL, Harrell FE Jr, Moons KGM, et al. Minimum sample size for developing a multivariable prediction model: PART II - binary and time-to-event outcomes. Stat Med. 2019;38(7):1276-1296. doi:10.1002/sim.7992.

18. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

Published

2026-06-01