Consistency Verification After Storage Virtual Machine Failover: Development and Evaluation of a Reproducible Test Framework

Authors

  • Sofia Rossi Department of Computer Science and Cloud Systems, University of Michigan Author
  • Rohan Kapoor Department of Computer Science and Cloud Systems, Lund University Author

Abstract

Post-failover checks are frequently implemented as ad hoc scripts, making it difficult to demonstrate that a storage virtual machine has recovered consistently across data, configuration, identity, and network layers. We developed a reproducible verification framework containing 37 automated assertions covering volume availability, export or share configuration, name resolution, authentication, replication state, snapshot consistency, and application read/write tests. The framework was evaluated in 120 planned and fault-induced failover events across four storage configurations. It detected 94 of 101 deliberately introduced inconsistencies, yielding a sensitivity of 93.1%, while 6 false alerts were generated during transient service stabilisation. Median verification time was 6.4 minutes compared with 22.7 minutes for the existing manual checklist. Repeated execution produced identical classifications in 98% of events. Most missed defects involved application-level semantics not visible to infrastructure checks. The results show that post-failover validation can be standardised and substantially accelerated through a layered assertion framework. Automated verification should complement, rather than replace, application-specific acceptance tests when business consistency depends on domain-level transactions.

References

1. Nazir M. SVM-based disaster recovery for hybrid enterprise storage: effects on recovery time and service availability. Journal of Computing, Intelligence and Information Sciences. 2022. Available from: https://jciis.com/index.php/jciis/article/view/2022-01-05

2. Patterson RH, Manley S, Federwisch M, Hitz D, Kleiman S, Owara S. SnapMirror: file-system-based asynchronous mirroring for disaster recovery. In: Proceedings of the 1st USENIX Conference on File and Storage Technologies. 2002. p. 117-129. Available from: https://www.usenix.org/conference/fast-02/snapmirror-file-system-based-asynchronous-mirroring-disaster-recovery

3. Cully B, Lefebvre G, Meyer D, Feeley M, Hutchinson N, Warfield A. Remus: high availability via asynchronous virtual machine replication. In: Proceedings of the 5th USENIX Symposium on Networked Systems Design and Implementation. 2008. p. 161-174. Available from: https://www.usenix.org/legacy/events/nsdi08/tech/full_papers/cully/cully.pdf

4. Gray J, Lamport L. Consensus on transaction commit. ACM Trans Database Syst. 2006;31(1):133-160. doi:10.1145/1132863.1132867.

5. Gilbert S, Lynch N. Brewer's conjecture and the feasibility of consistent, available, partition-tolerant web services. ACM SIGACT News. 2002;33(2):51-59. doi:10.1145/564585.564601.

6. Ghemawat S, Gobioff H, Leung ST. The Google file system. In: Proceedings of the 19th ACM Symposium on Operating Systems Principles. 2003. p. 29-43. doi:10.1145/945445.945450.

7. DeCandia G, Hastorun D, Jampani M, Kakulapati G, Lakshman A, Pilchin A, et al. Dynamo: Amazon's highly available key-value store. In: Proceedings of the 21st ACM Symposium on Operating Systems Principles. 2007. p. 205-220. doi:10.1145/1294261.1294281.

8. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80. doi:10.1145/2408776.2408794.

9. NetApp. Learn about ONTAP SnapMirror asynchronous disaster recovery [Internet]. [cited 2026 Sep 22]. Available from: https://docs.netapp.com/us-en/ontap/data-protection/snapmirror-disaster-recovery-concept.html

10. Swanson M, Bowen P, Phillips AW, Gallup D, Lynes D. Contingency planning guide for federal information systems. Gaithersburg (MD): National Institute of Standards and Technology. 2010;NIST SP 800-34 Rev. 1. doi:10.6028/NIST.SP.800-34r1.

11. Grance T, Nolan T, Burke K, Dudley R, White G, Good T. Guide to test, training, and exercise programs for IT plans and capabilities. Gaithersburg (MD): National Institute of Standards and Technology. 2006;NIST SP 800-84. doi:10.6028/NIST.SP.800-84.

12. Mytkowicz T, Diwan A, Hauswirth M, Sweeney PF. Producing wrong data without doing anything obviously wrong. In: Proceedings of the 14th International Conference on Architectural Support for Programming Languages and Operating Systems. 2009. p. 265-276. doi:10.1145/1508244.1508275.

13. Kalibera T, Jones R. Rigorous benchmarking in reasonable time. In: Proceedings of the 2013 International Symposium on Memory Management. 2013. p. 63-74. doi:10.1145/2464157.2464160.

14. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

15. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.

16. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.

17. Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863.

18. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.

Published

2026-06-01