Adaptive Snapshot Scheduling for Asynchronous Replication Under Bursty Write Workloads: A Simulation Study of Recovery-Point Compliance
Abstract
Fixed snapshot intervals used for asynchronous replication can either waste resources during quiet periods or violate recovery-point objectives during bursts of write activity. We evaluated an adaptive scheduling policy that adjusted snapshot frequency according to recent changed-data rate, replication backlog, and available bandwidth. A discrete-event simulation replayed 90 production-like workload traces covering databases, virtual machines, and file services under four network-capacity profiles. Recovery-point compliance, snapshot count, replication traffic, and backlog duration were compared with 5-, 15-, and 30-minute fixed schedules. The adaptive policy maintained the target recovery point in 96.8% of evaluated intervals, compared with 88.1% for the best fixed schedule under bursty conditions. It reduced snapshot operations by 24% relative to the 5-minute policy and lowered prolonged backlog episodes by 37%. Performance deteriorated when bandwidth estimates lagged sudden network degradation, but incorporating a conservative backlog margin mitigated this effect. Adaptive snapshot scheduling can improve recovery-point compliance without permanently operating at the most aggressive interval. Practical implementations should combine workload-aware timing with explicit safeguards for estimation error and network volatility.
References
1. Nazir M. RTO and RPO optimization using asynchronous replication in multi-petabyte enterprise storage environments. Journal of Computing, Intelligence and Information Sciences. 2021. Available from: https://jciis.com/index.php/jciis/article/view/2020-01-04
2. Nazir M. Automated SnapMirror migration from on-premises storage to Cloud Volumes ONTAP: a replication performance study. Journal of Multidisciplinary Research and Integrated Sciences. 2022;2(1). Available from: https://jmris.com/index.php/jmris/article/view/2022V1P6
3. Patterson RH, Manley S, Federwisch M, Hitz D, Kleiman S, Owara S. SnapMirror: file-system-based asynchronous mirroring for disaster recovery. In: Proceedings of the 1st USENIX Conference on File and Storage Technologies. 2002. p. 117-129. Available from: https://www.usenix.org/conference/fast-02/snapmirror-file-system-based-asynchronous-mirroring-disaster-recovery
4. Cully B, Lefebvre G, Meyer D, Feeley M, Hutchinson N, Warfield A. Remus: high availability via asynchronous virtual machine replication. In: Proceedings of the 5th USENIX Symposium on Networked Systems Design and Implementation. 2008. p. 161-174. Available from: https://www.usenix.org/legacy/events/nsdi08/tech/full_papers/cully/cully.pdf
5. Gray J, Lamport L. Consensus on transaction commit. ACM Trans Database Syst. 2006;31(1):133-160. doi:10.1145/1132863.1132867.
6. Gilbert S, Lynch N. Brewer's conjecture and the feasibility of consistent, available, partition-tolerant web services. ACM SIGACT News. 2002;33(2):51-59. doi:10.1145/564585.564601.
7. Ghemawat S, Gobioff H, Leung ST. The Google file system. In: Proceedings of the 19th ACM Symposium on Operating Systems Principles. 2003. p. 29-43. doi:10.1145/945445.945450.
8. DeCandia G, Hastorun D, Jampani M, Kakulapati G, Lakshman A, Pilchin A, et al. Dynamo: Amazon's highly available key-value store. In: Proceedings of the 21st ACM Symposium on Operating Systems Principles. 2007. p. 205-220. doi:10.1145/1294261.1294281.
9. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80. doi:10.1145/2408776.2408794.
10. NetApp. Learn about ONTAP SnapMirror asynchronous disaster recovery [Internet]. [cited 2026 Sep 22]. Available from: https://docs.netapp.com/us-en/ontap/data-protection/snapmirror-disaster-recovery-concept.html
11. Swanson M, Bowen P, Phillips AW, Gallup D, Lynes D. Contingency planning guide for federal information systems. Gaithersburg (MD): National Institute of Standards and Technology. 2010;NIST SP 800-34 Rev. 1. doi:10.6028/NIST.SP.800-34r1.
12. Grance T, Nolan T, Burke K, Dudley R, White G, Good T. Guide to test, training, and exercise programs for IT plans and capabilities. Gaithersburg (MD): National Institute of Standards and Technology. 2006;NIST SP 800-84. doi:10.6028/NIST.SP.800-84.
13. Mytkowicz T, Diwan A, Hauswirth M, Sweeney PF. Producing wrong data without doing anything obviously wrong. In: Proceedings of the 14th International Conference on Architectural Support for Programming Languages and Operating Systems. 2009. p. 265-276. doi:10.1145/1508244.1508275.
14. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.
15. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.
16. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.
17. Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863.
18. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.
Published
Issue
Section
License
Authors retain copyright. Articles published under the Creative Commons Attribution 4.0 International licence (CC BY 4.0) may be shared and adapted for any purpose, including commercially, provided appropriate credit is given, a link to the licence is supplied, and changes are indicated. No additional legal or technological restrictions may be imposed. Third-party material is included only where its credit line permits. Licence: https://creativecommons.org/licenses/by/4.0/. Earlier publications remain subject to their stated licence and author agreements unless the rights holder authorizes a change.