Snapshot Retention, Deduplication, and Reclaimable Capacity: A Factorial Evaluation of Enterprise Storage Policies

Authors

  • Maya Aziz Department of Computer Science and Cloud Systems, University College London (UCL) Author
  • Bilal Nawaz Department of Computer Science and Cloud Systems, University of Pretoria Author

Abstract

Snapshot retention and deduplication policies are often tuned independently, although both determine how much provisioned storage can ultimately be reclaimed. A 3×3 factorial experiment examined the interaction between snapshot retention length and deduplication intensity across 270 workload runs representing databases, virtual desktops, document repositories, and engineering files. Capacity consumed, unique-block ratio, reclaimable space, deletion latency, and read/write performance were measured over repeated snapshot cycles. Longer retention predictably increased logical snapshot footprint, but the effect on physical consumption varied substantially by workload. High deduplication reduced physical growth by 28–46% for repetitive workloads but by less than 9% for already compressed data. Shortening retention from 30 to 7 recovery points increased reclaimable capacity by a median of 21%, although gains were smaller where deduplication ratios were high. No meaningful change in median read latency was observed, while snapshot deletion caused brief p99 write-latency increases under the most aggressive policy combination. The results support joint optimisation of retention and data-reduction settings. Policy decisions based solely on logical snapshot size may misrepresent actual capacity pressure and reclamation potential.

References

1. Nazir M. Capacity reclamation and thin-provisioning efficiency in large-scale enterprise storage systems. International Journal of Business & Computational Sciences. 2021;1(1). Available from: https://ijbcs.org/index.php/IJBCS/article/view/2021P4

2. Patterson DA, Gibson G, Katz RH. A case for redundant arrays of inexpensive disks (RAID). In: Proceedings of the 1988 ACM SIGMOD International Conference on Management of Data. 1988. p. 109-116. doi:10.1145/50202.50214.

3. Meyer DT, Bolosky WJ. A study of practical deduplication. In: Proceedings of the 9th USENIX Conference on File and Storage Technologies. 2011. p. 1-13. Available from: https://www.usenix.org/legacy/events/fast11/tech/techAbstracts.html#Meyer

4. Linux kernel contributors. Thin provisioning [Internet]. [cited 2026 Sep 22]. Available from: https://docs.kernel.org/admin-guide/device-mapper/thin-provisioning.html

5. util-linux contributors. fstrim(8) - discard unused blocks on a mounted filesystem [Internet]. [cited 2026 Sep 22]. Available from: https://man7.org/linux/man-pages/man8/fstrim.8.html

6. Axboe J, fio contributors. fio documentation [Internet]. [cited 2026 Sep 22]. Available from: https://fio.readthedocs.io/en/latest/fio_doc.html

7. Mytkowicz T, Diwan A, Hauswirth M, Sweeney PF. Producing wrong data without doing anything obviously wrong. In: Proceedings of the 14th International Conference on Architectural Support for Programming Languages and Operating Systems. 2009. p. 265-276. doi:10.1145/1508244.1508275.

8. Kalibera T, Jones R. Rigorous benchmarking in reasonable time. In: Proceedings of the 2013 International Symposium on Memory Management. 2013. p. 63-74. doi:10.1145/2464157.2464160.

9. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80. doi:10.1145/2408776.2408794.

10. Ghemawat S, Gobioff H, Leung ST. The Google file system. In: Proceedings of the 19th ACM Symposium on Operating Systems Principles. 2003. p. 29-43. doi:10.1145/945445.945450.

11. DeCandia G, Hastorun D, Jampani M, Kakulapati G, Lakshman A, Pilchin A, et al. Dynamo: Amazon's highly available key-value store. In: Proceedings of the 21st ACM Symposium on Operating Systems Principles. 2007. p. 205-220. doi:10.1145/1294261.1294281.

12. Swanson M, Bowen P, Phillips AW, Gallup D, Lynes D. Contingency planning guide for federal information systems. Gaithersburg (MD): National Institute of Standards and Technology. 2010;NIST SP 800-34 Rev. 1. doi:10.6028/NIST.SP.800-34r1.

13. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

14. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.

15. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.

16. Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863.

17. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.

18. Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Series B Stat Methodol. 1995;57(1):289-300. doi:10.1111/j.2517-6161.1995.tb02031.x.

Published

2026-06-01