Forecast-Aware Versus Threshold-Based Capacity Expansion in Thin-Provisioned Storage: A Trace-Driven Simulation

Authors

  • Ibrahim Farouk Department of Computer Science and Cloud Systems, Imperial College London Author
  • Zain Haddad Department of Computer Science and Cloud Systems, University of Cambridge Author
  • Omar Malik Department of Computer Science and Cloud Systems, Aga Khan University Author

Abstract

Threshold-based expansion is simple to operate but can react too late when thin-provisioned storage experiences rapid growth. We compared fixed-threshold expansion with a forecast-aware policy using 12 months of anonymised capacity traces from 48 enterprise volumes. A rolling forecasting model estimated seven-day utilisation trajectories and triggered expansion when projected headroom breached a risk-adjusted limit. Both policies were replayed under identical workload traces with expansion delay, demand volatility, and forecast error varied in sensitivity analyses. Forecast-aware expansion reduced critical-capacity events by 63% and lowered the number of emergency expansions from 94 to 31. It also decreased unused allocated capacity by 12.7% relative to a conservative 80% fixed threshold. Benefits were greatest in workloads with recurring growth patterns, while highly irregular volumes generated more false-positive expansions. Forecast errors above 20% materially reduced the advantage of the predictive policy. A hybrid strategy combining minimum safety thresholds with short-horizon forecasts provided the most stable performance across workloads. These results indicate that capacity expansion can be made more efficient when predictive signals complement, rather than replace, deterministic safety controls.

References

1. Nazir M. Capacity reclamation and thin-provisioning efficiency in large-scale enterprise storage systems. International Journal of Business & Computational Sciences. 2021;1(1). Available from: https://ijbcs.org/index.php/IJBCS/article/view/2021P4

2. Patterson DA, Gibson G, Katz RH. A case for redundant arrays of inexpensive disks (RAID). In: Proceedings of the 1988 ACM SIGMOD International Conference on Management of Data. 1988. p. 109-116. doi:10.1145/50202.50214.

3. Meyer DT, Bolosky WJ. A study of practical deduplication. In: Proceedings of the 9th USENIX Conference on File and Storage Technologies. 2011. p. 1-13. Available from: https://www.usenix.org/legacy/events/fast11/tech/techAbstracts.html#Meyer

4. Linux kernel contributors. Thin provisioning [Internet]. [cited 2026 Sep 22]. Available from: https://docs.kernel.org/admin-guide/device-mapper/thin-provisioning.html

5. util-linux contributors. fstrim(8) - discard unused blocks on a mounted filesystem [Internet]. [cited 2026 Sep 22]. Available from: https://man7.org/linux/man-pages/man8/fstrim.8.html

6. Axboe J, fio contributors. fio documentation [Internet]. [cited 2026 Sep 22]. Available from: https://fio.readthedocs.io/en/latest/fio_doc.html

7. Mytkowicz T, Diwan A, Hauswirth M, Sweeney PF. Producing wrong data without doing anything obviously wrong. In: Proceedings of the 14th International Conference on Architectural Support for Programming Languages and Operating Systems. 2009. p. 265-276. doi:10.1145/1508244.1508275.

8. Kalibera T, Jones R. Rigorous benchmarking in reasonable time. In: Proceedings of the 2013 International Symposium on Memory Management. 2013. p. 63-74. doi:10.1145/2464157.2464160.

9. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80. doi:10.1145/2408776.2408794.

10. Ghemawat S, Gobioff H, Leung ST. The Google file system. In: Proceedings of the 19th ACM Symposium on Operating Systems Principles. 2003. p. 29-43. doi:10.1145/945445.945450.

11. DeCandia G, Hastorun D, Jampani M, Kakulapati G, Lakshman A, Pilchin A, et al. Dynamo: Amazon's highly available key-value store. In: Proceedings of the 21st ACM Symposium on Operating Systems Principles. 2007. p. 205-220. doi:10.1145/1294261.1294281.

12. Swanson M, Bowen P, Phillips AW, Gallup D, Lynes D. Contingency planning guide for federal information systems. Gaithersburg (MD): National Institute of Standards and Technology. 2010;NIST SP 800-34 Rev. 1. doi:10.6028/NIST.SP.800-34r1.

13. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

14. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.

15. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.

16. Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863.

17. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.

18. Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Series B Stat Methodol. 1995;57(1):289-300. doi:10.1111/j.2517-6161.1995.tb02031.x.

Published

2026-06-01