Metadata-Intensive Versus Throughput-Intensive Workloads in Scale-Out File Storage: A FlexGroup Performance Evaluation

Authors

  • Rohan Kapoor Department of Computer Science and Cloud Systems, Lund University Author
  • Zain Haddad Department of Computer Science and Cloud Systems, University of Cambridge Author

Abstract

Scale-out file systems are expected to support both metadata-intensive and bandwidth-intensive applications, although their scaling behaviour differs substantially between workload classes. We evaluated FlexGroup performance using directory-creation, small-file, mixed metadata, sequential-read, and large-file write workloads across configurations ranging from 4 to 32 constituent volumes. Throughput, operations per second, latency distribution, CPU utilisation, and data-placement balance were measured under increasing client concurrency. Large-file throughput scaled near linearly through 16 constituents before network saturation became dominant. Metadata-intensive workloads improved more modestly and showed greater p99 latency variability, particularly when directory hot spots concentrated activity on a subset of constituents. Increasing parallelism without adjusting client distribution produced diminishing returns beyond 64 concurrent workers. A placement-aware test configuration reduced metadata tail latency by 18% compared with the default workload generator. The results emphasise that scale-out performance claims should report workload semantics and tail behaviour, not only aggregate throughput. Tuning strategies that benefit sequential data transfer may provide limited advantage for namespace-heavy applications.

References

1. Nazir M. Performance comparison of AWS EBS, EFS and S3 for enterprise cloud-storage workloads. International Journal of Business & Computational Sciences. 2022;2(1). Available from: https://ijbcs.org/index.php/IJBCS/article/view/2022-01-05

2. Nazir M. FlexGroup and FlexCache optimization for AWS-based CAE/HPC file workloads: a latency and throughput evaluation. International Journal of Business & Computational Sciences. 2023;3(1). Available from: https://ijbcs.org/index.php/IJBCS/article/view/2023-01-04

3. Ghemawat S, Gobioff H, Leung ST. The Google file system. In: Proceedings of the 19th ACM Symposium on Operating Systems Principles. 2003. p. 29-43. doi:10.1145/945445.945450.

4. DeCandia G, Hastorun D, Jampani M, Kakulapati G, Lakshman A, Pilchin A, et al. Dynamo: Amazon's highly available key-value store. In: Proceedings of the 21st ACM Symposium on Operating Systems Principles. 2007. p. 205-220. doi:10.1145/1294261.1294281.

5. Patterson DA, Gibson G, Katz RH. A case for redundant arrays of inexpensive disks (RAID). In: Proceedings of the 1988 ACM SIGMOD International Conference on Management of Data. 1988. p. 109-116. doi:10.1145/50202.50214.

6. Meyer DT, Bolosky WJ. A study of practical deduplication. In: Proceedings of the 9th USENIX Conference on File and Storage Technologies. 2011. p. 1-13. Available from: https://www.usenix.org/legacy/events/fast11/tech/techAbstracts.html#Meyer

7. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80. doi:10.1145/2408776.2408794.

8. Mytkowicz T, Diwan A, Hauswirth M, Sweeney PF. Producing wrong data without doing anything obviously wrong. In: Proceedings of the 14th International Conference on Architectural Support for Programming Languages and Operating Systems. 2009. p. 265-276. doi:10.1145/1508244.1508275.

9. Kalibera T, Jones R. Rigorous benchmarking in reasonable time. In: Proceedings of the 2013 International Symposium on Memory Management. 2013. p. 63-74. doi:10.1145/2464157.2464160.

10. Amazon Web Services. Amazon EBS volume types [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/ebs/latest/userguide/ebs-volume-types.html

11. Amazon Web Services. Amazon EFS performance [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/efs/latest/ug/performance.html

12. Amazon Web Services. Best practices design patterns: optimizing Amazon S3 performance [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimizing-performance.html

13. Axboe J, fio contributors. fio documentation [Internet]. [cited 2026 Sep 22]. Available from: https://fio.readthedocs.io/en/latest/fio_doc.html

14. IOR contributors. IOR documentation [Internet]. [cited 2026 Sep 22]. Available from: https://ior.readthedocs.io/en/latest/

15. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

16. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.

17. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.

18. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.

Published

2026-06-01