Metadata-Intensive Versus Throughput-Intensive Workloads in Scale-Out File Storage: A FlexGroup Performance Evaluation
Abstract
Scale-out file systems are expected to support both metadata-intensive and bandwidth-intensive applications, although their scaling behaviour differs substantially between workload classes. We evaluated FlexGroup performance using directory-creation, small-file, mixed metadata, sequential-read, and large-file write workloads across configurations ranging from 4 to 32 constituent volumes. Throughput, operations per second, latency distribution, CPU utilisation, and data-placement balance were measured under increasing client concurrency. Large-file throughput scaled near linearly through 16 constituents before network saturation became dominant. Metadata-intensive workloads improved more modestly and showed greater p99 latency variability, particularly when directory hot spots concentrated activity on a subset of constituents. Increasing parallelism without adjusting client distribution produced diminishing returns beyond 64 concurrent workers. A placement-aware test configuration reduced metadata tail latency by 18% compared with the default workload generator. The results emphasise that scale-out performance claims should report workload semantics and tail behaviour, not only aggregate throughput. Tuning strategies that benefit sequential data transfer may provide limited advantage for namespace-heavy applications.
References
1. Nazir M. Performance comparison of AWS EBS, EFS and S3 for enterprise cloud-storage workloads. International Journal of Business & Computational Sciences. 2022;2(1). Available from: https://ijbcs.org/index.php/IJBCS/article/view/2022-01-05
2. Nazir M. FlexGroup and FlexCache optimization for AWS-based CAE/HPC file workloads: a latency and throughput evaluation. International Journal of Business & Computational Sciences. 2023;3(1). Available from: https://ijbcs.org/index.php/IJBCS/article/view/2023-01-04
3. Ghemawat S, Gobioff H, Leung ST. The Google file system. In: Proceedings of the 19th ACM Symposium on Operating Systems Principles. 2003. p. 29-43. doi:10.1145/945445.945450.
4. DeCandia G, Hastorun D, Jampani M, Kakulapati G, Lakshman A, Pilchin A, et al. Dynamo: Amazon's highly available key-value store. In: Proceedings of the 21st ACM Symposium on Operating Systems Principles. 2007. p. 205-220. doi:10.1145/1294261.1294281.
5. Patterson DA, Gibson G, Katz RH. A case for redundant arrays of inexpensive disks (RAID). In: Proceedings of the 1988 ACM SIGMOD International Conference on Management of Data. 1988. p. 109-116. doi:10.1145/50202.50214.
6. Meyer DT, Bolosky WJ. A study of practical deduplication. In: Proceedings of the 9th USENIX Conference on File and Storage Technologies. 2011. p. 1-13. Available from: https://www.usenix.org/legacy/events/fast11/tech/techAbstracts.html#Meyer
7. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80. doi:10.1145/2408776.2408794.
8. Mytkowicz T, Diwan A, Hauswirth M, Sweeney PF. Producing wrong data without doing anything obviously wrong. In: Proceedings of the 14th International Conference on Architectural Support for Programming Languages and Operating Systems. 2009. p. 265-276. doi:10.1145/1508244.1508275.
9. Kalibera T, Jones R. Rigorous benchmarking in reasonable time. In: Proceedings of the 2013 International Symposium on Memory Management. 2013. p. 63-74. doi:10.1145/2464157.2464160.
10. Amazon Web Services. Amazon EBS volume types [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/ebs/latest/userguide/ebs-volume-types.html
11. Amazon Web Services. Amazon EFS performance [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/efs/latest/ug/performance.html
12. Amazon Web Services. Best practices design patterns: optimizing Amazon S3 performance [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimizing-performance.html
13. Axboe J, fio contributors. fio documentation [Internet]. [cited 2026 Sep 22]. Available from: https://fio.readthedocs.io/en/latest/fio_doc.html
14. IOR contributors. IOR documentation [Internet]. [cited 2026 Sep 22]. Available from: https://ior.readthedocs.io/en/latest/
15. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.
16. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.
17. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.
18. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.
Published
Issue
Section
License
Authors retain copyright. Articles published under the Creative Commons Attribution 4.0 International licence (CC BY 4.0) may be shared and adapted for any purpose, including commercially, provided appropriate credit is given, a link to the licence is supplied, and changes are indicated. No additional legal or technological restrictions may be imposed. Third-party material is included only where its credit line permits. Licence: https://creativecommons.org/licenses/by/4.0/. Earlier publications remain subject to their stated licence and author agreements unless the rights holder authorizes a change.