Block, File, and Object Storage for Enterprise Analytics Pipelines: A Semantics-Aware AWS Performance Benchmark
Abstract
Enterprise analytics pipelines increasingly span block, file, and object storage, but benchmark comparisons often ignore the access semantics that determine how each service is actually used. We evaluated all three storage classes on AWS using six analytics workloads covering sequential scans, metadata-heavy transforms, checkpointing, random updates, and large-object exchange. Tests were repeated across three data scales and two concurrency levels. Object storage delivered the highest cost efficiency for large sequential reads but incurred substantial overhead for fine-grained update patterns. File storage provided the most consistent end-to-end runtime for workflows requiring shared namespace semantics, while block storage achieved the lowest p99 latency for transactional scratch operations. At high concurrency, metadata-bound file workloads degraded more sharply than throughput-oriented object workloads. No storage class was uniformly superior across the pipeline. A semantics-aware placement strategy that matched stages to their dominant access pattern reduced total workflow time by 17% and estimated storage cost by 13% relative to a single-service baseline. Storage selection for analytics should therefore be based on operation semantics and pipeline stage rather than headline throughput figures alone.
References
1. Nazir M. Performance comparison of AWS EBS, EFS and S3 for enterprise cloud-storage workloads. International Journal of Business & Computational Sciences. 2022;2(1). Available from: https://ijbcs.org/index.php/IJBCS/article/view/2022-01-05
2. Ghemawat S, Gobioff H, Leung ST. The Google file system. In: Proceedings of the 19th ACM Symposium on Operating Systems Principles. 2003. p. 29-43. doi:10.1145/945445.945450.
3. DeCandia G, Hastorun D, Jampani M, Kakulapati G, Lakshman A, Pilchin A, et al. Dynamo: Amazon's highly available key-value store. In: Proceedings of the 21st ACM Symposium on Operating Systems Principles. 2007. p. 205-220. doi:10.1145/1294261.1294281.
4. Patterson DA, Gibson G, Katz RH. A case for redundant arrays of inexpensive disks (RAID). In: Proceedings of the 1988 ACM SIGMOD International Conference on Management of Data. 1988. p. 109-116. doi:10.1145/50202.50214.
5. Meyer DT, Bolosky WJ. A study of practical deduplication. In: Proceedings of the 9th USENIX Conference on File and Storage Technologies. 2011. p. 1-13. Available from: https://www.usenix.org/legacy/events/fast11/tech/techAbstracts.html#Meyer
6. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80. doi:10.1145/2408776.2408794.
7. Mytkowicz T, Diwan A, Hauswirth M, Sweeney PF. Producing wrong data without doing anything obviously wrong. In: Proceedings of the 14th International Conference on Architectural Support for Programming Languages and Operating Systems. 2009. p. 265-276. doi:10.1145/1508244.1508275.
8. Kalibera T, Jones R. Rigorous benchmarking in reasonable time. In: Proceedings of the 2013 International Symposium on Memory Management. 2013. p. 63-74. doi:10.1145/2464157.2464160.
9. Amazon Web Services. Amazon EBS volume types [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/ebs/latest/userguide/ebs-volume-types.html
10. Amazon Web Services. Amazon EFS performance [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/efs/latest/ug/performance.html
11. Amazon Web Services. Best practices design patterns: optimizing Amazon S3 performance [Internet]. [cited 2026 Sep 22]. Available from: https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimizing-performance.html
12. Axboe J, fio contributors. fio documentation [Internet]. [cited 2026 Sep 22]. Available from: https://fio.readthedocs.io/en/latest/fio_doc.html
13. IOR contributors. IOR documentation [Internet]. [cited 2026 Sep 22]. Available from: https://ior.readthedocs.io/en/latest/
14. NetApp. Learn about ONTAP FlexGroup volumes [Internet]. [cited 2026 Sep 22]. Available from: https://docs.netapp.com/us-en/ontap/flexgroup/index.html
15. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.
16. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.
17. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.
18. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.
Published
Issue
Section
License
Authors retain copyright. Articles published under the Creative Commons Attribution 4.0 International licence (CC BY 4.0) may be shared and adapted for any purpose, including commercially, provided appropriate credit is given, a link to the licence is supplied, and changes are indicated. No additional legal or technological restrictions may be imposed. Third-party material is included only where its credit line permits. Licence: https://creativecommons.org/licenses/by/4.0/. Earlier publications remain subject to their stated licence and author agreements unless the rights holder authorizes a change.