article
The growing adoption of big data across sectors has triggered a significant transformation in data architecture, shifting from monolithic systems to more dynamic and scalable ecosystems. Initially dominated by Hadoop-based frameworks relying on tools like HDFS and Spark big data processing has since evolved towards more modular architectures. The modern data stack introduces flexibility and tool diversity, while cloud-native platforms redefine scalability and simplify integration. Yet, selecting the appropriate stack remains a complex task, as most prior research focuses on isolated components rather than holistic, practical comparisons. This paper addresses that gap by evaluating three prevalent data architecture paradigms: Hadoop, Modern, and Cloud-Based stacks. We conduct a thorough end-to-end comparison across the full big data value chain, from ingestion to visualization. Moreover, beyond architectural assessment, we implement each stack on a real-world scenario using the Amazon Books Reviews dataset where each implementation is evaluated based on key metrics such as scalability, performance, ease, and deployment costs. Our findings aim to provide data professionals with a practical reference for selecting suitable data architectures in big data environments.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/icoa66896.2025.11236898
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.