Performance and energy efficiency of big data applications in cloud environments: A Hadoop case study
Journal of Parallel and Distributed ComputingPublished 9 January 2015Open access
Eugen Feller, Lavanya Ramakrishnan, Christine Morin
Citations68
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
An energy efficiency evaluation of Hadoop on physical and virtual clusters in different configurations and a discussion on the implications of using cloud environments for big data analyses are presented.
Abstract
International audience
Keywords
Computer Science
INTERNATIONAL JOURNAL OF RESEARCH AND ENGINEERINGMapReduce: Simplified Data Processing on Large Cluster
2,972 Citations2018Jay B. Dean
The implementation of MapReduce runs on a large cluster of commodity machines and is highly scalable: a typical MapReduce computation processes many terabytes of data on thousands of machines.
eScholarship (California Digital Library)Ceph: a scalable, high-performance distributed file system
1,430 Citations2006Sage A. Weil, Scott Brandt +3 more
Performance measurements under a variety of workloads show that Ceph has excellent I/O performance and scalable metadata management, supporting more than 250,000 metadata operations per second.
The International Journal of High Performance Computing ApplicationsGrid'5000: A Large Scale And Highly Reconfigurable Experimental Grid Testbed
461 Citations2006Raphaël Bolze, Franck Cappello +15 more
This paper presents Grid'5000, a 5000 CPU nation-wide infrastructure for research in Grid computing, designed to provide a scientific tool for computer scientists similar to the large-scale instruments used by physicists, astronomers, and biologists.
Proceedings of the VLDB EndowmentThe performance of MapReduce
394 Citations2010Dawei Jiang, Beng Chin Ooi +2 more
By carefully tuning these factors, the overall performance of Hadoop can be improved by a factor of 2.5 to 3.5, and is thus more comparable to that of parallel database systems.
Dynamic voltage and frequency scaling: the laws of diminishing returns
385 Citations2010Etienne Le Sueur, Gernot Heiser
It is found that while DVFS is effective on the older platforms, it actually increases energy usage on the most recent platform, even for highly memory-bound workloads.
MapReduce for Data Intensive Scientific Analyses
378 Citations2008Jaliya Ekanayake, Shrideep Pallickara +1 more
This paper presents CGL-MapReduce, a streaming-based MapReduce implementation and compares its performance with Hadoop, and presents the experience in applying the MapReduced technique for two scientific data analyses: high energy physics data analyses; and K-means clustering.
ACM SIGOPS Operating Systems ReviewOn the energy (in)efficiency of Hadoop clusters
339 Citations2010Jacob Leverich, Christos Kozyrakis
It is found that running Hadoop clusters in fractional configurations can save between 9% and 50% of energy consumption, and that there is a tradeoff between performance energy consumption.
The Hadoop distributed filesystem: Balancing portability and performance
296 Citations2010Jeffrey Shafer, Scott Rixner +1 more
The performance of HDFS is analyzed and several performance issues are uncovered, including architectural bottlenecks exist in the Hadoop implementation that result in inefficient HDFS usage due to delays in scheduling new MapReduce tasks.
GreenHadoop
295 Citations2012Íñigo Goiri, Kien Le +4 more
GreenHadoop is proposed, a MapReduce framework for a datacenter powered by a photovoltaic solar array and the electrical grid (as a backup) and can significantly increase green energy consumption and decrease electricity cost, compared to Hadoop.
USENIX Annual Technical ConferenceHigh performance VMM-bypass I/O in virtual machines
272 Citations2006Jiuxing Liu, Wei Huang +2 more
VMM-bypass allows time-critical I/O operations to be carried out directly in guest VMs without involvement of the VMM and/or a privileged VM by exploiting the intelligence found in modern high speed network interfaces.
Proceedings of the VLDB EndowmentEnergy management for MapReduce clusters
217 Citations2010Willis Lang, Jignesh M. Patel
This paper develops a framework for systematically considering various MapReduce node power down strategies, and their impact on the overall energy consumption and workload response time and proposes the All-In Strategy (AIS), which is often the right energy saving strategy.
Energy efficiency for large-scale MapReduce workloads with significant interactive analysis
174 Citations2012Yanpei Chen, Sara Alspaugh +2 more
The key insight is that although MIA clusters host huge data volumes, the interactive jobs operate on a small fraction of the data, and thus can be served by a small pool of dedicated machines; the less time-sensitive jobs can run on the rest of the cluster in a batch fashion.
Lecture notes in computer scienceEvaluating MapReduce on Virtual Machines: The Hadoop Case
132 Citations2009Shadi Ibrahim, Hai Jin +4 more
A series of experiments are conducted to measure and analyze the performance of Hadoop on VMs and outline several issues that will need to be considered when implementing MapReduce to fit completely in the cloud.
Snooze: A Scalable and Autonomic Virtual Machine Management Framework for Private Clouds
132 Citations2012Eugen Feller, Louis Rilling +1 more
The design, implementation and evaluation of a novel scalable and autonomic virtual machine (VM) management framework called Snooze, which utilizes a self-organizing hierarchical architecture and performs distributed VM management, shows that the fault tolerance features of the framework do not impact application performance.
Improving MapReduce energy efficiency for computation intensive workloads
81 Citations2011Thomas Wirtz, Rong Ge
This paper studies the performance and energy efficiency of the Hadoop implementation of MapReduce under the context of energy-proportional computing, and indicates significant energy savings can be achieved from judicious resource allocation and intelligent DVFS scheduling for computation intensive applications.
I/O performance of virtualized cloud environments
77 Citations2011Devarshi Ghoshal, R. S. Canon +1 more
The results in benchmarking the I/O performance over different cloud and HPC platforms to identify the major bottlenecks in existing infrastructure will help applications decide between the different storage options enabling applications to make effective choices.
Resilin: Elastic MapReduce over Multiple Clouds
51 Citations2013Andrei Vasile Iordache, Christine Morin +3 more
Resilin goes one step beyond Amazon's proprietary EMR solution and allows users to leverage resources from one or multiple public and/or private clouds and gives Resilin users the opportunity to perform MapReduce computations over a large number of potentially geographically distributed resources.
Evaluating Hadoop for Data-Intensive Scientific Operations
33 Citations2012Zacharia Fadika, Madhusudhan Govindaraju +2 more
Evaluating Hadoop specifically for data-intensive scientific operations -- filter, merge and reorder-- in the context of High Performance Computing (HPC) environments to understand the impact of the file system, network and programming modes on performance.
Big data @ facebook
30 Citations2012Aravind Menon
The Data Warehousing and Analytics platform of Facebook that provides support for batch-oriented analytics applications and a rich set of tools for different users to perform analytics queries on Facebook data is focused on.
Lecture notes in computer scienceMapReduce Applications in the Cloud: A Cost Evaluation of Computation and Storage
6 Citations2012Diana Moise, Alexandra Carpen-Amarie
The overhead implied by the execution of MapReduce jobs in the Cloud, compared to an execution on a Grid, and the actual costs of renting the corresponding Cloud resources are looked into.
Lecture notes in computer scienceData Management in Cloud, Grid and P2P Systems
1 Citations2012Abdelkader Hameurlain, A Min Tjoa
The proposed solution permits to handle transactions rapidly while using few middleware resources to reduce financial costs and adjusts the optimal number of nodes according to the workload variation.
