articleJan 1, 2010Closed access

The HiBench benchmark suite: Characterization of the MapReduce-based data analysis

Indexed incrossref

Abstract

The MapReduce model is becoming prominent for the large-scale data analysis in the cloud. In this paper, we present the benchmarking, evaluation and characterization of Hadoop, an open-source implementation of MapReduce. We first introduce HiBench, a new benchmark suite for Hadoop. It consists of a set of Hadoop programs, including both synthetic micro-benchmarks and real-world Hadoop applications. We then evaluate and characterize the Hadoop framework using HiBench, in terms of speed (i.e., job running time), throughput (i.e., the number of tasks completed per minute), HDFS bandwidth, system resource (e.g., CPU, memory and I/O) utilizations, and data access patterns.

Citation impact

711
total citations
FWCI
47.13
Percentile
100%
References
6
Citations per year

Authors

5

Topics & keywords

Keywords
  • Computer science
  • Suite
  • Benchmarking
  • Benchmark (surveying)
  • Cloud computing
  • Operating system
  • Big data
  • Parallel computing
No related works found for this paper.