login

Wide area cluster monitoring with Ganglia

Published 1 January 2003
Sacerdoti, Katz, Massie, Culler
Citations180

TL;DR

A structure for monitoring a large set of computational clusters comprised of many clusters while keeping processing requirements low is presented and emphasis is placed on scalability, fast query response, fault tolerance, and grid compatibility.

Abstract

In this paper, we present a structure for monitoring a large set of computational clusters. We illustrate methods for scaling a monitor network comprised of many clusters while keeping processing requirements low. A design for presenting high-level Web-based summaries of the monitor network is provided, along with a generalization to a distributed, multiple-resolution monitoring tree. Emphasis is placed on scalability, fast query response, fault tolerance, and grid compatibility. Experimental evidence is presented that demonstrates the performance of our design.

Keywords

Computer Science