login

Reliability challenges in large systems

Future Generation Computer SystemsPublished 4 January 2005
Daniel A. Reed, Charng‐Da Lu, Celso L. Mendes
Citations60
SJR quartileQ1
SJR score1.55
SNIP2.23

TL;DR

This paper presents techniques for detecting imminent failures in the environment and that allow an application to run successfully despite such failures and shows how intelligent and adaptive software can lead to failure resilience and efficient system usage.

Abstract

Clusters built from commodity PCs dominate high-performance computing today, with systems containing thousands of processors now being deployed. As node counts for multi-teraflop systems grow to tens of thousands, with proposed petaflop system likely to contain hundreds of thousands of nodes, the assumption of fully reliable hardware and software becomes much less credible. In this paper, after presenting examples and experimental data that quantify the reliability of current systems, we describe possible approaches for effective system use. In particular, we present techniques for detecting imminent failures in the environment and that allow an application to run successfully despite such failures. We also show how intelligent and adaptive software can lead to failure resilience and efficient system usage.

Keywords

Computer ScienceEngineering