Embracing Failure: A Case for Recovery-Oriented Computing (ROC)
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
There is a fundamental mismatch between traditional high-availability approaches—fault-tolerant hardware, careful software testing, vendor-supplied technicians—and the realities of modern heterogeneous, distributed server environments, like those backing e-commerce and e-business sites.
Abstract
Motivated by the lack of availability demonstrated by current approaches to building servers for the Internet environment, we argue for a new approach to building highly-available systems that better reflects the realities of the modern server environment, namely that failures of hardware, software, and humans are inevitable. Our approach, denoted recovery-oriented computing (ROC), recognizes the inevitability of unanticipated failure and thus emphasizes recovery and repair rather than simple fault-tolerance. We define the properties that a ROC system must provide, and briefly consider how they might be achieved.
