Recovery-oriented computing: building multitier dependability
ComputerPublished 1 November 2004Open access
George Candea, Aaron Brown, Armando Fox, David A. Patterson
Citations136
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Building systems to recover fast may be more productive than aiming for systems that never fail, and the authors advocate multiple lines of defense in managing failures.
Abstract
Building systems to recover fast may be more productive than aiming for systems that never fail. Because recovery is not immune to failure either, the authors advocate multiple lines of defense in managing failures
Keywords
Computer Science
ACM Computing SurveysA survey of rollback-recovery protocols in message-passing systems
1,795 Citations2002E. N. Mootaz Elnozahy, Lorenzo Alvisi +2 more
This survey covers rollback-recovery techniques that do not require special language constructs and distinguishes between checkpoint-based and log-based protocols, which rely solely on checkpointing for system state restoration.
IEEE Transactions on ComputersOn Evaluating the Performability of Degradable Computing Systems
714 Citations1980Meyer
A hierarchical modeling scheme is used to formulate the capability function and capability is used, in turn, to evaluate performability, and techniques are illustrated for a specific application: the performability evaluation of an aircraft computer in the environment of an air transport mission.
arXiv (Cornell University)Microreboot -- A Technique for Cheap Recovery
366 Citations2004George Candea, Shinichi Kawamoto +3 more
Reliable Computer Systems
298 Citations1998Daniel P. Siewiorek, Robert S. Swarz
Find the secret to improve the quality of life by reading this reliable computer systems design and evaluation third edition.
Performance and scalability of EJB applications
230 Citations2002Emmanuel Cecchet, Julie Marguerite +1 more
This work investigates the combined effect of application implementation method, container design, and efficiency of communication layers on the performance scalability of J2EE application servers by detailed measurement and profiling of an auction site server.
Crash-only software
133 Citations2003George Candea, Armando Fox
This paper presents ideas on how to build such crash-only Internet services, taking successful techniques to their logical extreme, and shows that it can lead to more reliable, predictable code and faster, more effective recovery.
Undo for operators: building an undoable e-mail store
101 Citations2003Aaron Brown, David A. Patterson
Performance and scalability of EJB applications
66 Citations2002Emmanuel Cecchet, Julie Marguerite +1 more
