The Datacenter as a Computer: An Introduction to the Design of Warehouse-Scale Machines, Second edition
Synthesis lectures on computer architecturePublished 31 July 2013Open access
Luiz André Barroso, Jimmy Clidaras, Urs Hölzle
Citations491
SJR quartileQ4
SJR score0.18
SNIP0.48
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The architecture of WSCs is described, the main factors influencing their design, operation, and cost structure, and the characteristics of their software base are described.
Abstract
This book describes warehouse-scale computers (WSCs), the computing platforms that power cloud computing and all the great web services we use every day. It discusses how these new systems treat the d
Keywords
Computer Science
Communications of the ACMMapReduce
18,538 Citations2008Jay B. Dean, Sanjay Ghemawat
This presentation explains how the underlying runtime system automatically parallelizes the computation across large-scale clusters of machines, handles machine failures, and schedules inter-machine communication to make efficient use of the network and disks.
Computer Networks and ISDN SystemsThe anatomy of a large-scale hypertextual Web search engine
15,828 Citations1998Sergey Brin, Lawrence M. Page
This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.
Communications of the ACMA view of cloud computing
8,997 Citations2010Michael Armbrust, Armando Fox +9 more
The clouds are clearing the clouds away from the true potential and obstacles posed by this computing capability.
ACM SIGOPS Operating Systems ReviewXen and the art of virtualization
5,948 Citations2003Paul Barham, Boris Dragovic +7 more
ACM SIGOPS Operating Systems ReviewThe Google file system
5,003 Citations2003Sanjay Ghemawat, Howard Gobioff +1 more
IEEE Journal of Solid-State CircuitsDesign of ion-implanted MOSFET's with very small physical dimensions
3,477 Citations1974R.H. Dennard, F.H. Gaensslen +4 more
This paper considers the design, fabrication, and characterization of very small Mosfet switching devices suitable for digital integrated circuits, using dimensions of the order of 1 /spl mu/.
Dynamo
3,454 Citations2007Giuseppe DeCandia, Deniz Hastorun +7 more
D Dynamo is presented, a highly available key-value storage system that some of Amazon's core services use to provide an "always-on" experience and makes extensive use of object versioning and application-assisted conflict resolution in a manner that provides a novel interface for developers to use.
ACM Transactions on Computer SystemsBigtable
3,432 Citations2008Fay W. Chang, Jay B. Dean +7 more
The simple data model provided by Bigtable is described, which gives clients dynamic control over data layout and format, and the design and implementation of Bigtable are described.
ComputerThe Case for Energy-Proportional Computing
2,479 Citations2007Luiz André Barroso, Urs Hölzle
Energy-proportional designs would enable large energy savings in servers, potentially doubling their efficiency in real-life use, particularly the memory and disk subsystems.
Dryad
2,460 Citations2007Michael Isard, Mihai Budiu +3 more
The Dryad execution engine handles all the difficult problems of creating a large distributed, concurrent application: scheduling the use of computers and their CPUs, recovering from communication or computer failures, and transporting data between vertices.
Bigtable: a distributed storage system for structured data
1,955 Citations2006Fay W. Chang, Jay B. Dean +7 more
Power provisioning for a warehouse-sized computer
1,801 Citations2007Xiaobo Fan, Wolf-Dietrich Weber +1 more
This paper presents the aggregate power usage characteristics of large collections of servers for different classes of applications over a period of approximately six months, and uses the modelling framework to estimate the potential of power management schemes to reduce peak power and energy usage.
B4
1,779 Citations2013Sushant Jain, Alok Kumar +12 more
This work presents the design, implementation, and evaluation of B4, a private WAN connecting Google's data centers across the planet, using OpenFlow to control relatively simple switches built from merchant silicon.
Communications of the ACMThe tail at scale
1,740 Citations2013Jay B. Dean, Luiz André Barroso
Software techniques that tolerate latency variability are vital to building responsive large-scale Web services.
UC BerkeleyMesos: a platform for fine-grained resource sharing in the data center
1,593 Citations2011Benjamin Hindman, Andy Konwinski +6 more
The results show that Mesos can achieve near-optimal data locality when sharing the cluster among diverse frameworks, can scale to 50,000 (emulated) nodes, and is resilient to failures.
The Google file system
1,366 Citations2003Sanjay Ghemawat, Howard Gobioff +1 more
This paper presents file system interface extensions designed to support distributed applications, discusses many aspects of the design, and reports measurements from both micro-benchmarks and real world use.
Managing energy and server resources in hosting centers
1,285 Citations2001Jeffrey S. Chase, Darrell C. Anderson +3 more
Experimental results from a prototype confirm that the system adapts to offered load and resource availability, and can reduce server energy usage by 29% or more for a typical Web workload.
Communications of the ACMParallel database systems
1,281 Citations1992David J. DeWitt, Jim Gray
Over the last decade 'Eradata, Tandem, and a host of startup companies have successfully developed and marketed highly parallel machines that refutes a 1983 paper predicting the demise of database machines.
ACM SIGOPS Operating Systems ReviewDynamo
1,109 Citations2007Giuseppe DeCandia, Deniz Hastorun +7 more
IEEE MicroWeb search for a planet: the google cluster architecture
1,026 Citations2003Luiz André Barroso, Jay B. Dean +1 more
Googless architecture features clusters of more than 15,000 commodity-class PCs with fault tolerant software that achieves superior performance at a fraction of the cost of a system built from fewer, but more expensive, high-end servers.
Communications of the ACMEventually consistent
1,005 Citations2008Werner Vogels
Building reliable distributed systems at a worldwide scale demands trade-offs between consistency and availability.
Environmental Research LettersWorldwide electricity used in data centers
961 Citations2008Jonathan Koomey
This study estimates historical electricity use by data centers worldwide and regionally on the basis of more detailed data than were available for previous assessments, including electricity used by servers, data center communications, and storage equipment.
PortLand
936 Citations2009Radhika Niranjan Mysore, Andreas Pamboris +6 more
Through the design and implementation of PortLand, a scalable, fault tolerant layer 2 routing and forwarding protocol for data center environments, it is shown that PortLand holds promise for supporting a ``plug-and-play" large-scale, data center network.
PowerNap
900 Citations2009David Meisner, Brian T. Gold +1 more
The PowerNap concept, an energy-conservation approach where the entire system transitions rapidly between a high-performance active state and a near-zero-power idle state in response to instantaneous load, is proposed and the Redundant Array for Inexpensive Load Sharing (RAILS) is introduced.
Operating Systems Design and ImplementationThe Chubby lock service for loosely-coupled distributed systems
899 Citations2006Mike Burrows
The paper describes the initial design and expected use, compares it with actual use, and explains how the design had to be modified to accommodate the differences.
Spanner: Google's globally-distributed database
827 Citations2012James C. Corbett, Jay B. Dean +24 more
Disk failures in the real world: what does an MTTF of 1,000,000 hours mean to you?
773 Citations2007Bianca Schroeder, Garth A. Gibson
This paper presents and analyzes field-gathered disk replacement data from a number of large production systems, including high-performance computing sites and internet services sites, finding evidence, based on records of disk replacements in the field, that failure rate is not constant with age, and that, rather than a significant infant mortality effect, the authors see a significant early onset of wear-out degradation.
Proceedings of the VLDB EndowmentDremel
709 Citations2010Sergey Melnik, Andrey Gubarev +5 more
This paper describes the architecture and implementation of Dremel, and explains how it complements MapReduce-based computing, and presents a novel columnar storage representation for nested records.
Failure trends in a large disk drive population
709 Citations2007Eduardo Pinheiro, Wolf-Dietrich Weber +1 more
It is found that temperature and activity levels were much less correlated with drive failures than previously reported, and models based on SMART parameters alone are unlikely to be useful for predicting individual drive failures.
Megastore: Providing Scalable, Highly Available Storage for Interactive Services
654 Citations2011Jason D. Baker, Chris T. Bond +8 more
Megastore provides fully serializable ACID semantics within ne-grained partitions of data, which allows us to synchronously replicate each write across a wide area network with reasonable latency and support seamless failover between datacenters.
Energy-aware server provisioning and load dispatching for connection-intensive internet services
645 Citations2008Chen Gong, Wenbo He +5 more
This paper characterize unique properties, performance, and power models of connection servers, based on a real data trace collected from the deployed Windows Live Messenger, and shows that these algorithms can save a significant amount of energy without sacrificing user experiences.
Omega
638 Citations2013Malte Schwarzkopf, Andy Konwinski +2 more
This work presents a novel approach to address increasing scale and the need for rapid response to changing requirements using parallelism, shared state, and lock-free optimistic concurrency control to address monolithic cluster scheduler architectures.
USENIX Annual Technical ConferenceMaking scheduling cool: temperature-aware workload placement in data centers
630 Citations2005Justin Moore, Jeff Chase +2 more
This paper examines a theoretic thermodynamic formulation that uses information about steady state hot spots and cold spots in the data center and develops real-world scheduling algorithms, and develops an alternate approach to address the problem of heat management through temperature-aware workload placement.
No "power" struggles
621 Citations2008Ramya Raghavendra, Parthasarathy Ranganathan +3 more
This paper proposes and validate a power management solution that coordinates different individual approaches and performs a detailed quantitative sensitivity analysis to draw conclusions about the impact of different architectures, implementations, workloads, and system design choices.
Scientific ProgrammingInterpreting the Data: Parallel Analysis with Sawzall
616 Citations2005Rob Pike, Sean Dorward +2 more
The design -- including the separation into two phases, the form of the programming language, and the properties of the aggregators -- exploits the parallelism inherent in having data and computation distributed across many machines.
UC BerkeleyWhy Do Internet Services Fail, and What Can Be Done About It?
603 Citations2002David Oppenheimer, Archana Ganapathi +1 more
It is found that operator errors are their primary cause, operator error is the most difficult type of failure to mask, service front-ends are responsible for more problems than service back-ends but fewer minutes of unavailability, and that online testing and more thoroughly exposing and detecting component failures could reduce system failure rates for at least one service.
Managing server energy and operational costs in hosting centers
560 Citations2005Yiyu Chen, Amitayu Das +4 more
This paper proposes three new online solution strategies based on steady state queuing analysis, feedback control theory, and a hybrid mechanism borrowing ideas from these two that are more adaptive to workload behavior when performing server provisioning and speed control than earlier heuristics towards minimizing operational costs while meeting the SLAs.
Operating Systems Design and ImplementationAvailability in globally distributed storage systems
556 Citations2010Daniel Alexander Ford, François Labelle +6 more
This work characterize the availability properties of cloud storage systems based on an extensive one year study of Google's main storage infrastructure and presents statistical models that enable further insight into the impact of multiple design choices, such as data placement and replication strategies.
DRAM errors in the wild
555 Citations2009Bianca Schroeder, Eduardo Pinheiro +1 more
Measurements of memory errors in a large fleet of commodity servers over a period of 2.5 years provide strong evidence that memory errors are dominated by hard errors, rather than soft errors, which previous work suspects to be the dominant error mode.
IEEE Annals of the History of ComputingImplications of Historical Trends in the Electrical Efficiency of Computing
479 Citations2010Jonathan Koomey, Stephen Berard +2 more
SibFU Digital Repository (Siberian Federal University)A rapid in vitro protocol for callus production in Piper aduncum L. propagation of East Kalimantan supplemented with gradient sucrose solutions
475 Citations2018Sudrajat Sudrajat
Dapper, a Large-Scale Distributed Systems Tracing Infrastructure
470 Citations2010Benjamin H. Sigelman, Luiz André Barroso +6 more
The design of Dapper is introduced, Google’s production distributed systems tracing infrastructure is described, and how its design goals of low overhead, application-level transparency, and ubiquitous deployment on a very large scale system were met are described.
ACM SIGOPS Operating Systems ReviewThe case for RAMClouds
451 Citations2010John K. Ousterhout, Parag Agrawal +11 more
This paper argues for a new approach to datacenter storage called RAMCloud, where information is kept entirely in DRAM and large-scale systems are created by aggregating the main memories of thousands of commodity servers.
Energy proportional datacenter networks
414 Citations2010Dennis Abts, Michael R. Marty +3 more
It is demonstrated that energy proportional datacenter communication is indeed possible and that there is a significant power advantage to having independent control of each unidirectional channel comprising a network link.
X-trace: a pervasive network tracing framework
408 Citations2007Rodrigo Fonseca, George Porter +3 more
This paper proposes X-Trace, a tracing framework that provides such a comprehensive view of service behavior for systems that adopt it, and discusses how it works in three deployed scenarios: DNS resolution, a three-tiered photo-hosting website, and a service accessed through an overlay network.
Power management of online data-intensive services
405 Citations2011David Meisner, Christopher Sadler +3 more
This work evaluates the applicability of active and idle low-power modes to reduce the power consumed by the primary server components (processor, memory, and disk), while maintaining tight response time constraints, particularly on 95th-percentile latency.
Towards highly reliable enterprise network services via inference of multi-level dependencies
373 Citations2007Paramvir Bahl, Ranveer Chandra +4 more
An Inference Graph model is introduced, which is well-adapted to user-perceptible problems rooted in conditions giving rise to both partial service degradation and hard faults, and takes into account multi-level structure, which leads to a 30% improvement in fault localization, as compared to two-level approaches.
Recovery Oriented Computing (ROC): Motivation, Definition, Techniques, and Case Studies
371 Citations2002David A. Patterson, Aaron Brown +13 more
Recovery Oriented Computing (ROC) takes the perspective that hardware faults, software bugs, and operator errors are facts to be coped with, not problems to be solved, and thus offers higher availability.
Journal of Physics Conference SeriesUnderstanding failures in petascale computers
361 Citations2007Bianca Schroeder, Garth A. Gibson
This paper reviews sources of failure information for compute clusters and storage systems, projects failure rates and the corresponding decrease in application effectiveness, and discusses coping strategies such as application-level checkpoint compression and system level process-pairs fault-tolerance for supercomputing.
IEEE Transactions on ReliabilityA census of Tandem system availability between 1985 and 1990
324 Citations1990Jodi Gray
A census of customer outages reported to Tandem indicates that software is now the major source of reported outages, followed by system operations, a dramatic shift from the statistics for 1985.
IEEE Communications MagazineIEEE 802.3az: the road to energy efficient ethernet
315 Citations2010Ken Christensen, Pedro Reviriego +4 more
The development of the EEE standard and how energy savings resulting from the adoption of EEE may exceed $400 million per year in the U.S. alone are described and results show that packet coalescing can significantly improve energy efficiency while keeping absolute packet delays to tolerable bounds are presented.
An analysis of latent sector errors in disk drives
295 Citations2007Lakshmi N. Bairavasundaram, Garth R. Goodson +2 more
This is the first study of such large scale the sample size is at least an order of magnitude larger than previously published studies and the first one to focus specifically on latent sector errors and their implications on the design and reliability of storage systems.
Conserving disk energy in network servers
285 Citations2003Enrique V. Carrera, Eduardo Pinheiro +1 more
The results for Web and proxy servers show that the fourth approach can provide energy savings of up to 23%, in comparison to conventional servers, without any degradation in server performance.
IEEE Internet ComputingLessons from giant-scale services
248 Citations2001Kyle Palos
The article looks at the basic model for such services, focusing on the key real-world challenges they face (high availability, evolution, and growth), and developing some principles for attacking these problems.
Cosmic rays don't strike twice
243 Citations2012Andy A. Hwang, Ioan Stefanovici +1 more
It is found that a large fraction of DRAM errors in the field can be attributed to hard errors and a detailed analytical study of their characteristics is provided, including both hard and soft errors.
Best Practices for Data Centers: Lessons Learned from Benchmarking 22 Data Centers
218 Citations2006Steve Greenberg, Evan Mills +2 more
Energy storage in datacenters
207 Citations2012Di Wang, Chuangang Ren +3 more
This paper intends to fill a critical void in the extensive design space involving multiple ESD technology provisioning and placement options by presenting a theoretical framework for capturing important characteristics of different ESD technologies, the trade-off of placing them at different levels of the power hierarchy, and quantifying the resulting cost-benefit trade-offs as a function of workload properties.
ACM SIGOPS Operating Systems ReviewAutopilot
203 Citations2007Michael Isard
The first version of Autopilot is described, the automatic data center management infrastructure developed within Microsoft over the last few years, responsible for automating software provisioning and deployment; system monitoring; and carrying out repair actions to deal with faulty software and hardware.
Mercury and freon
202 Citations2006Taliver Heath, Ana Paula Centeno +4 more
Mercury, a software suite that accurately emulating temperatures based on simple layout, hardware, and componentutilization data, is introduced that runs the entire software stack natively, enables repeatable experiments, and allows the study of thermal emergencies without harming hardware reliability.
ACM SIGARCH Computer Architecture NewsUnderstanding and Designing New Server Architectures for Emerging Warehouse-Computing Environments
201 Citations2008Kevin Lim, Parthasarathy Ranganathan +4 more
A new solution that incorporates volume non-server-class components in novel packaging solutions, with memory sharing and flash-based disk caching, has promise, with a 2X improvement on average in performance-per-dollar for the benchmark suite.
Temperature management in data centers
197 Citations2012Nosayba El-Sayed, Ioan Stefanovici +3 more
A multi-faceted study of temperature management in data centers using a large collection of field data from different production environments to study the impact of temperature on hardware reliability, including the reliability of the storage subsystem, the memory subsystem and server reliability as a whole.
IEEE MicroGoogle-Wide Profiling: A Continuous Profiling Infrastructure for Data Centers
191 Citations2010Gang Ren, Eric Tune +4 more
Google-Wide Profiling (GWP), a continuous profiling infrastructure for data centers, provides performance insights for cloud applications and introduces novel applications of its profiles, such as application-platform affinity measurements and identification of platform-specific, microarchitectural peculiarities.
ACM SIGMETRICS Performance Evaluation ReviewManaging server energy and operational costs in hosting centers
186 Citations2005Yiyu Chen, Amitayu Das +4 more
The growing cost of tuning and managing computer systems is leading to out-sourcing of commercial services to hosting centers, which provision thousands of dense servers within a relatively small area.
ACM SIGARCH Computer Architecture NewsManaging distributed ups energy for effective power capping in data centers
180 Citations2012Vasileios Kontorinis, Liuyi Eric Zhang +6 more
This work presents an architecture for distributed per-server UPSs that stores energy during low activity periods and uses this energy during power spikes, which leverages the distributed nature of the UPS batteries and develops policies that prolong the duration of their usage.
Thin servers with smart pipes
176 Citations2013Kevin Lim, David Meisner +3 more
This work argues for an alternate architecture---Thin Servers with Smart Pipes (TSSP)---for cost-effective high-performance memcached deployment, and demonstrates the potential benefits of the TSSP architecture through an FPGA prototyping platform, and shows the potential for a 6X-16X power-performance improvement over conventional server baselines.
Thermal considerations in cooling large scale high compute density data centers
173 Citations2003Chandrakant Patel, Ratnesh Sharma +2 more
Understanding and Designing New Server Architectures for Emerging Warehouse-Computing Environments
171 Citations2008Kevin Lim, Parthasarathy Ranganathan +4 more
Journal of Applied PhysiologyImproved muscular efficiency displayed as Tour de France champion matures
169 Citations2005Edward F. Coyle
This case describes the physiological maturation from ages 21 to 28 yr of the bicyclist who has now become the six-time consecutive Grand Champion of the Tour de France, at ages 27-32 yr, and it appears that an 8% improvement in muscular efficiency and thus power production when cycling at a given oxygen uptake (Vo(2)) is the characteristic that improved most.
ACM SIGARCH Computer Architecture NewsEnergy proportional datacenter networks
167 Citations2010Dennis Abts, Michael R. Marty +3 more
Communications of the ACMThe case for RAMCloud
159 Citations2011John K. Ousterhout, Parag Agrawal +12 more
The Future of Computing Performance: Game Over or Next Level?
147 Citations2011Samuel H. Fuller, Lynette I. Millett
The factors that have led to the future limitations on growth for single processors that are based on complementary metal oxide semiconductor (CMOS) technology are described and challenges inherent in parallel computing and architecture are explored, including ever-increasing power consumption and the escalated requirements for heat dissipation.
Failure data analysis of a LAN of Windows NT based computers
147 Citations2003M. Kalyanakrishnam, Zbigniew Kalbarczyk +1 more
The key observations from this study are: most of the problems that lead to reboots are software related, and the average availability is not a good measure to characterize this type of network service.
IEEE Transactions on Computer-Aided Design of Integrated Circuits and SystemsEnergy-Efficient Datacenters
143 Citations2012Massoud Pedram
The goal of this paper is to provide an introduction to resource provisioning and power or thermal management problems in datacenters, and to review strategies that maximize the datacenter energy efficiency subject to peak or total power consumption and thermal constraints, while meeting stipulated service level agreements in terms of task throughput and/or response time.
IEEE/ACM Transactions on NetworkingEnd-to-end WAN service availability
141 Citations2003Michael Dahlin, Bharat Chandra +2 more
It is found that caching alone is seldom effective at insulating services from failures but that the combination of mobile extension code and prefetching can improve average unavailability by as much as an order of magnitude for classes of service whose semantics support disconnected operation.
IEEE MicroScale-Out Networking in the Data Center
131 Citations2010Amin Vahdat, Mohammad Al-Fares +4 more
Through the UCSD Triton network architecture, the authors explore issues in managing the network as a single plug-and-play virtualizable fabric scalable to hundreds of thousands of ports and petabits per second of aggregate bandwidth.
WAP5
125 Citations2006Patrick Reynolds, Janet L. Wiener +3 more
This paper presents a new algorithm for reconstructing application structure in both local- and wide-area distributed systems, an infrastructure for gathering application traces in PlanetLab, and an experiences tracing and analyzing three systems.
Disk-locality in datacenter computing considered irrelevant
124 Citations2011Ganesh Ananthanarayanan, Ali Ghodsi +2 more
Data center computing is becoming pervasive in many organizations, and a considerable work has been done to improve the efficiency of computing frameworks such as MapReduce, Hadoop and Dryad.
ACM SIGCOMM Computer Communication ReviewTowards highly reliable enterprise network services via inference of multi-level dependencies
121 Citations2007Paramvir Bahl, Ranveer Chandra +4 more
Boosting Data Center Performance Through Non-Uniform Power Allocation
119 Citations2005Mark E. Femal, Vincent W. Freeh
A non-uniform power allocation scheme increases throughput by over 16% versus a uniform power allocation mechanism, which is useful for those data centers that cannot expand the number of power circuits or seek effective usage of their available power budget due to unplanned power fluctuations.
ParasitologyNeedle in a haystack: involvement of the copepod <i>Paracartia grani</i> in the life-cycle of the oyster pathogen <i>Marteilia refringens</i>
109 Citations2002Corinne Audemard, Frédérique Le Roux +8 more
It is shown that the copepod Paracartia(Acartia) grani is a host of M. refringens and the presence of the parasite in the ovarian tissues was demonstrated using in situ hybridization, providing evidence that P. grani can be infected from infected flat oysters.
ACM SIGMETRICS Performance Evaluation ReviewEnergy storage in datacenters
106 Citations2012Di Wang, Chuangang Ren +3 more
Energy storage - in the form of UPS units - in a datacenter has been primarily used to fail-over to diesel generators upon power outages, but recent interest in using these Energy Storag...
On designing and deploying internet-scale services
104 Citations2007James R. Hamilton
This paper summarizes the best practices accumulated over many years in scaling some of the largest services at MSN and Windows Live.
ComputerModels and Metrics to Enable Energy-Efficiency Optimizations
88 Citations2007Suzanne Rivoire, Mehul A. Shah +3 more
Researchers and system designers need benchmarks that characterize energy efficiency to evaluate systems and identify promising new technologies and to predict the effects of new designs and configurations, and need accurate methods of modeling power consumption.
Warehouse-Scale Computing: Entering the Teenage Decade
82 Citations2011Luiz André Barroso
Wimpy node clusters
77 Citations2010Willis Lang, Jignesh M. Patel +1 more
Results show that in most cases, computationally complex queries exhibit disproportionate scaleup characteristics which potentially makes scale-out with low-end nodes an expensive and lower performance solution.
FigsharePower Capping Via Forced Idleness
76 Citations2018Anshul Gandhi, Mor Harchol‐Balter +3 more
A novel power capping technique that achieves higher effective server frequency for a given power constraint than existing techniques is introduced, and it is argued how IdleCap applies to next-generation servers using DVFS and advanced idle states.
Communications of the ACMA guided tour of data-center networking
74 Citations2012Dennis Abts, Bob Felderman
A good user experience depends on predictable performance within the data-center network, so it is important to have a good understanding of how the network works and how to improve the user experience.
QueueGFS: Evolution on Fast-forward
71 Citations2009Marshall Kirk McKusick, Sean Quinlan
During the early stages of development at Google, the initial thinking did not include plans for building a new file system, but while work was still being done on one of the earliest versions of the company’s crawl and indexing system, however, it became quite clear to the core engineers that they really had no other choice, and GFS (Google File System) was born.
Server class disk drives: how reliable are they?
67 Citations2004J.G. Elerath, Sandeep Shah
This paper further elaborates on these four causes of variability and explains how each is responsible for a possible gap between expected and measured drive reliability.
Brawny cores still beat wimpy cores, most of the time
60 Citations2010Urs Hölzle
Multicore architectures are great for warehouse-scale systems because they provide ample parallelism in the request stream as well as data parallelism for search or analysis over petabyte data sets, but wimpy-core systems can require applications to be explicitly parallelized or otherwise optimized for acceptable performance.
Lecture notes in computer scienceSafe Overprovisioning: Using Power Limits to Increase Aggregate Throughput
48 Citations2005Mark E. Femal, Vincent W. Freeh
Host-based and network-centric models are proposed to monitor and coordinate the distribution of power with the fundamental goal of increasing throughput and initial results with a synthetic benchmark indicate throughput increases of nearly 6% from a staticly assigned, power managed environment and over 30% from an unmanaged environment.
Communications of the ACMFAWN
48 Citations2011David G. Andersen, Jason Franklin +4 more
The design centers around purely log-structured datastores that provide the basis for high performance on flash storage, as well as for replication and consistency obtained using chain replication on a consistent hashing ring.
ACM SIGARCH Computer Architecture NewsPiranha
43 Citations2000Luiz André Barroso, Kourosh Gharachorloo +7 more
High Performance Datacenter Networks: Architectures, Algorithms, & Opportunities
41 Citations2011Dennis Abts, John Kim
ACM SIGARCH Computer Architecture NewsCosmic rays don't strike twice
37 Citations2012Andy A. Hwang, Ioan Stefanovici +1 more
…
