(Re)Designing Data-Centric Data Centers
IEEE MicroPublished 1 January 2012
Parthasarathy Ranganathan, Jichuan Chang
Citations10
SJR quartileQ1
SJR score0.94
SNIP1.40
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A grand challenge facing computing today is to design future systems to efficiently manage and harvest useful insights from the growing amount of information through novel redesign at both the hardware and software levels.
Abstract
A grand challenge facing computing today is to design future systems to efficiently manage and harvest useful insights from the growing amount of information. Emerging technology inflections provide a unique opportunity to address this challenge through novel redesign at both the hardware and software levels.
Keywords
Computer Science
Dark silicon and the end of multicore scaling
1,511 Citations2011Hadi Esmaeilzadeh, Emily Blem +3 more
The study shows that regardless of chip organization and topology, multicore scaling is power limited to a degree not widely appreciated by the computing community.
Scalable high performance main memory system using phase-change memory technology
1,297 Citations2009Moinuddin K. Qureshi, Vijayalakshmi Srinivasan +1 more
This paper analyzes a PCM-based hybrid main memory system using an architecture level model of PCM and proposes simple organizational and management solutions of the hybrid memory that reduces the write traffic to PCM, boosting its lifetime from 3 years to 9.7 years.
Better I/O through byte-addressable, persistent memory
828 Citations2009Jeremy Condit, Edmund B. Nightingale +5 more
A file system and a hardware architecture that are designed around the properties of persistent, byteaddressable memory, which provides strong reliability guarantees and offers better performance than traditional file systems, even when both are run on top of byte-addressable, persistent memory.
Mnemosyne
720 Citations2011Haris Volos, Andres Jaan Tack +1 more
In tests emulating the performance characteristics of forthcoming SCMs, Mnemosyne can persist data as fast as 3 microseconds and can be up to 1400% faster than alternative persistence strategies, such as Berkeley DB or Boost serialization, that are designed for disks.
IEEE MicroA case for intelligent RAM
668 Citations1997David A. Patterson, Thomas E. Anderson +6 more
The state of microprocessors and DRAMs today is reviewed, some of the opportunities and challenges for IRAMs are explored, and performance and energy efficiency of three IRAM designs are estimated.
NV-Heaps
658 Citations2011Joel Coburn, Adrian M. Caulfield +5 more
A lightweight, high-performance persistent object system called NV-heaps is implemented that provides transactional semantics while preventing these errors and providing a model for persistence that is easy to use and reason about.
FAWN
563 Citations2009David G. Andersen, Jason Franklin +4 more
This paper presents a new cluster architecture for low-power data-intensive computing that couples low-power embedded CPUs to small amounts of local flash storage, and balances computation and I/O capabilities to enable efficient, massively parallel access to data.
IEEE MicroDark Silicon and the End of Multicore Scaling
489 Citations2012Hadi Esmaeilzadeh, Emily Blem +3 more
A comprehensive study that projects the speedup potential of future multicores and examines the underutilization of integration capacity-dark silicon-is timely and crucial.
IEEE MicroPhase-Change Technology and the Future of Main Memory
404 Citations2010Benjamin C. Lee, Ping Zhou +6 more
This article discusses how to mitigate limitations through buffer sizing, row caching, write reduction, and wear leveling, to make PCM a viable dream alternative for scalable main memories.
Relaxing non-volatility for fast and energy-efficient STT-RAM caches
394 Citations2011Clinton W. Smullen, Vidyabhushan Mohan +3 more
It is found that a pure STT-RAM cache hierarchy provides the best energy efficiency, though a hybrid design of SRAM-based L1 caches with reduced-retention STt-RAM L2 and L3 caches eliminates performance loss while still reducing the energy-delay product by more than 70%.
File and Storage TechnologiesConsistent and durable data structures for non-volatile byte-addressable memory
365 Citations2011Shivaram Venkataraman, Niraj H. Tolia +2 more
This paper presents Consistent and Durable Data Structures (CDDSs), a single-level data store that, on current hardware, allows programmers to safely exploit the low-latency and non-volatile aspects of new memory technologies.
Disaggregated memory for expansion and sharing in blade servers
364 Citations2009Kevin Lim, Jichuan Chang +4 more
It is demonstrated that memory disaggregation can provide substantial performance benefits (on average 10X) in memory constrained environments, while the sharing enabled by the solutions can improve performance-per-dollar by up to 57% when optimizing memory provisioning across multiple servers.
Hybrid cache architecture with disparate memory technologies
333 Citations2009Xiaoxia Wu, Jian Li +4 more
This paper discusses and evaluates two types of hybrid cache architectures: inter cache Level HCA (LHCA), in which the levels in a cache hierarchy can be made of disparate memory technologies; and intra cache level or cache Region based H CA (RHCA), where a single level of cache can be partitioned into multiple regions, each of a different memory technology.
IEEE MicroToward Dark Silicon in Servers
311 Citations2011Nikos Hardavellas, Michael Ferdman +2 more
Server chips will not scale beyond a few tens to low hundreds of cores, and an increasing fraction of the chip in future technologies will be dark silicon that the authors cannot afford to power.
Conservation cores
273 Citations2010Ganesh Venkatesh, Jack Sampson +6 more
A toolchain for automatically synthesizing c-cores from application source code is presented and it is demonstrated that they can significantly reduce energy and energy-delay for a wide range of applications, and patching can extend the useful lifetime of individual c-Cores to match that of conventional processors.
Active Storage for Large-Scale Data Mining and Multimedia
273 Citations1998Erik Riedel, Garth A. Gibson +1 more
Gordon
248 Citations2009Adrian M. Caulfield, Laura M. Grupp +1 more
The paper presents an exhaustive analysis of the design space of Gordon systems, focusing on the trade-offs between power, energy, and performance that Gordon must make, and describes a novel flash translation layer tailored to data intensive workloads and large flash storage arrays.
Rethinking Database Algorithms for Phase Change Memory
225 Citations2011Shimin Chen, Phillip B. Gibbons +1 more
Improved algorithms that reduce both execution time and energy on PCM while increasing write endurance are presented, and current approaches for common database algorithms such as B + -trees and Hash Joins are suboptimal for PCM.
Dynamically Specialized Datapaths for energy efficient computing
204 Citations2011Venkatraman Govindaraju, Chen-Han Ho +1 more
D Dynamically Specialized Datapaths are proposed to improve the energy efficiency of general purpose programmable processors and show that in most cases two DySER blocks can achieve the same performance as having a specialized hardware module for each path-tree.
ACM SIGARCH Computer Architecture NewsUnderstanding and Designing New Server Architectures for Emerging Warehouse-Computing Environments
201 Citations2008Kevin Lim, Parthasarathy Ranganathan +4 more
A new solution that incorporates volume non-server-class components in novel packaging solutions, with memory sharing and flash-based disk caching, has promise, with a 2X improvement on average in performance-per-dollar for the benchmark suite.
The Rio file cache
195 Citations1996Peter M. Chen, Wee Teck Ng +4 more
The goal of the Rio (RAM I/O) file cache is to make ordinary main memory safe for persistent storage by enabling memory to survive operating system crashes by protecting memory during a crash and restoring it during a reboot.
Understanding and Designing New Server Architectures for Emerging Warehouse-Computing Environments
171 Citations2008Kevin Lim, Parthasarathy Ranganathan +4 more
Communications of the ACMThe case for RAMCloud
159 Citations2011John K. Ousterhout, Parag Agrawal +12 more
A Case for Intelligent RAM: IRAM
155 Citations1997David A. Patterson, Thomas E. Anderson +6 more
This paper reviews the state of microprocessors and DRAMs today, explores some of the opportunities and challenges for IRAMs, and finally estimates performance and energy effi- ciency of three IRAM designs.
ComputerFrom Microprocessors to Nanostores: Rethinking Data-Centric Systems
103 Citations2011Parthasarathy Ranganathan
The confluence of emerging technologies and new data-centric workloads offers a unique opportunity to rethink traditional system architectures and memory hierarchies in future designs.
ComputerFrom Synapses to Circuitry: Using Memristive Memory to Explore the Electronic Brain
96 Citations2011Greg Snider, Rick Amerson +10 more
In a synchronous digital platform for building large cognitive models, memristive nanodevices form dense, resistive memories that can be placed close to conventional processing circuitry and through adaptive transformations, the devices can interact with the world in real time.
Warehouse-Scale Computing: Entering the Teenage Decade
82 Citations2011Luiz André Barroso
IEEE Journal of Solid-State CircuitsA parallel processing chip with embedded DRAM macros
25 Citations1996T. Sunaga, H. Miyatake +3 more
A combined DRAM and logic chip has been developed for massively parallel processing (MPP) applications that delivers 50-MIPS of performance at 2.7 W and contains eight 16-b CPUs and some broadcast logic circuits.
Flash in a DBMS: Where and How?
25 Citations2010Manos Athanassoulis, Anastasia Ailamaki +3 more
Techniques for making effective use of flash in three contexts are described: as a log device for transaction processing on memory-resident data, as the main data store for transactionprocessing, and as an update cache for HDD- resident data warehouses.
Digital Access to Scholarship at Harvard (DASH) (Harvard University)Multicore OSes: looking forward from 1991, er, 2011
8 Citations2011David A. Holland, Margo Seltzer
What adopting the lightweight messages and channels programming model entails, the architecture of an OS based on this model, and a few likely implementation challenges are discussed.
Operating systems must support GPU abstractions
7 Citations2011Christopher J. Rossbach, Jon Currey +1 more
It is argued that lack of OS support for GPU abstractions fundamentally limits the usability of GPUs in many application domains and proposed new kernel abstractions to support GPUs and other accelerator devices as first class computing resources are proposed.
