Using destination-set prediction to improve the latency/bandwidth tradeoff in shared-memory multiprocessors
Published 1 January 2003
Milo M. K. Martin, Pacia J. Harper, Daniel J. Sorin, Mark D. Hill, David A. Wood
Citations124
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. In most cases, these works may not be reposted without the explicit permission of the copyright holder.
Keywords
Computer Science
ComputerSimics: A full system simulation platform
2,115 Citations2002Peter Magnusson, Magnus Christensson +7 more
Simics is a platform for full system simulation that can run actual firmware and completely unmodified kernel and driver code, and it provides both functional accuracy for running commercial workloads and sufficient timing accuracy to interface to detailed hardware models.
The SPLASH-2 programs: characterization and methodological considerations
1,599 Citations2002Sandi Woo, Moriyoshi Ohara +3 more
The SGI Origin
779 Citations1997James Laudon, Daniel Lenoski
The motivation for building the Origin 2000 is discussed and the architecture and implementation of the multiprocessor is described, and performance results are presented for the NAS Parallel Benchmarks V2.2 and the SPLASH2 applications.
AlgorithmicaCompetitive snoopy caching
612 Citations1988Anna R. Karlin, Mark S. Manasse +2 more
This work presents new on-line algorithms to be used by the caches of snoopy cache multiprocessor systems to decide which blocks to retain and which to drop in order to minimize communication over the bus.
Memory system characterization of commercial workloads
379 Citations1998Luiz André Barroso, Kourosh Gharachorloo +1 more
Token coherence
276 Citations2003Milo M. K. Martin, Mark D. Hill +1 more
Performance of database workloads on shared-memory systems with out-of-order processors
199 Citations1998Parthasarathy Ranganathan, Kourosh Gharachorloo +2 more
The results show that the combination of out-of-order execution and multiple instruction issue is effective in improving performance of database workloads, providing gains of 1.5 and 2.6 times over an in-order single-issue processor for OLTP and DSS, respectively.
An adaptive cache coherence protocol optimized for migratory sharing
190 Citations1993Per Stenström, Mats Brorsson +1 more
Dynamic self-invalidation
187 Citations1995Alvin R. Lebeck, David A. Wood
The results show that DSI reduces execution time of a sequentially consistent full-map coherence protocol by as much as 41%.
Adaptive cache coherency for detecting migratory shared data
174 Citations1993Anna L. Cox, Robert J. Fowler
ComputerSimulating a $2M commercial server on a $2K PC
158 Citations2003Alaa R. Alameldeen, Milo M. K. Martin +6 more
The authors have developed a simulation methodology that uses multiple simulations, pays careful attention to the effects of scaling on workload behavior, and extends Virtutech AB's Simics full system functional simulator with detailed timing models.
Selective, accurate, and timely self-invalidation using last-touch prediction
153 Citations2000An-Chow Lai, Babak Falsafi
IEEE Transactions on ComputersCache invalidation patterns in shared-memory multiprocessors
144 Citations1992Aman Gupta, W.-D. Weber
The cache invalidation patterns of several parallel applications are analyzed and a classification scheme for data objects found in parallel programs is proposed, indicating that cache line sizes in the 32-byte range yield the lowest data and invalidation traffic.
Full-system timing-first simulation
136 Citations2002Carl J. Mauer, Mark D. Hill +1 more
This paper advocates decoupled simulator organizations that separate functional and performance concerns, and defines an approach, called timing-first simulation, that uses an augmented timing simulator to execute instructions important to performance in conjunction with a functional simulator to insure correctness.
WildFire: a scalable path for SMPs
136 Citations1999Erik Hägersten, Michael Köster
This paper proposes new scalable techniques that allow large SMPs to be tied together efficiently, while maintaining the compatibility with, and performance characteristics of, an SMP.
ACM SIGARCH Computer Architecture NewsAdaptive software cache management for distributed shared memory architectures
135 Citations1990John K. Bennett, John B. Carter +1 more
Memory system characterization of commercial workloads
129 Citations2002Luiz André Barroso, Kourosh Gharachorloo +1 more
This study characterizes the memory system behavior of these workloads through a large number of architectural experiments on Alpha multiprocessors augmented with full system simulations to determine the impact of architectural trends.
ACM SIGPLAN NoticesArchitecture and design of AlphaServer GS320
114 Citations2000Kourosh Gharachorloo, Madhu Sharma +2 more
This paper describes the architecture and implementation of the AlphaServer GS320, a cache-coherent non-uniform memory access multiprocessor developed at Compaq and incorporates a couple of innovative techniques that extend previous approaches for efficiently implementing memory consistency models.
Reactive NUMA
109 Citations1997Babak Falsafi, David A. Wood
This paper proposes and evaluates a new approach to directory-based cache coherence protocols called Reactive NUMA (R-NUMA), which combines a conventional CC-N UMA coherence protocol with a more-recent Simple-COMA (S- COMA) protocol.
Using prediction to accelerate coherence protocols
106 Citations1998Shubhendu S. Mukherjee, Mark D. Hill
This paper takes the first step toward using general prediction to accelerate coherence protocols by developing and evaluating the Cosmos coherence message predictor, an extension of Yeh and Patt's two-level PAp branch predictor.
Memory sharing predictor: the key to a speculative coherent DSM
98 Citations1999An-Chow Lai, Babak Falsafi
The memory performance of DSS commercial workloads in shared-memory multiprocessors
93 Citations2002Pedro Trancoso, Josep-L. Larriba-Pey +2 more
This paper analyzes in detail the memory access patterns of several queries that are representative of Decision Support System (DSS) databases and shows that both Index and Sequential queries exhibit spatial locality and, therefore, can benefit from relatively long cache lines.
Multicast snooping: a new coherence method using a multicast address network
93 Citations1999E. Ender Bilir, Ross M. Dickson +5 more
Improving CC-NUMA performance using Instruction-based Prediction
87 Citations1999Stefanos Kaxiras, James Goodman
It is shown that for the first two optimizations, instruction-based prediction, using few predictor entries per node, outpaces address based schemes, and for the producer consumer optimization which uses speculative execution, low mis speculation rates show promise for performance improvements.
Dynamic self-invalidation: reducing coherence overhead in shared-memory multiprocessors
83 Citations2002Alvin R. Lebeck, David Wood
Evaluating Non-deterministic Multi-threaded Commercial Workloads
79 Citations2001Alaa R. Alameldeen, Carl J. Mauer +6 more
This paper introduces non-deterministic workload behavior as another potential challenge in timing simulation and proposes a methodology that uses pseudo-random perturbations and standard statistical techniques to compensate for these non-Deterministic effects.
IEEE Transactions on Parallel and Distributed SystemsSpecifying and verifying a broadcast and a multicast snooping cache coherence protocol
73 Citations2002Daniel J. Sorin, Manoj Plakal +4 more
This work develops a specification methodology that documents and specifies a cache coherence protocol in eight tables: the states, events, actions, and transitions of the cache and memory controllers, and demonstrates the utility of the table-based specification methodology.
Bandwidth adaptive snooping
71 Citations2004Milo M. K. Martin, Daniel J. Sorin +2 more
Bandwidth Adaptive Snooping Hybrid (BASH), a hybrid protocol that ranges from behaving like snooping when excess bandwidth is available to behaving like a directory protocol (by unicasting requests) when bandwidth is limited, is proposed.
Coherence communication prediction in shared-memory multiprocessors
64 Citations2002Stefanos Kaxiras, Cliff Young
This work presents a taxonomy of prediction schemes that includes all previously-proposed prediction schemes in a uniform space, and discovers prediction schemes more accurate than those previously proposed.
Conference on High Performance Computing (Supercomputing)Owner Prediction for Accelerating Cache-to-Cache Transfer Misses in a cc-NUMA Architecture
55 Citations2002Manuel E. Acacio, José González +2 more
Results indicate that owner prediction can significantly reduce the latency of cache-to-cache transfer misses, which translates into speed-ups on application performance up to 12% and the inclusion of a small and fast directory cache in every node is evaluated.
Token Coherence: decoupling performance and correctness
50 Citations2004Milo M. K. Martin, M.D. Hill +1 more
TokenB, a specific Token Coherence performance protocol that allows a glueless multiprocessor to both exploit a low latency unordered interconnect and avoid indirection and can significantly outperform traditional snooping and directory protocols.
Full-system timing-first simulation
46 Citations2002Carl J. Mauer, Mark D. Hill +1 more
The use of prediction for accelerating upgrade misses in cc-NUMA multiprocessors
43 Citations2002Manuel E. Acacio, José María Faci González +2 more
Adaptive software cache management for distributed shared memory architectures
40 Citations2002J.K. Bennett, J.B. Carter +1 more
IEEE Transactions on Parallel and Distributed SystemsDesign of an adaptive cache coherence protocol for large scale multiprocessors
38 Citations1992Qing Yang, G. Thangadurai +1 more
A large scale, cache-based multiprocessor that is interconnected by a hierarchical network such as hierarchical buses or a multistage interconnection network (MIN) is considered and an adaptive cache coherence scheme is proposed based on a hardware approach that handles multiple shared reads efficiently.
Boosting the performance of hybrid snooping cache protocols
35 Citations1995Fredrik Dahlgren
A hybrid protocol, extended with write caches as well as read snarfing, manages to reduce the number of coherence misses by between 83% and 95% as compared to a write-invalidate protocol for all five applications in this study.
Distance-adaptive update protocols for scalable shared-memory multiprocessors
28 Citations2002A. Raynaud, Zheng Zhang +1 more
A model of sharing that is key to investigating the performance of optimized update protocols: the update distance model is presented, which gives insight into the update patterns that optimized protocols need to handle.
Two adaptive hybrid cache coherency protocols
27 Citations2002C. Anderson, Anna R. Karlin
Performance measurements across a range of parallel applications indicate that the adaptive protocols presented perform well compared to both write-invalidate and write-update protocols.
The use of prediction for accelerating upgrade misses in cc-NUMA multiprocessors
26 Citations2003Manuel E. Acacio, Jesús González +2 more
Using execution-driven simulations, it is shown that the use of prediction can significantly accelerate upgrade misses (latency reductions of more than 40% in some cases) and translate into speed-ups on application performance up to 14%.
IEEE MicroSystem optimization for OLTP workloads
22 Citations1999Steven R. Kunkel, Brian Armstrong +1 more
This article connects recent AS/400 hardware advances with the corresponding approaches used to tune the system performance for large online transaction processing (OLTP) workloads, particularly those tuning efforts that affect the memory system.
IEEE Transactions on Parallel and Distributed SystemsThe potential of compile-time analysis to adapt the cache coherence enforcement strategy to the data sharing characteristics
20 Citations1995Farnaz Mounes-Toussi, David J. Lilja
A combined hardware-software strategy that uses the predictive capability of the compiler to select updating or invalidating for each write reference and can potentially reduce the miss ratio and reduce the generated network traffic with an ideal compiler.
Multicast snooping: a new coherence method using a multicast address network
19 Citations2003E.E. Bilir, R.M. Dickson +5 more
Preliminary performance numbers with mostly SPLASH-2 benchmarks running on 32 processors show that multicast snooping can obtain data directly but apply to larger systems (like directories).
Improving performance of load-store sequences for transaction processing workloads on multiprocessors
19 Citations2003Jonas Nilsson, Fredrik Dahlgren
Two conceptually different approaches for detecting load-store sequences are explored; data-centric and instruction-centric, and the goal is to load the cache block exclusively already at the load instruction so that the store instruction can perform locally.
Reducing ownership overhead for load-store sequences in cache-coherent multiprocessors
10 Citations2002Jonas Nilsson, Fredrik Dahlgren
This paper proposes a new hardware-based approach far performing this optimization by targeting load-store sequences, which is a super-set of migrator sharing and shows that the technique is able to reduce write-related latency and network traffic more than previousHardware-based techniques.
ACM SIGOPS Operating Systems ReviewPerformance of database workloads on shared-memory systems with out-of-order processors
7 Citations1998Parthasarathy Ranganathan, Kourosh Gharachorloo +2 more
ACM SIGARCH Computer Architecture NewsArchitecture and design of AlphaServer GS320
3 Citations2000Kourosh Gharachorloo, Madhu Sharma +2 more
Boosting the performance of hybrid snooping cache protocols
1 Citations2002F. Dahlgren
