4GHz+ low-latency fixed-point and binary floating-point execution units for the POWER6 processor
Published 1 January 2006
Brian Curran, B.D. McCredie, L. Sigal, E. Schwarz, Bruce Fleischer, Yiu-Hing Chan
Citations18
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A 1-pipe stage, low-latency, 13 FO4, 64b fixed-point execution unit, implemented in a 65nm SOI CMOS process, allows back-to-back execution of data dependent adds, subtracts, compares, shifts, rotates, and logical operations.
Abstract
A 1-pipe stage, low-latency, 13 FO4, 64b fixed-point execution unit, implemented in a 65nm SOI CMOS process, allows back-to-back execution of data dependent adds, subtracts, compares, shifts, rotates, and logical operations. A 7-pipe stage, 91 FO4, double-precision floating-point unit allows forwarding of dependent results after 6 cycles in most cases
Keywords
Computer Science
IBM Journal of Research and DevelopmentPOWER4 system microarchitecture
649 Citations2002Joel M. Tendler, J. S. Dodson +3 more
The processor microarchitecture as well as the interconnection architecture employed to form systems up to a 32-way symmetric multiprocessor are described.
IEEE Transactions on ComputersFPU Implementations with Denormalized Numbers
39 Citations2005E. Schwarz, M.S. Schmookler +1 more
This paper summarizes the little known techniques for handling denormalized numbers and underflows, and most of the techniques described here only appear in filed or pending patent applications.
