Big data, but are we ready?
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work states that computational analysis in biology is high-dimensional, and predicts that peta-bytes, even exabytes, of data will be soon stored and analysed, and illustrates, through a simple calculation, how suitable current computational technologies really are for such large volumes of data.
Abstract
, 647–657 (2010)) , which presents cloud and heterogeneous computing as solutions for tackling large-scale and high-dimensional data sets. These technologies have been around for years, raising the question: why are they not used more often in bioinformatics? The answer is that, apart from introducing complexity, they quickly break down when a large amount of data is communicated between computing nodes.In their Review, Schadt and colleagues state that computational analysis in biology is high-dimensional, and predict that peta-bytes, even exabytes, of data will be soon stored and analysed. We agree with this predicted scenario and illustrate, through a simple calculation, how suitable current computational technologies really are for such large volumes of data.Currently, it takes minimally 9 hours for each of 1,000 cloud nodes to process 500 GB, at a cost of US$3,000 (500 GB to 500 TB of total data). The bottleneck in this process is the input/output (IO) hardware that links data storage to the calculation node
