Big Data Applications Using Workflows for Data Parallel Computing
Computing in Science & EngineeringPublished 16 April 2014Open access
Jianwu Wang, Daniel Crawl, İlkay Altıntaş, Weizhong Li
Citations33
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
An easy-to-use, scalable approach to build and execute Big Data applications using actor-oriented modeling in data parallel computing is presented.
Abstract
In the Big Data era, workflow systems need to embrace data parallel computing techniques for efficient data analysis and analytics. Here, the authors present an easy-to-use, scalable approach to build and execute Big Data applications using actor-oriented modeling in data parallel computing. They use two bioinformatics use cases for next-generation sequencing data analysis to verify the feasibility of their approach.
Keywords
Computer ScienceDecision Sciences
Communications of the ACMMapReduce
18,538 Citations2008Jay B. Dean, Sanjay Ghemawat
This presentation explains how the underlying runtime system automatically parallelizes the computation across large-scale clusters of machines, handles machine failures, and schedules inter-machine communication to make efficient use of the network and disks.
Computer NetworksThe Internet of Things: A survey
15,290 Citations2010Luigi Atzori, Antonio Iera +1 more
This survey is directed to those who want to approach this complex discipline and contribute to its development, and finds that still major issues shall be faced by the research community.
Concurrency and Computation Practice and ExperienceScientific workflow management and the Kepler system
1,696 Citations2005Bertram Ludäscher, İlkay Altıntaş +7 more
Characteristics of and requirements for scientific workflows as identified in a number of application projects are described, and some key features of Kepler and its underlying Ptolemy II system, planned extensions, and areas of future research are described.
BioinformaticsCloudBurst: highly sensitive read mapping with MapReduce
631 Citations2009Michael C. Schatz
CloudBurst is a new parallel read-mapping algorithm optimized for mapping next-generation sequence data to the human genome and other reference genomes, for use in a variety of biological analyses including SNP discovery, genotyping and personal genomics.
Workflows For E-Science: Scientific Workflows For Grids
512 Citations2007Ian Taylor, Ewa Deelman +2 more
This is a timely book presenting an overview of the current state-of-the-art within established projects, presenting many different aspects of workflow from users to tool builders.
Nephele/PACTs
263 Citations2010Dominic Battré, Stephan Ewen +4 more
The PACT programming model is a generalization of the well-known map/reduce programming model, extending it with further second-order functions, as well as with Output Contracts that give guarantees about the behavior of a function.
BMC BioinformaticsAnalysis and comparison of very large metagenomes with fast clustering and functional annotation
130 Citations2009Weizhong Li
RAMMCAP is a very fast method that can cluster and annotate one million metagenomic reads in only hundreds of CPU hours.
Nova
116 Citations2011Christopher Olston, Greg Chiou +11 more
A workflow manager developed and deployed at Yahoo called Nova is described, which pushes continually-arriving data through graphs of Pig programs executing on Hadoop clusters, which is a good fit for a large fraction of Yahoo's data processing use-cases.
Oozie
86 Citations2012Mohammad Samiul Islam, Angelo K. Huang +6 more
This paper finds that conventional workflow management tools lack at least one of these qualities of Scalability, Security, Multi-tenancy, and Operability, and therefore presents Apache Oozie, a workflow management system specialized for Hadoop.
Future Generation Computer SystemsHeterogeneous composition of models of computation
61 Citations2008Antoon Goderis, Christopher Brooks +3 more
This paper explains how MoCs are combined in Kepler and Ptolemy II and analyzes which combinations of MoC are currently possible and useful and demonstrates the approach by combining MoCs involving dataflow and finite state machines.
Challenges and approaches for distributed workflow-driven analysis of large-scale biological data
27 Citations2012İlkay Altıntaş, Jianwu Wang +2 more
The challenges related to next-generation sequencing data are discussed, the approaches taken in bioKepler to help with analysis of such data are explained, and preliminary results demonstrating these approaches are presented.
Maryland Shared Open Access Repository (USMAI Consortium)Comparison of Distributed Data-Parallelization Patterns for Big Data Analysis: A Bioinformatics Case Study
6 Citations2013Wang, Jianwu, Crawl, Daniel +3 more
