login

Recording and Reasoning over Data Provenance in Web and Grid Services

Lecture notes in computer sciencePublished 1 January 2003
Martin Szomszor, Luc Moreau
Citations104
SJR quartileQ2
SJR score0.35
SNIP0.55

TL;DR

This work proposes an infrastructure level support for a provenance recording capability for service-oriented architectures such as the Grid and Web Services and provides a mechanism by which provenance is used to determine whether previous computed results are still up to date.

Abstract

Large-scale, dynamic and open environments such as the Grid and Web Services build upon existing computing infrastructures to supply dependable and consistent large-scale computational systems. This kind of architecture has been adopted by the business and scientific communities allowing them to exploit extensive and diverse computing resources to perform complex data processing tasks. In such systems, results are often derived by composing multiple, geographically distributed, heterogeneous services as specified by intricate workflow management. This leads to the undesirable situation where the results are known, but the means by which they were achieved is not. With both scientific experiments and business transactions, the notion of lineage and dataset derivation is of paramount importance since without it, information is potentially worthless. We address the issue of {\\em data provenance\\/}, the description of the origin of a piece of data, in these environments showing the requirements, uses and implementation difficulties. We propose an infrastructure level support for a provenance recording capability for service-oriented architectures such as the Grid and Web Services. We also offer services to view and retrieve provenance and we provide a mechanism by which provenance is used to determine whether previous computed results are still up to date.

Keywords

Computer ScienceDecision Sciences