login

Efficient provenance storage

Published 9 June 2008
Adriane Chapman, H. V. Jagadish, Prakash Ramanan
Citations191

TL;DR

This work has used the techniques described in this work to significantly reduce the provenance storage costs associated with constructing MiMI, a warehouse of data regarding protein interactions, as well as two provenance stores, Karma and PReServ, produced through workflow execution.

Abstract

As the world is increasingly networked and digitized, the data we store has more and more frequently been chopped, baked, diced and stewed. In consequence, there is an increasing need to store and manage provenance for each data item stored in a database, describing exactly where it came from, and what manipulations have been applied to it. Storage of the complete provenance of each data item can become prohibitively expensive. In this paper, we identify important properties of provenance that can be used to considerably reduce the amount of storage required.

Keywords

Computer ScienceDecision Sciences