Efficient provenance storage
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work has used the techniques described in this work to significantly reduce the provenance storage costs associated with constructing MiMI, a warehouse of data regarding protein interactions, as well as two provenance stores, Karma and PReServ, produced through workflow execution.
Abstract
As the world is increasingly networked and digitized, the data we store has more and more frequently been chopped, baked, diced and stewed. In consequence, there is an increasing need to store and manage provenance for each data item stored in a database, describing exactly where it came from, and what manipulations have been applied to it. Storage of the complete provenance of each data item can become prohibitively expensive. In this paper, we identify important properties of provenance that can be used to considerably reduce the amount of storage required.
