login

Profiling linked open data with ProLOD

Published 1 March 2010
Christoph Böhm, Felix Naumann, Ziawasch Abedjan, Dandy Fenz, Toni Grütze, Daniel Hefenbrock
Citations49

TL;DR

A suite of methods ranging from the domain level (clustering, labeling), via the schema level (matching, disambiguation), to the data level (data type detection, pattern detection, value distribution) are proposed, packaged into an interactive, web-based tool that allows iterative exploration and discovery of new LOD sources.

Abstract

Abstract — Linked open data (LOD), as provided by a quickly growing number of sources constitutes a wealth of easily acces-sible information. However, this data is not easy to understand. It is usually provided as a set of (RDF) triples, often enough in the form of enormous files covering many domains. What is more, the data usually has a loose structure when it is derived from end-user generated sources, such as Wikipedia. Finally, the quality of the actual data is also worrisome, because it may be incomplete, poorly formatted, inconsistent, etc. To understand and profile such linked open data, traditional data profiling methods do not suffice. With ProLOD, we propose a suite of methods ranging from the domain level (clustering, labeling), via the schema level (matching, disambiguation), to the data level (data type detection, pattern detection, value distribution). Packaged into an interactive, web-based tool, they

Keywords

Computer ScienceBiochemistry, Genetics and Molecular Biology