login

Privacy-preserving <i>k</i> -means clustering over vertically partitioned data

Published 24 August 2003
Jaideep Vaidya, Chris Clifton
Citations631

TL;DR

This work presents a method for k-means clustering when different sites contain different attributes for a common set of entities, where each site learns the cluster of each entity, but learns nothing about the attributes at other sites.

Abstract

Privacy and security concerns can prevent sharing of data, derailing data mining projects. Distributed knowledge discovery, if done correctly, can alleviate this problem. The key is to obtain valid results, while providing guarantees on the (non)disclosure of data. We present a method for k-means clustering when different sites contain different attributes for a common set of entities. Each site learns the cluster of each entity, but learns nothing about the attributes at other sites.

Keywords

Computer ScienceDecision Sciences