login

Automatic de-identification of textual documents in the electronic health record: a review of recent research

BMC Medical Research MethodologyPublished 2 August 2010Open access
Stéphane M. Meystre, F Jeffrey Friedlin, Brett R. South, Shuying Shen, Matthew H. Samore
Citations322
SJR quartileQ1
SJR score1.67
SNIP1.73
View PDF

TL;DR

A review of recent research in automated de-identification of narrative text documents from the electronic health record finds methods based on dictionaries performed better with PHI that is rarely mentioned in clinical text, but are more difficult to generalize.

Abstract

In general, methods based on dictionaries performed better with PHI that is rarely mentioned in clinical text, but are more difficult to generalize. Methods based on machine learning tend to perform better, especially with PHI that is not mentioned in the dictionaries used. Finally, the issues of anonymization, sufficient performance, and "over-scrubbing" are discussed in this publication.

Keywords

Computer ScienceDecision SciencesHealth Professions