The Challenges of Creating a Gold Standard for De-identification Research.
PubMedPublished 1 January 2014Open access
Allen C. Browne, Mehmet Kayaalp, Zeyno A Dodd, Pamela Sagan, Clement J. McDonald
Citations6
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A Gold Standard corpus comprised over 20,000 records of annotated narrative clinical reports for use in the training and evaluation of NLM Scrubber, a de-identification software system for medical records.
Abstract
We created a Gold Standard corpus comprised over 20,000 records of annotated narrative clinical reports for use in the training and evaluation of NLM Scrubber, a de-identification software system for medical records. Our experience with designing the corpus demonstrated the conceptual complexity of the task.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
PubMedDe-identification of Address, Date, and Alphanumeric Identifiers in Narrative Clinical Reports.
28 Citations2014Mehmet Kayaalp, Allen C. Browne +3 more
NLM Scrubber's sensitivity on de-identifying patient names, alphanumeric identifiers, addresses and dates was 99%.
Journal of the American Medical Informatics AssociationThe pattern of name tokens in narrative clinical text and a comparison of five systems for redacting them
17 Citations2013Mehmet Kayaalp, Allen C. Browne +5 more
The nature and size of name lists have substantial influences on scrubing success, and the use of very large name lists with frequency statistics accounts for much of NLM-NS scrubbing success.
