A Reliability Study for Evaluating Information Extraction from Radiology Reports
Journal of the American Medical Informatics AssociationPublished 1 March 1999Open access
George Hripcsak, Gilad J. Kuperman, C Friedman, Daniel F. Heitjan
Citations49
SJR quartileQ1
SJR score2.04
SNIP1.95
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
In these evaluations, physician raters were able to judge very reliably the presence of clinical conditions based on text reports. Once the reliability of a specific rater is confirmed, it would be possible for that rater to create a reference standard reliable enough to assess aggregate measures on a system. Six raters would be needed to create a reference standard sufficient to assess a system on a case-by-case basis. These results should help evaluators design future information extraction studies for natural language processors and other knowledge-based systems.
Keywords
Computer ScienceMedicineBiochemistry, Genetics and Molecular Biology
The Design and Analysis of Clinical Experiments
4,314 Citations1999Joseph L. Fleiss
JAMAIncidence of adverse drug events and potential adverse drug events. Implications for prevention. ADE Prevention Study Group
2,509 Citations1995David W. Bates
Adverse drug events were common and often preventable; serious ADEs were more likely to be preventable and prevention strategies should target both stages of the drug delivery process.
PubMedIncidence of adverse drug events and potential adverse drug events. Implications for prevention. ADE Prevention Study Group.
1,109 Citations1995D W Bates, Cullen Dj +8 more
Computers and medicineEvaluation Methods in Medical Informatics
452 Citations1997Charles P. Friedman, Jeremy C Wyatt
This chapter discusses Subjectivist Approaches to Evaluation, the design and conduct of Subjectivist Studies, and Organizational Evaluation of Medical Information Resources.
Annals of Internal MedicineUnlocking Clinical Data from Narrative Reports: A Study of Natural Language Processing
338 Citations1995George Hripcsak, Carol Friedman +4 more
A general-purpose processor is evaluated that is intended to cover various clinical reports and converts narrative reports that are available in electronic form either through word processors or electronic scanning to coded descriptions that are appropriate for automated systems.
New England Journal of MedicinePerformance of Four Computer-Based Diagnostic Systems
326 Citations1994Eta S. Berner, George D. Webster +12 more
A profile of the strengths and limitations of the diagnostic capabilities of four internal medicine diagnostic systems: Dxplain, Iliad, Meditel, and QMR is provided.
Statistical Methods in Medical ResearchReview papers : Design and analysis of reliability studies
326 Citations1992Graham Dunn
This review covers the design and analysis of essentially two types of reliability study: method comparison studies and generalizability experiments, which are expected to be used to supplement the simpler traditional methods.
Journal of the American Medical Informatics AssociationDevelopment and Initial Validation of an Instrument to Measure Physicians' Use of, Knowledge about, and Attitudes Toward Computers
149 Citations1998R. D. Cork, William M. Detmer +1 more
The four scales of the questionnaire appear to measure with adequate reliability five attributes of academic physicians' attitudes toward computers in medical care: computer use, self-reported computer knowledge, demand for computer functionality, demandFor computer usability, and computer optimism.
Natural Language EngineeringNatural language processing in an operational clinical information system
124 Citations1995C Friedman, George Hripcsak +3 more
How the natural language system was made compatible with the existing CIS is described and engineering issues which involve performance, robustness, and accessibility of the data from the end users' viewpoint are discussed.
PubMedIdentification of suspected tuberculosis patients based on natural language processing of chest radiograph reports.
96 Citations1996Nilesh Jain, Charles Knirsch +2 more
A retrospective study to determine if MedLEE can identify patients at risk for having tuberculosis (TB) based on their admission chest radiographs to determine patient eligibility for computerized clinical practice guidelines.
Methods of Information in MedicineExtracting Findings from Narrative Reports: Software Transferability and Sources of Physician Disagreement
78 Citations1998G.J. Kuperman, C Friedman +1 more
MedLEE, a general-purpose natural language processor developed for Columbia-Presbyterian Medical Center, was compared to physicians' ability to detect seven clinical conditions in 200 Brigham and Women's Hospital chest radiograph reports, and performance improved so that it was indistinguishable from physicians.
Infection Control and Hospital EpidemiologyRespiratory Isolation of Tuberculosis Patients Using Clinical Guidelines and an Automated Clinical Decision Support System
74 Citations1998Charles Knirsch, Nilesh Jain +3 more
Journal of the American Medical Informatics AssociationAn Experiment Comparing Lexical and Statistical Methods for Extracting MeSH Terms from Clinical Free Text
66 Citations1998Gregory F. Cooper, Rosemary Anne O'Connor Miller
The results suggest possible approaches to reduce the number of terms output while maintaining the percentage of terms captured, including the use of UMLS semantic types to constrain the output list to contain only clinically relevant MeSH terms.
The LancetComparison of computer-aided and human review of general practitioners' management of hypertension
55 Citations1991Johan van der Lei, Jan H. van Bemmel +3 more
A computer program called hypercritic is written that audits general practitioners' management of patients with essential hypertension by taking patient-specific data from the ELIAS system, and automated review of computer-based medical records compares favourably with review by physicians.
Computers and Biomedical ResearchValidation of the medical expert system PNEUMON-IA
41 Citations1992A Verdaguer, Alex Patak +3 more
The method used to validate PNEUMON-IA could prove useful to assess the performance of expert systems in fields in fields where no gold standard is available and distances between arrays of etiological possibilities given by specialists and by PNEumon-IA were considered as an agreement measure between diagnoses.
PubMedAn evaluation of natural language processing methodologies.
22 Citations1998C Friedman, George Hripcsak +1 more
An existing MLP system MedLEE was modified and results from a previous study were used and results showed that the two methods based on obtaining the largest well-formed segment within a sentence had significantly higher sensitivity than the others.
PubMedKnowledge discovery and data mining to assist natural language understanding.
22 Citations1998Adam Wilcox, George Hripcsak
C5.0, a decision tree generator, was used to create a rule base for a natural language understanding system, and the generated rule base performed as well as lay persons, but worse than physicians.
Comparing human and machine performance for natural language information extraction
17 Citations1993Craig A. Will
Results are presented from a comparative study of human and machine performance for one of the information extraction tasks used in the MUC-5/Tipster evaluation that can help assess the maturity and applicability of the technology.
