Data Cleaning: Detecting, Diagnosing, and Editing Data Abnormalities
PLoS MedicinePublished 6 September 2005Open access
Jan Van den Broeck, Solveig A. Cunningham, R. Eeckels, Kobus Herbst
Citations435
SJR quartileQ1
SJR score4.28
SNIP3.21
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
In this policy forum the authors argue that data cleaning is an essential part of the research process, and should be incorporated into study design.
Keywords
Computer ScienceDecision SciencesMedicine
Neural Networks: A Comprehensive Foundation
29,816 Citations1998Simon Haykin
Thorough, well-organized, and completely up to date, this book examines all the important aspects of this emerging technology, including the learning process, back-propagation learning, radial-basis function networks, self-organizing systems, modular networks, temporal processing and neurodynamics, and VLSI implementation of neural networks.
TechnometricsAnalysis of Incomplete Multivariate Data
5,644 Citations2000David E. Booth, Joseph L. Schafer
Journal of the American Statistical AssociationClassical and Modern Regression With Applications.
3,716 Citations1988R. Kirk Steinhorst, Raymond H. Myers
The Annals of StatisticsThe Dip Test of Unimodality
1,906 Citations1985J. A. Hartigan, Pamela Hartigan
Statistics with Confidence
1,444 Citations2000Michael Smithson
Communications of the ACMA product perspective on total data quality management
898 Citations1998Richard Y. Wang
The purpose of this TDQM methodology is to deliver highquality information products (IP) to information consumers and aims to facilitate the implementation of an organization’s overall data quality policy formally expressed by top management.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Identifying Multiple Outliers in Multivariate Data
772 Citations1992Ali S. Hadi
BMJPost-randomisation exclusions: the intention to treat principle and excluding patients from analysis
694 Citations2002Dean Fergusson, Shawn D Aaron +2 more
The authors consider the circumstances when it may be possible to exclude patients from the analysis of data in clinical trials, even in an intention to treat trial.
Journal of Clinical EpidemiologyAttrition in longitudinal studies
512 Citations2002Jos W. R. Twisk, Wieke de Vente
This study showed that the theoretically more valid multiple imputation method did not lead to different point estimates than the more simple (longitudinal) imputation methods, and the estimated standard errors appeared to be theoretically more adequate, because they reflect the uncertainty in estimation caused by missing values.
Clinical ChemistryEffect of Outliers and Nonhealthy Individuals on Reference Interval Estimation
301 Citations2001Paul S. Horn, Lan Feng +2 more
Combining traditional and robust statistical techniques provide a good method of identifying outliers in a reference interval setting, even in healthy samples, and there is a large deviation among analytes.
Elsevier eBooksINFLUENCE FUNCTIONS AND REGRESSION DIAGNOSTICS
99 Citations1982Roy E. Welsch
The chapter presents the exploratory approach that considers an influence measure as a batch of n numbers and used the techniques of exploratory data analysis including stem-and-leaf plots, box plots, and transformations to symmetry to identify unusual observations.
Clinical Data Management
45 Citations1999
The first comprehensive volume on the subject of clinical data management, this book contains concise, well-researched information covering all aspects of data management from handling early phase I studies in volunteers to the presentation of final reports for regulatory purposes.
Journal of the American Statistical AssociationStandards for Discussion and Presentation of Errors in Survey and Census Data
15 Citations1975María E. González, Jack L. Ogus +2 more
This supplement includes illustrations of alternative methods of presenting sampling and nonsampling errors, and presents guidelines for consideration in preparing statistical reports, based on the guidelines developed by the Bureau in its Technical Paper 32.
Journal of Biopharmaceutical StatisticsThe impact of outlying subjects on decising of bioequivalence
15 Citations1995Fanny Y. C. Ki, Jen‐pei Liu +2 more
The impact of a statistically identified outlying subject on the decision of bioEquivalence through a simulation study under the structure of a standard two-way crossover design based on interval hypotheses for bioequivalence is examined.
American Journal of EpidemiologyEditing Data: What Difference Do Consistency Checks Make?
13 Citations2000Ursula E. Bauer, Tammie M. Johnson
The authors examined five possible approaches to handling data inconsistencies and the effect that each has on point estimates of current cigarette use in a self-administered school-based survey of tobacco use, attitudes, and behaviors in Florida.
Journal of the American Statistical AssociationData Base Error Trapping and Prediction
11 Citations1991Mike West, Robert L. Winkler
This work develops and analyzes models for a class of problems involving inferences about uncertain numbers of errors in data bases and generates inferences in terms of predictive distributions for the numbers of undetected errors.
