Unpredictable bias when using the missing indicator method or complete case analysis for missing confounder values: an empirical example
Journal of Clinical EpidemiologyPublished 26 March 2010
Mirjam J. Knol, Kristel J.M. Janssen, A. Rogier T. Donders, Toine C. G. Egberts, Eibert R. Heerdink, Diederick E. Grobbee
Citations181
SJR quartileQ1
SJR score3.15
SNIP2.66
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
MIM should not be used in handling missing confounder data because it gives unpredictable bias of the odds ratio even with small percentages of missing values, and CC can be used when missing values are completely random, but it gives loss of statistical power.
Abstract
MIM should not be used in handling missing confounder data because it gives unpredictable bias of the odds ratio even with small percentages of missing values. CC can be used when missing values are completely random, but it gives loss of statistical power.
Keywords
Social SciencesMathematics
PLoS MedicineThe Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) Statement: Guidelines for Reporting Observational Studies
21,612 Citations2007Erik von Elm, Douglas G. Altman +5 more
The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) Initiative developed recommendations on what should be included in an accurate and complete report of an observational study, resulting in a checklist of 22 items that relate to the title, abstract, introduction, methods, results, and discussion sections of articles.
Wiley series in probability and statisticsMultiple Imputation for Nonresponse in Surveys
20,606 Citations1987Donald B. Rubin
TechnometricsApplied Regression Analysis
18,042 Citations2005
This tutorial discusses simple and multiple linear regression, diagnostics, model selection, models with categorical variables, and nonlinear models; logistic regression.
Journal of the American Statistical AssociationStatistical Analysis With Missing Data
17,489 Citations1989Maureen Lahiff, Roderick J. A. Little +1 more
Psychological MethodsMissing data: Our view of the state of the art.
10,918 Citations2002Joseph L. Schafer, J. A. Graham
2 general approaches that come highly recommended: maximum likelihood (ML) and Bayesian multiple imputation (MI) are presented and may eventually extend the ML and MI methods that currently represent the state of the art.
Journal of the American Statistical AssociationMultiple Imputation after 18+ Years
2,949 Citations1996Donald B. Rubin
Archives of General PsychiatryThe Composite International Diagnostic Interview
2,901 Citations1988Lee N. Robins
The design and development of the CIDI is described and the current field testing of a slightly reduced "core" version is described, allowing investigators reliably to assess mental disorders according to the most widely accepted nomenclatures in many different populations and cultures.
Journal of the American Statistical AssociationMultiple Imputation After 18+ Years
2,693 Citations1996Donald B. Rubin
A description of the assumed context and objectives of multiple imputation is provided, and a review of the multiple imputations framework and its standard results are reviewed.
EpidemiologyStrengthening the Reporting of Observational Studies in Epidemiology (STROBE)
2,666 Citations2007Jan P. Vandenbroucke, Erik von Elm +7 more
Journal of Clinical EpidemiologyReview: A gentle introduction to imputation of missing values
2,581 Citations2006A. Rogier T. Donders, Geert J. M. G. van der Heijden +2 more
Journal of the Royal Statistical Society Series A (Statistics in Society)Statistical Analysis with Missing Data.
2,233 Citations1988Chris Chatfield, Roderick J. A. Little +1 more
Statistics in MedicineMultiple imputation of missing blood pressure covariates in survival analysis
2,156 Citations1999Stef van Buuren, Hendriek C. Boshuizen +1 more
A non-response problem in survival analysis where the occurrence of missing data in the risk factor is related to mortality is studied, and multiple imputation is used to impute missing blood pressure and then analyse the data under a variety of non- response models.
Statistics in MedicineMultiple imputation in health‐are databases: An overview and some applications
1,572 Citations1991Donald B. Rubin, Nathaniel Schenker
This paper provides an overview of methods for creating and analysing multiply-imputed data sets, and illustrates the dramatic improvements possible when using multiple rather than single imputation.
Journal of Clinical EpidemiologyUsing the outcome for imputation of missing predictor values was preferred
1,015 Citations2006Karel G.M. Moons, Rogier Donders +2 more
Journal of the American Statistical AssociationRegression with Missing<i>X</i>'s: A Review
974 Citations1992Roderick J. A. Little
Regression With Missing X's: A Review Author(s): Roderick J. A.
American Journal of EpidemiologyA Critical Look at Methods for Handling Missing Covariates in Epidemiologic Regression Analyses
886 Citations1995Sander Greenland, William D. Finkle
The authors recommend that epidemiologists avoid using the missing-indicator method and use more sophisticated methods whenever a large proportion of data are missing, and contrast the results of multiple imputation to simple methods in the analysis of a case-control study of endometrial cancer.
The American StatisticianMuch Ado About Nothing
760 Citations2007Nicholas J. Horton, Ken Kleinman
These routines to incorporate observations with incomplete variables in regression models are reviewed in the context of a motivating example from a large health services research dataset, and it is feasible to incorporate partially observed values.
Journal of the American Statistical AssociationRegression With Missing X's: A Review
622 Citations1992Roderick J. A. Little
Journal of Clinical EpidemiologyImputation of missing values is superior to complete case analysis and the missing-indicator method in multivariable diagnostic research: A clinical example
602 Citations2006Geert J. M. G. van der Heijden, A. Rogier T. Donders +2 more
In multivariable diagnostic research complete case analysis and the use of the missing-indicator method should be avoided, even when data are missing completely at random.
Statistics in MedicineAdjusting for partially missing baseline measurements in randomized trials
381 Citations2004Ian R. White, Simon G. Thompson
Joint modelling of baseline and outcome is the most efficient method, subject to three conditions, which are illustrated in a randomized trial in community psychiatry.
Journal of the American Statistical AssociationIndicator and Stratification Methods for Missing Explanatory Variables in Multiple Linear Regression
333 Citations1996Michael P. Jones
British Journal of CancerMissing covariate data within cancer prognostic studies: a review of current reporting and proposed guidelines
222 Citations2004Andrea Burton, Douglas G. Altman
American Journal of EpidemiologyBiased Estimation of the Odds Ratio in Case-Control Studies due to the Use of Ad Hoc Methods of Correcting for Missing Values for Confounding Variables
172 Citations1991Werner Vach, Mana Blettner
It is suggested that investigators should carry out validation studies to understand whether the missing values occur randomly across the study population or occur more frequently in specific subgroups.
Psychosomatic MedicineDepressive Symptoms in Subjects With Diagnosed and Undiagnosed Type 2 Diabetes
141 Citations2007Mirjam J. Knol, Eibert R. Heerdink +7 more
The findings suggest that disturbed glucose homeostasis is not associated with depressive symptoms, and suggests that depressive symptoms might be a consequence of the burden of diabetes.
American Journal of EpidemiologyMultiple Imputation of Baseline Data in the Cardiovascular Health Study
120 Citations2002Alice M. Arnold
The implementation of the software to impute missing baseline data in the setting of the Cardiovascular Health Study, a large, observational study, is described and an increase in power was evident and variable selection simplified when using the imputed data sets.
Journal of Clinical EpidemiologyBias arising from missing data in predictive models
117 Citations2006Marc H. Gorelick
All three methods of handling large amounts of missing data can lead to biased estimates of the OR and of model performance in predictive models.
Journal of Clinical EpidemiologyA comparison of analytic methods for non-random missingness of outcome data
110 Citations1995Sybil L. Crawford
Data from a study of patterns of care in disabled elders were used to evaluate several common methods when missingness of the outcome was nonrandom, and multiple model-based imputation provided an easily implemented method of adjustment for non-random non-response in both univariate and multivariate analyses.
BMC Public HealthPrediction of depression in European general practice attendees: the PREDICT study
98 Citations2006Michael King, Scott Weich +20 more
PubMedGene-covariate interaction between dysplastic nevi and the CDKN2A gene in American melanoma-prone families.
40 Citations2000Alisa M. Goldstein, María Martínez +2 more
There was a significant improvement in the likelihood when DN, total nevi or both covariates were added to the base model, which included dominant transmission of the CDKN2A gene and a linear increase of risk with the logarithm of age on the logit scale.
