Missing Data
Published 1 January 1995
Roderick J. A. Little, Nathaniel Schenker
Citations153
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Studies in the social and behavioral sciences frequently suffer from missing data. For instance, sample surveys often have some individuals who either refuse to participate or do not supply answers to certain questions, and panel studies often have incomplete data due to attrition. Recent comprehensive treatments of the subject of missing data include three volumes produced by the Panel on Incomplete Data of the Committee on National Statistics (Madow, Nisselson, and Olkin 1983; Madow and Olkin 1983; Madow, Olkin, and Rubin 1983) and Little and Rubin (1987).
Keywords
Social SciencesMathematics
BiometrikaThe central role of the propensity score in observational studies for causal effects
30,963 Citations1983Paul R. Rosenbaum, Donald B. Rubin
Wiley series in probability and statisticsMultiple Imputation for Nonresponse in Surveys
20,606 Citations1987Donald B. Rubin
IEEE Transactions on Pattern Analysis and Machine IntelligenceStochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images
17,980 Citations1984Stuart Geman, Donald Geman
The analogy between images and statistical mechanics systems is made and the analogous operation under the posterior distribution yields the maximum a posteriori (MAP) estimate of the image given the degraded observations, creating a highly parallel ``relaxation'' algorithm for MAP estimation.
The Annals of StatisticsBootstrap Methods: Another Look at the Jackknife
17,446 Citations1979B. Efron
Statistical ScienceInference from Iterative Simulation Using Multiple Sequences
16,747 Citations1992Andrew Gelman, Donald B. Rubin
The focus is on applied inference for Bayesian posterior distributions in real problems, which often tend toward normal- ity after transformations and marginalization, and the results are derived as normal-theory approximations to exact Bayesian inference, conditional on the observed simulations.
BiometricsRandom-Effects Models for Longitudinal Data
8,915 Citations1982Nan M. Laird, James H. Ware
A unified approach to fitting these models, based on a combination of empirical Bayes and maximum likelihood estimation of model parameters and using the EM algorithm, is discussed.
Journal of the American Statistical AssociationA Test of Missing Completely at Random for Multivariate Data with Missing Values
8,301 Citations1988Roderick J. A. Little
BiometricsLongitudinal Data Analysis for Discrete and Continuous Outcomes
7,834 Citations1986Scott L. Zeger, Kung‐Yee Liang
A class of generalized estimating equations (GEEs) for the regression parameters is proposed, extensions of those used in quasi-likelihood methods which have solutions which are consistent and asymptotically Gaussian even when the time dependence is misspecified as the authors often expect.
Journal of the American Statistical AssociationSampling-Based Approaches to Calculating Marginal Densities
6,616 Citations1990Alan E. Gelfand, A. F. M. Smith
Stochastic substitution, the Gibbs sampler, and the sampling-importance-resampling algorithm can be viewed as three alternative sampling- (or Monte Carlo-) based approaches to the calculation of numerical estimates of marginal probability distributions.
Digital Access to Scholarship at Harvard (DASH) (Harvard University)Maximum Likelihood from Incomplete Data via the EM Algorithm
4,523 Citations1977Dempster, Arthur P., Laird, Nan M. +1 more
Journal of the American Statistical AssociationThe Calculation of Posterior Distributions by Data Augmentation
3,782 Citations1987Martin A. Tanner, Wing Hung Wong
If data augmentation can be used in the calculation of the maximum likelihood estimate, then in the same cases one ought to be able to use it in the computation of the posterior distribution of parameters of interest.
The American StatisticianExplaining the Gibbs Sampler
2,355 Citations1992George Casella, Edward I. George
A simple explanation of how and why the Gibbs sampler works is given and analytically establish its properties in a simple case and insight is provided for more complicated cases.
Statistical SciencePractical Markov Chain Monte Carlo
2,028 Citations1992Charles J. Geyer
The case is made for basing all inference on one long run of the Markov chain and estimating the Monte Carlo error by standard nonparametric methods well-known in the time-series and operations research literature.
BiometrikaMaximum likelihood estimation via the ECM algorithm: A general framework
1,814 Citations1993Xiao‐Li Meng, Donald B. Rubin
Statistics in MedicineMultiple imputation in health‐are databases: An overview and some applications
1,572 Citations1991Donald B. Rubin, Nathaniel Schenker
This paper provides an overview of methods for creating and analysing multiply-imputed data sets, and illustrates the dramatic improvements possible when using multiple rather than single imputation.
BiometricsUnbalanced Repeated-Measures Models with Structured Covariance Matrices
1,346 Citations1986Robert I. Jennrich, Mark Schluchter
This work addresses the question of how to analyze unbalanced or incomplete repeated-measures data through maximum likelihood analysis using a general linear model for expected responses and arbitrary structural models for the within-subject covariances.
Proceedings of the Edinburgh Mathematical SocietyApplications of Mathematics to Medical Problems
1,177 Citations1925A. G. M'kendrick
If one thinks of individuals, be they human beings or be they cells, as moving in all sorts of dimensions, reversibly or irreversibly, continuously or discontinuously, by unit stages or per saltum, then the method of their movement becomes a study in kinetics, and can be approached by the methods ordinarily adopted in the study of such systems.
Journal of the American Statistical AssociationIllustration of Bayesian Inference in Normal Data Models Using Gibbs Sampling
966 Citations1990Alan E. Gelfand, Susan E. Hills +2 more
The use of the Gibbs sampler as a method for calculating Bayesian marginal posterior and predictive densities is reviewed and illustrated with a range of normal data models, including variance components, unordered and ordered means, hierarchical growth curves, and missing data in a crossover trial.
Journal of the American Statistical AssociationPattern-Mixture Models for Multivariate Incomplete Data
961 Citations1993Roderick J. A. Little
Journal of Business and Economic StatisticsMissing-Data Adjustments in Large Surveys
944 Citations1988Roderick J. A. Little
Useful properties of a general-purpose imputation method for numerical data are suggested and discussed in the context of several large government surveys and weighting-based analogs to predictive mean matching are outlined.
PsychometrikaOn Structural Equation Modeling with Data that are not Missing Completely at Random
889 Citations1987Bengt Muthén, David M. Kaplan +1 more
A general latent variable model is given which includes the specification of a missing data mechanism which allows for an elucidating discussion of existing general multivariate theory bearing on maximum likelihood estimation with missing data.
Journal of the American Statistical AssociationNewton—Raphson and EM Algorithms for Linear Mixed-Effects Models for Repeated-Measures Data
879 Citations1988Mary J. Lindstrom, Douglas M. Bates
Journal of the American Statistical AssociationMultiple Imputation for Interval Estimation from Simple Random Samples with Ignorable Nonresponse
722 Citations1986Donald B. Rubin, Nathaniel Schenker
BiometricsEstimating Equations for Parameters in Means and Covariances of Multivariate Discrete and Continuous Responses
555 Citations1991R. L. Prentice, Lue Ping Zhao
A class of quadratic exponential models is used to develop joint estimating equations for mean and covariance parameters in a more systematic fashion, and proposals for the use of such equations are developed.
Journal of the American Statistical AssociationUsing EM to Obtain Asymptotic Variance-Covariance Matrices: The SEM Algorithm
543 Citations1991Xiao‐Li Meng, Donald B. Rubin
This article defines and illustrates a procedure that obtains numerically stable asymptotic variance–covariance matrices using only the code for computing the complete-data variance-covarance matrix, the code of the expectation maximization algorithm, and code for standard matrix operations.
International Statistical ReviewSurvey Nonresponse Adjustments for Estimates of Means
522 Citations1986Roderick J. A. Little
The Review of Economic StudiesSome Approaches to the Correction of Selectivity Bias
484 Citations1982Lung‐fei Lee
Journal of Business and Economic StatisticsStatistical Matching Using File Concatenation With Adjusted Weights and Multiple Imputations
477 Citations1986Donald B. Rubin
The method of file concatenation with adjusted weights and multiple imputations is described and illustrated on an artificial example, showing the ability to display sensitivity of inference to untestable assumptions being made when creating the matched file.
A MISSING INFORMATION PRINCIPLE: THEORY AND APPLICATIONS
456 Citations1972Terence Orchard, Max A. Woodbury
The problem that a relatively simple analysis is changed into a complex one just because some of the information is missing, is one which faces most practicing statisticians at some point in their career.
Journal of the American Statistical AssociationApproximate Bayesian Inference in Conditionally Independent Hierarchical Models (Parametric Empirical Bayes Models)
450 Citations1989Robert E. Kass, Duane Steffey
Journal of the American Statistical AssociationMaximum Likelihood Estimates for a Multivariate Normal Distribution when Some Observations are Missing
395 Citations1957T. W. Anderson
BiometrikaA class of pattern-mixture models for normal incomplete data
387 Citations1994Roderick J. A. Little
The Annals of Mathematical StatisticsMultivariate Correlation Models with Mixed Discrete and Continuous Variables
378 Citations1961Ingram Olkin, R. F. Tate
Journal of the American Statistical AssociationFormalizing Subjective Notions about the Effect of Nonrespondents in Sample Surveys
357 Citations1977Donald B. Rubin
Journal of the American Statistical AssociationPost-Stratification: A Modeler's Perspective
288 Citations1993Roderick J. A. Little
Sociological Methods & ResearchThe Treatment of Missing Data in Multivariate Analysis
285 Citations1977Jae-On Kim, James Curry
How to assess the nature of missing data especially with regard to randomness, a comparison of listwise and pairwise deletion, and methods for using maximum information to estimate parameters or missing values are covered.
Communication in Statistics- Theory and MethodsCommunications in statistics: 1972-1976
256 Citations1978Richard F. Gunst, Tsushung A. Hua
Journal of the American Statistical AssociationImputation of Missing Values When the Probability of Response Depends on the Variable Being Imputed
241 Citations1982John Greenlees, William S. Reece +1 more
BiometrikaMaximum likelihood estimation for mixed continuous and categorical data with missing values
241 Citations1985Roderick J. A. Little, Mark Schluchter
Journal of the Royal Statistical Society Series C (Applied Statistics)Multiple Imputation for the Fatal Accident Reporting System
228 Citations1991Daniel F. Heitjan, Roderick J. A. Little
Two specific methods of multiple imputation based on predictive mean matching are described and applied to a sample of the FARS data and a simulation study compares the frequency properties of the methods.
Journal of the American Statistical AssociationRegression Analysis for Categorical Variables with Outcome Subject to Nonignorable Nonresponse
214 Citations1988Stuart G. Baker, Nan M. Laird
Journal of Political EconomyWhat Do We Really Know about Wages? The Importance of Nonreporting and Census Imputation
212 Citations1986Lee A. Lillard, James P. Smith +1 more
Sociological Methods & ResearchTheory Testing in a World of Constrained Research Design
192 Citations1990Ross M. Stolzenberg, Daniel A. Relles
BiometrikaCharacterizing the effect of matching using linear propensity score methods with normal distributions
188 Citations1992Donald B. Rubin, Neal Thomas
Journal of the American Statistical AssociationMultiple Imputation of Industry and Occupation Codes in Census Public-use Samples Using Bayesian Logistic Regression
183 Citations1991Clifford C. Clogg, Donald B. Rubin +3 more
Selection Modeling Versus Mixture Modeling with Nonignorable Nonresponse
169 Citations1986Robert J. Glynn, Nan M. Laird +1 more
Journal of the American Statistical AssociationMaximum Likelihood Estimation and Model Selection in Contingency Tables with Missing Data
146 Citations1982Camil Fuchs
Journal of the American Statistical AssociationAlternative Methods for CPS Income Imputation
130 Citations1986Martin David, Roderick J. A. Little +2 more
Journal of the American Statistical AssociationCausal Models for Patterns of Nonresponse
130 Citations1986Robert E. Fay
Journal of Food ScienceImputation of the 1989 Survey of Consumer Finances: Stochastic Relaxation and Multiple Imputation
123 Citations1997Arthur B. Kennickell
The author thanks the advice and support of Fritz Scheuren, comments by Roderick Little and Donald Rubin on an earlier description of the imputation model presented here, the patience and creativity of Louise Woodburn in serving as sounding board for many of the ideas, and Gerhard Fries for valuable help in debugging the software.
Journal of Statistical Computation and SimulationImputation using markov chains
116 Citations1988Kim-hung Li
In this paper, an iterative imputation procedure, based on the idea of Markov chain, is proposed, and examples are presented to illustrate its applications.
The Annals of StatisticsAsymptotic Results for Multiple Imputation
109 Citations1988Nathaniel Schenker, A. H. Welsh
BiometricsTwo-Dimensional Contingency Tables with Both Completely and Partially Cross-Classified Data
102 Citations1974Tar Chen, Stephen E. Fienberg
Journal of the American Statistical AssociationLinear Regression Analysis with Missing Observations among the Independent Variables
92 Citations1964Marvin Glasser
BiometricsAnalyzing Repeated Measures on Generalized Linear Models via the Bootstrap
77 Citations1989Lawrence H. Moulton, Scott L. Zeger
In this paper simple methods are introduced for the class of generalized linear models (GLMs).
Journal of the American Statistical AssociationMissing Observations in Multivariate Statistics II. Point Estimation in Simple Linear Regression
73 Citations1967A. A. Afifi, Robert M. Elashoff
Journal of Business and Economic StatisticsProjecting From Advance Data Using Propensity Modeling: An Application to Income and Tax Statistics
72 Citations1992John L. Czajka, Sharon M. Hirabayashi +2 more
Statistics in MedicineMultiple imputation for threshold‐crossing data with interval censoring
67 Citations1993Frederick J. Dorey, Roderick J. A. Little +1 more
New methods based on multiple imputation of the threshold-crossing time with use of models that take into account values recorded at the times of visits are presented.
Statistics in MedicineAnalysis of incomplete multivariate data using linear models with structured covariance matrices
66 Citations1988Mark Schluchter
A variety of applications of the model are discussed, including univariate and multivariate analysis of incomplete repeated measures data, analysis of growth curves with missing data using random effects and time-series models, and applications to unbalanced longitudinal data.
Communication in Statistics- Theory and MethodsUsing the jackknife to estimate the variance of regression estimators from repeated measures studies
61 Citations1990Stuart R. Lipsitz, Nan M. Laird +1 more
PsychometrikaLeast-Squares Theory Based on General Distributional Assumptions with an Application to the Incomplete Observations Problem
37 Citations1985B. M. S. Van Praag, Theo K. Dijkstra +1 more
Journal of the Royal Statistical Society Series C (Applied Statistics)Analyses of Public Use Decennial Census Data with Multiply Imputed Industry and Occupation Codes
29 Citations1993Nathaniel Schenker, Donald J. Treiman +1 more
A recently completed project in which multiple imputation was used to recalibrate industry and occupation codes in 1970 U.S. census public use samples to the 1980 standard is described.
Statistics in MedicineEstimation of parameters and missing values under a regression model with non‐normally distributed and non‐randomly incomplete data
27 Citations1989Stanley P. Azen, Michael Van Guilder +1 more
It is found that the EM and complete cases algorithms performed equally well regardless of the correlational structure, when the percentage of incomplete data was only 5 per cent and when this percentage increased to 25 per cent, the EM algorithm was generally best for estimation, but the complete cases algorithm was safe and conservative.
Sociological MethodologyEvaluating a Multiple-Imputation Method for Recalibrating 1970 U.S. Census Detailed Industry Codes to the 1980 Standard
24 Citations1988Donald J. Treiman, William T. Bielby +1 more
A multiple-imputation procedure for converting data coded into the 1970 U.S. census detailed classification of industries to the categories of the 1980 classification is evaluated.
TechnometricsComparing Regressions When Some Predictor Values Are Missing
23 Citations1976Donald B. Rubin
