A Contingency-Table Model for Imputing Data Satisfying Analytic Constraints
Published 1 January 2002
William E. Winkler
Citations19
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This paper describes a method for imputation in general contingency tables when the imputations are subject to both analytic (edit) constraints and probabilistic distributional constraints. The model extends edit ideas in Fellegi and Holt (1976) and Winkler and Chen (2002). The model extends missing-at-random imputation ideas in Little and Rubin (1987). Some of the ideas are related to Friedman (2001) and Thibaudeau and Winkler (2002).
Keywords
Computer ScienceMathematics
The Annals of StatisticsGreedy function approximation: A gradient boosting machine.
28,973 Citations2001Jerome H. Friedman
A general gradient descent boosting paradigm is developed for additive expansions based on any fitting criterion, and specific algorithms are presented for least-squares, least absolute deviation, and Huber-M loss functions for regression, and multiclass logistic likelihood for classification.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
19,345 Citations2013Trevor Hastie, Robert Tibshirani +1 more
Journal of the American Statistical AssociationStatistical Analysis With Missing Data
17,489 Citations1989Maureen Lahiff, Roderick J. A. Little +1 more
Journal of the Royal Statistical Society Series B (Statistical Methodology)Local Computations with Probabilities on Graphical Structures and Their Application to Expert Systems
3,970 Citations1988Steffen L. Lauritzen, David J. Spiegelhalter
This work exploits a range of local representations for the joint probability distribution, combined with topological changes to the original network termed 'marrying' and 'filling-in', which allows efficient algorithms for transfer between representations, providing rapid absorption and propagation of evidence.
Learning Belief Networks in the Presence of Missing Values and Hidden Variables
391 Citations1997Nir Friedman
This paper proposes a new method for learning network structure from incomplete data based on an extension of the Expectation-Maximization (EM) algorithm for model selection problems that performs search for the best structure inside the EM procedure.
Journal of the American Statistical AssociationA Systematic Approach to Automatic Edit and Imputation
329 Citations1976Ivan P. Fellegi, Daniel T. Holt
A Systematic Approach to Automatic Edit and Imputation and its Applications to Mathematical Statistics.
Selectivity estimation using probabilistic models
265 Citations2001Lise Getoor, Benjamin Taskar +1 more
The approach produces more accurate estimates than standard approaches to selectivity estimation, using comparable space and time for both single-table multi-attribute queries and a general class of select-join queries.
Operations ResearchOptimal Imputation of Erroneous Data: Categorical Data, General Edits
50 Citations1986Robert Garfinkel, Anand S. Kunnathur +1 more
A model in which a response is modified to pass a set of edits with as little change as possible is developed, which is NP-hard for categorical data and general edits.
SET-COVERING AND EDITING DISCRETE DATA
19 Citations1998William E. Winkler
New set covering algorithms associated with the DISCRETE edit system correctly generate implicit edits for large subclasses and reduce computation during implicit-edit generation by as much as two orders of magnitude.
Extending the Fellegi-Holt Model of Statistical Data Editing
16 Citations2001William E. Winkler, Bor-Chung Chen
Set Covering Algorithms in Edit Generation
14 Citations1998Bor-Chung Chen, U. S. Bureau
