External validation of prognostic models for critically ill patients required substantial sample sizes
Journal of Clinical EpidemiologyPublished 6 February 2007
Niels Peek, Daniëlle G. T. Arts, Robert-Jan Bosman, Peter H. J. van der Voort, Nicolette F. de Keizer
Citations92
SJR quartileQ1
SJR score3.15
SNIP2.66
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Substantial sample sizes are required for performance assessment and model comparison in external validation of prognostic models for outcome at intensive care units (ICUs) andCalibration statistics and significance tests should not be used in these settings.
Abstract
Substantial sample sizes are required for performance assessment and model comparison in external validation. Calibration statistics and significance tests should not be used in these settings. Instead, a simple customization method to repair lack-of-fit problems is recommended.
Keywords
MathematicsMedicine
BiometricsComparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach
22,316 Citations1988Elizabeth R. DeLong, David M. DeLong +1 more
A nonparametric approach to the analysis of areas under correlated ROC curves is presented, by using the theory on generalized U-statistics to generate an estimated covariance matrix.
RadiologyThe meaning and use of the area under a receiver operating characteristic (ROC) curve.
21,812 Citations1982J A Hanley, Barbara J. McNeil
A representation and interpretation of the area under a receiver operating characteristic (ROC) curve obtained by the "rating" method, or by mathematical predictions based on patient characteristics, is presented and it is shown that in such a setting the area represents the probability that a randomly chosen diseased subject is (correctly) rated or ranked with greater suspicion than a random chosen non-diseased subject.
Critical Care MedicineAPACHE II
13,694 Citations1985William A. Knaus, Elizabeth A. Draper +2 more
The form and validation results of APACHE II, a severity of disease classification system that uses a point score based upon initial values of 12 routine physiologic measurements, age, and previous health status, are presented.
Critical Care MedicineAPACHE II-A Severity of Disease Classification System
13,396 Citations1986William A. Knaus, Elizabeth A. Draper +2 more
The form and validation results of APACHE II, a severity of disease classification system, are presented, showing an increasing score was closely correlated with the subsequent risk of hospital death for 5815 intensive care admissions from 13 hospitals.
JAMAA new Simplified Acute Physiology Score (SAPS II) based on a European/North American multicenter study
6,363 Citations1993J. R. Le Gall
The SAPS II, based on a large international sample of patients, provides an estimate of the risk of death without having to specify a primary diagnosis, and is a starting point for future evaluation of the efficiency of intensive care units.
Journal of the American Statistical AssociationRegression Modeling Strategies: With Applications to Linear Models, Logistic Regression, and Survival Analysis
1,975 Citations2003Sunil J Rao
The basic Bayesian framework must be constrained, use of the step function in computing the probability that a team would rank best or worst in a league, and implementation of a Dirichlet process prior are presented.
Statistics in MedicineWhat do we mean by validating a prognostic model?
1,425 Citations2000Douglas G. Altman, Patrick Royston
How to validate a model is considered and it is suggested that it is desirable to consider two rather different aspects - statistical and clinical validity - and some general approaches to validation are examined.
Annals of Internal MedicineAssessing the Generalizability of Prognostic Information
1,170 Citations1999Amy C. Justice, Kenneth E. Covinsky +1 more
JAMAMortality Probability Models (MPM II) Based on an International Cohort of Intensive Care Unit Patients
1,000 Citations1993Stanley Lemeshow
Among severity systems for intensive care patients, the MPM0 is the only model available for use at ICU admission and bothMPM0 and MPM24 are useful research tools and provide important clinical information when used alone or together.
Journal of Clinical EpidemiologyExternal validation is necessary in prediction research:
697 Citations2003Sacha E. Bleeker, Henriëtte A. Moll +5 more
For relatively small data sets, internal validation of prediction models by bootstrap techniques may not be sufficient and indicative for the model's performance in future patients.
Journal of Clinical EpidemiologyInternal and external validation of predictive models: A simulation study of bias and precision in small samples
578 Citations2003Ewout W. Steyerberg, Sacha E. Bleeker +3 more
Bootstrapping for internal validation gives reasonably valid estimates of the expected optimism in predictive performance provided that any selection of predictors is taken into account, and should be used for external validation.
Statistics in MedicineValidation and updating of predictive logistic regression models: a study on sample size and shrinkage
549 Citations2004Ewout W. Steyerberg, Gerard Borsboom +3 more
A logistic regression model may be used to provide predictions of outcome for individual patients at another centre than where the model was developed to improve predictions for future patients.
Statistics in MedicineValidation techniques for logistic regression models
319 Citations1991Michael E. Miller, Siu L. Hui +1 more
A model-based approach developed by Cox is adapted for use in model validation, which allows identification of problematic predictor variables in the prediction model as well as influential observations in the validation data that adversely affect the fit of the model.
Statistics in MedicineValidation, calibration, revision and combination of prognostic survival models
272 Citations2000Hans C. van Houwelingen
Methods are sketched to perform validation through 'calibration', that is by embedding the literature model in a larger calibration model, for x-year survival probabilities, Cox regression and general non-proportional hazards models.
Critical Care MedicineEvaluation of Acute Physiology and Chronic Health Evaluation III predictions of hospital mortality in an independent database
265 Citations1998Jack E. Zimmerman, Douglas P. Wagner +4 more
APACHE III accurately predicted aggregate hospital mortality in an independent sample of U.S. ICU admissions and further improvements in calibration can be achieved by more precise disease labeling, improved acquisition and weighting of neurologic abnormalities, adjustments that reflect changes in treatment outcomes over time, and a larger national database.
Intensive Care MedicineOutcome prediction in intensive care: results of a prospective, multicentre, Portuguese study
165 Citations1997Rui P. Moreno, P. Morais
SAPS II performed better than APACHE II in this independent database, but the results do not allow its use, at least without being customised, to analyse quality of care or performance among ICUs in the target population.
BMJABC of intensive care: Outcome data and scoring systems
153 Citations1999K. Gunning, Kathy Rowan
Scoring systems have been developed in response to an increasing emphasis on the evaluation and monitoring of health services to enable comparative audit and evaluative research of intensive care.
Critical Care MedicineFactors affecting the performance of the models in the Mortality Probability Model II system and strategies of customization
149 Citations1996Bao‐Ping Zhu, Stanley Lemeshow +4 more
Mortality Probability Model II models can be used to assess quality of care in ICUs, but the size of the sample should be considered when assessing calibration and discrimination, indicating that they are useful quality assurance tools.
Intensive Care MedicineQuality of data collected for severity of illness scores in the Dutch National Intensive Care Evaluation (NICE) registry
135 Citations2002Daniëlle G. T. Arts, Nicolette F. de Keizer +2 more
The current data quality of the NICE registry is good and justifies evaluative research, and positive results might be explained by the implementation of several quality assurance procedures in the Nice registry, such as training and automatic data checks.
Journal of Clinical EpidemiologyExternal validity of predictive models: a comparison of logistic regression, classification trees, and neural networks
131 Citations2003Norma Terrin, Christopher H. Schmid +3 more
A simulation study that compared the external validity of standard logistic regression with piecewise-linear and quadratic terms (LR2), classification trees, and neural networks (NNETs) highlights the necessity of external validation to test the transportability of predictive models.
JAMAMortality Probability Models (MPM II) based on an international cohort of intensive care unit patients
129 Citations1993Stanley Lemeshow
Methods of Information in MedicineTlie Measurement of Performance in Probabilistic Diagnosis
127 Citations1978Jørgen Hilden, J. Dik F. Habbema +1 more
The measurement of performance in probabilistic diagnosis using methods based on continuous functions of the diagnostic probabilities and its application in medicine.
Critical Care MedicineIntensive Care Societyʼs Acute Physiology and Chronic Health Evaluation (APACHE II) study in Britain and Ireland
124 Citations1994Kathy Rowan, John Kerr +4 more
APACHE II demonstrated a higher degree of overall goodness of fit, which was superior to MPM for groups of intensive care patients from Britain and Ireland, and even after modifications to theMPM for the assessment of coma, the performance of APACHE I was superior.
Critical Care MedicineImpact of different customization strategies in the performance of a general severity score
94 Citations1997Rui P. Moreno, Giovanni Apolone
In this ICU patient database, second-level customization was more effective than first- level customization in improving the overall goodness-of-fit of MPM II0 and should probably be chosen as the preferential strategy to improve the fit of a model when the sample size is large enough.
Intensive Care MedicinePrognostic performance and customization of the SAPS II: results of a multicenter Austrian study
80 Citations1999Philipp Metnitz, Andreas Valentin +6 more
SAPS II was not well calibrated when applied to all patients, however, it performed well for patients with cardiovascular diseases as the primary reason for admission and may thus be applied to these patients.
Critical Care MedicineSimplified Acute Physiology Score II for measuring severity of illness in intermediate care units
76 Citations1998I. Auriant, Isabelle Vinatier +3 more
The SAPS II assessment of severity of illness in patients admitted to an intermediate care unit is reliable and the efficiency of severity scores has been established in ICU patients, but not in the setting of intermediate care units.
Critical Care MedicineComparison of Acute Physiology and Chronic Health Evaluation II (APACHE II) and Simplified Acute Physiology Score II (SAPS II) scoring systems in a single Greek intensive care unit
68 Citations2000Stylianos Katsaragakis, Konstantinos Papadimitropoulos +4 more
APACHE II and SAPS II failed to predict mortality in a population sample other than the one used for their development and performed better than APAC HE II but had good discriminative power, with APACHE I performing better than SAPS I.
Critical CareTraining in data definitions improves quality of intensive care data
39 Citations2003Daniëlle G. T. Arts, Rob J. Bosman +3 more
Training in data definitions and data extraction guidelines is an effective way to improve quality of intensive care scoring data.
Medical CarePredicting Outcome in the Intensive Care Unit Using Scoring Systems
26 Citations1998Guido Bertolini, Roberto D’Amico +6 more
