Pisces did not have increased heart failure: data-driven comparisons of binary proportions between levels of a categorical variable can result in incorrect statistical significance levels
Journal of Clinical EpidemiologyPublished 25 September 2007Open access
Peter C. Austin, Meredith A. Goldwasser
Citations11
SJR quartileQ1
SJR score3.15
SNIP2.66
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The impact on statistical inference when a chi(2) test is used to compare the proportion of successes in the level of a categorical variable that has the highest observed proportion of success with the proportionof successes in all other levels of the categoricalVariable combined is examined.
Abstract
Post hoc comparisons of the proportions of successes across different levels of a categorical variable can result in incorrect inferences.
Keywords
Mathematics
Wiley series in probability and statisticsAn Introduction to Categorical Data Analysis
6,615 Citations2006Alan Agresti
The American StatisticianBootstrap Methods for Developing Predictive Models
620 Citations2004Peter C. Austin, Jack V. Tu
Using data on patients admitted to hospital with a heart attack, it is demonstrated that selecting those variables that were identified as independent predictors of mortality in at least 60%% of the bootstrap samples resulted in a parsimonious model with excellent predictive ability.
Annual Review of SociologyAn Introduction to Categorical Data Analysis
400 Citations1996Douglas M. Sloane, S. Philip Morgan
This paper reviews the basic log-linear strategy and illustrates key concepts and Citations are given to other articles on these topics, many of which are nontechnical and contain substantive sociological applications.
Journal of Clinical EpidemiologyAutomated variable selection methods for logistic regression produced unstable models for predicting acute myocardial infarction mortality
336 Citations2004Peter C. Austin, Jack V. Tu
The reproducibility of logistic regression models developed using automated variable selection methods are determined to be unstable and not reproducible because the variables selected as independent predictors are sensitive to random fluctuations in the data.
The American StatisticianThe Impact of Model Selection on Inference in Linear Regression
276 Citations1990Clifford M. Hurvich, Chih‐Ling Tsai
Statistics in MedicineInflation of the type I error rate when a continuous confounding variable is categorized in logistic regression analyses
248 Citations2004Peter C. Austin, Lawrence J. Brunner
It is found that the inflation of the type I error rate increases with increasing sample size, as the correlation between the risk factor and the confounding variable increases, and with a decrease in the number of categories into which the confounder is divided.
Journal of Clinical EpidemiologyTesting multiple statistical hypotheses resulted in spurious associations: a study of astrological signs and health
135 Citations2006Peter C. Austin, Muhammad Mamdani +2 more
A study of residents of Ontario to illustrate how the testing of multiple, non-prespecified hypotheses increases the likelihood of detecting implausible associations and has important implications for the analysis and interpretation of clinical studies.
The American StatisticianThe Impact of Model Selection on Inference in Linear Regression
81 Citations1990Clifford M. Hurvich, Chih‐Ling Tsai
Biometrical JournalMaximally Selected Chi-square Statistics for Ordinal Variables
39 Citations2006Anne‐Laure Boulesteix
This paper suggests an exact method to determine the finite-sample distribution of maximally selected chi-square statistics in this context and applies this method to a new data set describing pregnancy and birth for 811 babies.
Biometrical JournalMaximally Selected Chi‐Square Statistics and Binary Splits of Nominal Variables
32 Citations2006Anne‐Laure Boulesteix
BiometricsMaximally Selected x<sup>2</sup> Statistics for <i>k</i>× 2 Tables
26 Citations1999Rebecca A. Betensky, Daniel Rabinowitz
The asymptotic distributions of maximally selected chi2 statistics for association and for trend for the k x 2 table are derived and the methodology is illustrated with data from an AIDS clinical trial.
