login

Risk set sampling designs for proportional hazard models

Duo Research Archive (University of Oslo)Published 1 January 1997Open access
Ørnulf Borgan, Bryan Langholz
Citations4
View PDF

Abstract

Cox's regression model (Cox, 1972) and similar proportional hazards models are central to modern survival analysis, and they are the methods of choice when one wants to assess the in uence of risk factors and other covariates on mortality or morbidity. Estimation in such proportional hazards models is based on Cox's partial likelihood [see (2) below], which at each observed death or disease occurrence (failure) compares the covariate values of the failing individual to those of all individuals at risk at the time of the failure. In large epidemiologic cohort studies of a rare disease, (standard) use of proportional hazards models requires collection of covariate information on all individuals in the cohort even though only a small fraction of these actually get diseased. This may be very expensive, or even logistically impossible. Cohort sampling techniques, where covariate information is collected for all failing individuals (cases), but only for a sample of the non-failing individuals (controls) then o er useful alternatives which may drastically reduce the resources that need to be allocated to a study. Further, as most of the statistical information is contained in the cases, such studies may still be suAEcient to give reliable answers to the questions of interest. The most common cohort sampling design is nested case-control sampling, where for each case a small number of controls are selected at random from those at risk at the case's failure time, and where a new sample of controls is selected for each case. This risk set sampling technique was rst suggested by Thomas (1977), who proposed to base inference on a modi cation of Cox's partial likelihood. This suggestion was supported by the work of Prentice and Breslow (1978), who derived the same expression as a conditional likelihood for time-matched case-control sampling from an in nite population. A more decisive, but still heuristic, argument was provided by Oakes (1981), who showed that one indeed gets a partial likelihood when the sampling of controls is performed within the actual nite cohort. It took more than ten years, however, before Goldstein and Langholz (1992) proved rigorously that the estimator of the regression coeAEcients based on Oakes' partial likelihood enjoys similar large sample properties as ordinary maximum likelihood estimators. Goldstein and Langholz's paper initiated further work on risk set sampling methodology, and important progress has been achieved during the last few years both with respect to its theoretical foundation and the development of new methodology of practical importance. The

Keywords

Mathematics