Scinovex
reviewTop 1% cited

Estimating Causal Effects from Large Data Sets Using Propensity Scores

Annals of Internal Medicine · 1997 · Vol. 127(8_Part_2) · pp. 757–763
Donald B. Rubin

Abstract

The aim of many analyses of large databases is to draw causal inferences about the effects of actions, treatments, or interventions. Examples include the effects of various options available to a physician for treating a particular patient, the relative efficacies of various health care providers, and the consequences of implementing a new national health care policy. A complication of using large databases to achieve such aims is that their data are almost always observational rather than experimental. That is, the data in most large data sets are not based on the results of carefully conducted randomized clinical trials, but rather represent data collected through the observation of systems as they operate in normal practice without any interventions implemented by randomized assignment rules. Such data are relatively inexpensive to obtain, however, and often do represent the spectrum of medical practice better than the settings of randomized experiments. Consequently, it is sensible to try to estimate the effects of treatments from such large data sets, even if only to help design a new randomized experiment or shed light on the generalizability of results from existing randomized experiments. However, standard methods of analysis using available statistical software (such as linear or logistic regression) can be deceptive for these objectives because they provide no warnings about their propriety. Propensity score methods are more reliable tools for addressing such objectives because the assumptions needed to make their answers appropriate are more assessable and transparent to the investigator.

Advanced Causal Inference TechniquesHealth Systems, Economic Evaluations, Quality of LifeHealthcare Policy and ManagementObservational studyRandomized controlled trialMedicineGeneralizability theoryPsychological interventionPropensity score matchingHealth careCausal inferenceRandomized experimentData mining

MeSH terms

Breast NeoplasmsFemaleHumansResearch DesignSmokingConfounding Factors, EpidemiologicLinear ModelsSurvival AnalysisDatabases, Factual
Citations
2,895
FWCI
19.21
field-weighted impact
References
30
Percentile
100%
vs. same field & year
Citations per year
Cited by
Comparing apples and oranges
Journal of Thoracic and Cardiovascular Surgery · 2002 · 601 citations
Simpson's paradox in psychological science: a practical guide
Frontiers in Psychology · 2013 · 524 citations
Neighborhoods and health
Annals of the New York Academy of Sciences · 2010 · 2,701 citations
Variable Selection for Propensity Score Models
American Journal of Epidemiology · 2006 · 2,276 citations
Two internal thoracic artery grafts are better than one
Journal of Thoracic and Cardiovascular Surgery · 1999 · 917 citations
Recurrent mitral regurgitation after annuloplasty for functional ischemic mitral regurgitation
Journal of Thoracic and Cardiovascular Surgery · 2004 · 620 citations
References
Assessing Sensitivity to an Unobserved Binary Covariate in an Observational Study with Binary Outcome
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1983 · 1,119 citations
Reducing Bias in Observational Studies Using Subclassification on the Propensity Score
Journal of the American Statistical Association · 1984 · 3,240 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

Estimating Causal Effects from Large Data Sets Using Propensity Scores · Scinovex