articleJournal of Clinical EpidemiologyJan 22, 2015HYBRID OA

The number of subjects per variable required in linear regression analyses

Sunnybrook Research Institute · Sunnybrook Health Science Centre · +3 more institutions

PubMed
Indexed incrossrefpubmed

Abstract

Objectives

To determine the number of independent variables that can be included in a linear regression model. STUDY DESIGN AND SETTING: We used a series of Monte Carlo simulations to examine the impact of the number of subjects per variable (SPV) on the accuracy of estimated regression coefficients and standard errors, on the empirical coverage of estimated confidence intervals, and on the accuracy of the estimated R(2) of the fitted model.

Results

A minimum of approximately two SPV tended to result in estimation of regression coefficients with relative bias of less than 10%. Furthermore, with this minimum number of SPV, the standard errors of the regression coefficients were accurately estimated and estimated confidence intervals had approximately the advertised coverage rates. A much higher number of SPV were necessary to minimize bias in estimating the model R(2), although adjusted R(2) estimates behaved well. The bias in estimating the model R(2) statistic was inversely proportional to the magnitude of the proportion of variation explained by the population regression model.

Citation impact

1,030
total citations
FWCI
72.03
Percentile
100%
References
22
Citations per year

Authors

2

Topics & keywords

Keywords
  • Statistics
  • Linear regression
  • Confidence interval
  • Mathematics
  • Regression analysis
  • Regression dilution
  • Standard error
  • Regression
No related works found for this paper.