Entry
Reader's guide
Entries A-Z
R-Squared
The R-squared (R2) measures the explanatory or predictive power of a regression model. It is a GOODNESS-OF-FIT MEASURE, indicating how well the linear regression equation fits the data.
In regression analysis, it is important to evaluate the performance of the estimated regression equation. To what extent does it account for the phenomenon under study? The R2 is the leading performance measure for a simple or MULTIPLE REGRESSION model. Suppose a policy analyst is studying public school expenditures in the 50 American states. The analyst posits a simple regression model, Y =a +bX +e, where Y is the dependent variable of per-pupil public school expenditures (in thousands of dollars) in each state in the year 2000, X is the INDEPENDENT VARIABLE of urbanization (percentage of population in cities larger than 25,000, as of the 2000 census), and e is the error term. The model argues that state public school outlays are accounted for, in part, by urbanization. ORDINARY LEAST SQUARES (ols) estimation yields, hypothetically, the following results:Y =2.11 +.60X +e, suggesting that for every additional percentage of urban population, a state expects to spend $600 more. The R2, or coefficient of determination, for the equation is .42, which says that 42% of the variation in state public school expenditures is explained, or at least predicted, by urbanization.
The R2 shows the gain in predicting Y knowing X, as opposed to not knowing X. Suppose that the analyst knows the scores on Y, but not the states to which they are attached, and tries to predict school expenditures state by state. The best guess, the one that minimizes the error, would always be the average of Y, that is, Ym. This guess will be way off for most states, with a great distance from the observed expenditure score to the average expenditure score, that is, (Y −Ym). Adding all these distances together (after squaring to overcome the canceling out of the plus and minus signs), they represent the total prediction error not knowing X; that is, total sum of squared deviations (TSS) =(Y −Ym)(^Y =, is the improvement over the baseline prediction made possible by knowing X. This distance, squared and summed for all observations, is the reduction in prediction error attributable to application of the regression line; that is, regression sum of squared deviations (RSS) =(^2.
Regression analysis promises reduction of this prediction error through knowledge of X and its linear relationship to Y. For example, knowing a state’s score on X is 30% urban, the prediction, ^Y, is 20.21 2.11 + 18.10 = 20.21), not Ym. The distance from the predicted Y to the mean Y, (^Y −Ym)Y −Ym)2. Unless the regression line predicts each case perfectly, some error will still remain: the distance from the observed Y and the predicted Y. These distances, squared then summed, represent the variation in Y that remains unexplained; that is, error sum of squared deviations (ESS) =(Y − ^Y)2.
Total variation in the dependent variable, TSS, thus has two unique parts: RSS accounted for by the regression and ESS not accounted for by the regression. The R2 reflects the regression portion as a share of the total, RSS/TSS =R2. The statistic ranges from 1.0, when all variation is accounted for, to .00, when no variation is accounted for. With real-world data, the R2 seldom reaches these extreme values. In the above example, we may say that 42% of the variation in public school expenditures is explained by urbanization. It is important to remember that this explanation may be more “statistical” than “causal.” It could be that urbanization helps predict public school expenditure but does not really explain it in a theoretical sense. In such cases, it is more cautious, and perhaps more correct, to say that X merely “accounts for” so much variation in Y. In the bivariate regression case, the R2 equals the correlation coefficient squared (the r2).
...
- Analysis of Variance
- Association and Correlation
- Association
- Association Model
- Asymmetric Measures
- Biserial Correlation
- Canonical Correlation Analysis
- Correlation
- Correspondence Analysis
- Intraclass Correlation
- Multiple Correlation
- Part Correlation
- Partial Correlation
- Pearson's Correlation Coefficient
- Semipartial Correlation
- Simple Correlation (Regression)
- Spearman Correlation Coefficient
- Strength of Association
- Symmetric Measures
- Basic Qualitative Research
- Basic Statistics
- F Ratio
- N(n)
- t-Test
- X¯
- Y Variable
- z-Test
- Alternative Hypothesis
- Average
- Bar Graph
- Bell-Shaped Curve
- Bimodal
- Case
- Causal Modeling
- Cell
- Covariance
- Cumulative Frequency Polygon
- Data
- Dependent Variable
- Dispersion
- Exploratory Data Analysis
- Frequency Distribution
- Histogram
- Hypothesis
- Independent Variable
- Measures of Central Tendency
- Median
- Null Hypothesis
- Pie Chart
- Regression
- Standard Deviation
- Statistic
- Causal Modeling
- DISCOURSE/CONVERSATION ANALYSIS
- Econometrics
- Epistemology
- Ethnography
- Evaluation
- Event History Analysis
- Experimental Design
- Factor Analysis and Related Techniques
- Feminist Methodology
- Generalized Linear Models
- HISTORICAL/COMPARATIVE
- Interviewing in Qualitative Research
- Latent Variable Model
- LIFE HISTORY/BIOGRAPHY
- LOG-LINEAR MODELS (CATEGORICAL DEPENDENT VARIABLES)
- Longitudinal Analysis
- Mathematics and Formal Models
- Measurement Level
- Measurement Testing and Classification
- Multilevel Analysis
- Multiple Regression
- Qualitative Data Analysis
- Sampling in Qualitative Research
- Sampling in Surveys
- Scaling
- Significance Testing
- Simple Regression
- Survey Design
- Time Series
- ARIMA
- Box-Jenkins Modeling
- Cointegration
- Detrending
- Durbin-Watson Statistic
- Error Correction Models
- Forecasting
- Granger Causality
- Interrupted Time-Series Design
- Intervention Analysis
- Lag Structure
- Moving Average
- Periodicity
- Serial Correlation
- Spectral Analysis
- Time-Series Cross-Section (TSCS) Models
- Time-Series Data (Analysis/Design)
- Trend Analysis
Get a 30 day FREE TRIAL
-
Watch videos from a variety of sources bringing classroom topics to life
-
Read modern, diverse business cases
-
Explore hundreds of books and reference titles
Sage Recommends
We found other relevant content for you on other Sage platforms.
Have you created a personal profile? Login or create a profile so that you can save clips, playlists and searches