Skip to main content icon/video/no-internet

R2

R2 (pronounced as “R squared”), or the coefficient of determination, is a statistical measure that is interpreted as the proportion of variance of the dependent variable (DV, or regressand) that is explained by the independent variable (IV, or predictor or explanatory variable) or by the statistical model. In other words, the R2 value gives the information on how well the IV explains the outcome variable (accounts for the variability of the DV).

The value of R2 is standardized (ranges from 0 to 1), which makes it easy to interpret; therefore, it is commonly used in statistical models for research (e.g., in education, psychology, biology, or economy). In social sciences, the coefficient of determination might be used in studies aimed at predicting school grades or estimating the outcome theoretical model that takes into account various related measurements. The R2 value indicates how well a model fits to the set of observations or the difference between observed and expected values. The higher the R2, the better the model’s goodness of fit.

How to Interpret the R2

Because R2 is standardized, its value varies between 0 and 1. After multiplying that value by 100%, you obtain the percentage of variability explained by the IV or the statistical model. If the R2 value is equal to 0, you know that the IV does not explain the variance of DV at all (0 × 100% = 0% variability explained), whereas a value of 1 would mean that the variance of outcome variable is fully explained by the IV (1 × 100% = 100% variability explained). However, models that explain 100% of DV variance rarely happen.

For example, let us say you would like to predict the educational success of college students in one of the courses they attended. To assess it, you perform a research. You measure each student’s IQ, which accounts for intellectual abilities. As an outcome measure—DV—you measure each student’s grade at the end of the school year. After collecting data from the group of participants, you run a regression analysis using a statistical software program, in which the student’s grade is an outcome (dependent) variable and the IQ is an IV. The statistical program prints the output of the analysis, which provides you with information about your model. The R2 value is .5. What does this mean?

The IV explains 50% (0.5 × 100% = 50%) of variance of the student’s grade. But you might ask, where is the lacking 50% of explained variance? The 50% of the variance that this model did not account for are the variables that were not included in the model, such as student’s commitment to studying, socioeconomic status, or interest in the course.

Low R2 Versus High R2

A question that often arises in the context of estimating goodness of fit of the model is how does one know if the value on the output is “high enough.” The answer depends on the research area. Usually, social scientists obtain lower R2 values than biologists do due to the fact that behavior depends on a complex set of variables and the fact that social scientists fail to control all of them.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading