Skip to main content icon/video/no-internet

Multicollinearity

Multicollinearity refers to the linear dependence among two or more variables. Although collinearity technically refers to the linear dependence among two variables, multicollinearity and collinearity are often used interchangeably. When there is a perfect linear dependence among predictors, statistical analyses such as multiple linear regression cannot be conducted with all included variables as the regression equation becomes unsolvable.

Consider a scenario where a researcher wants to know whether variability in test scores is a function of hair color (i.e., brown, black, red, and blond). Although an analysis of variance would likely be the statistical analysis of choice, the general linear model indicates that multiple linear regression could also be used. Of course, the variable hair color could not be used in its original form given its categorical state. Therefore, the researcher would have to dummy code hair color into several variables.

Let’s pretend that the researcher created a dichotomous variable for each hair color (0 = is not the color, 1 = is the color). If a regression was conducted where brown, black, red, and blond were used as independent variables, the solution would be inadmissible, as there is a linear dependence among the independent variables (e.g., blond is known, given brown, black, and red). However, if one of the hair colors (e.g., black) was left out of the equation and interpreted as the intercept, the regression analysis would run just fine. Full multicollinearity can also occur when variables are perfectly corrected (e.g., rX1.X2 = 1). Although it is hard to imagine valid scenarios in the education field where predictors are perfectly correlated, it is much more likely that predictors are correlated but not to the point of achieving a perfect correlation (i.e., partial multicollinearity).

Imagine, for example, that a researcher wanted to determine how measures of engagement for first-year teachers related to their intent to stay. Measures of engagement could include vigor, dedication, and absorption, and while the researcher would not expect the correlations among the variables to be perfect, it would be very likely that there would be substantive correlations (e.g., r > .5) among the predictors. In this scenario, the regression analyses would run just fine. However, the results would be more difficult to interpret, as the regression coefficients would not indicate measures of relationship. Further, just because one variable (e.g., X1) had a low regression coefficient would not mean that it did not have a substantive relationship with the dependent variable. When a variable (e.g., X1) is correlated with another predictor (e.g., X2), some of one variable’s (e.g., X1) credit to the regression effect may be captured by another variable’s (e.g., X2) regression coefficient.

Situations like the one just described have led some researchers to suggest that multicollinearity is a problem in multiple regression and other general linear model analyses. However, if researchers analyze regression weights along with other measures of importance including structure coefficients (or bivariate correlations) and commonality analysis, multicollinearity is not a problem, as these techniques identify the presence, loci, and magnitude of multicollinearity. In the presence of multicollinearity, researchers should look beyond regression weights and fully interpret regression effects. Such results will ensure that regression results are not only properly interpreted but can also be used to inform theory and practice.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading