Skip to main content icon/video/no-internet

One of the standard assumptions of the classical linear regression model

y i = β 0 + β 1 x i 1 + β 2 x i 2 + + β k x i k + ε ; i = 1 N

is that the variance of the error term (εi) is the same for all observations, that is Var(|x1i,x2i,,xki)=σ2. The assumption of a constant error variance is known as homoskedasticity and its failure is referred to as heteroskedasticity, or unequal variance. Heteroskedasticity is expressed as Var(|x1i,x2i,,xki)=σi2, where an i subscript on σ2 indicates that the variance of the error is no longer constant but may vary from observation to observation. Note, the alternate spellings homoscedasticity and heteroscedasticity are also commonly used. This entry first reviews when heteroskedasticity typically arises. The consequences of heteroskedastic errors are then discussed, followed by sections describing the detection of and solutions for heteroskedasticity.

Heteroskedasticity is often encountered when using cross-section data, when observations are made at a given point in time. The data may include units of observation such as individuals, families, firms, industries, cities, or countries where the observations may be subject to wide variations such as low-, medium-, and high-income families. For example, if households are surveyed to determine how their income influences their consumption expenditures, we would expect less variation in spending patterns for low-income households compared to wealthy households. Low-income households spend their income on necessities such as food, clothing, and rent. Wealthy households have more choice in how they spend their income and can afford to spend money on luxury items with some choosing to do so and others not. Consequently, wealthy households have a larger dispersion around average consumption than low-income households. Heteroskedasticity may also arise when grouped data are used rather than individual data. An example is when data on industry averages is used because data on individual firms are not available. An additional case that results in heteroskedastic errors would occur when the dependent variable is a qualitative or binary variable. This type of model is known as the linear probability model.

Consequences

Under all the assumptions of the classical linear regression model, ordinary least squares (OLS) estimators are best linear unbiased estimators (BLUE). That is, within the class of linear unbiased estimators, OLS have minimum variance. Consequently, it is the most efficient estimator. It is also a consistent estimator, which means as the sample size gets larger and larger, it approaches the true value of the coefficient. If errors are heteroskedastic, the OLS estimators are still unbiased and consistent; however, they are now inefficient. This means they no longer have minimum variance. Furthermore, the usual estimated variances and covariances of the OLS estimates are generally biased. This leads to invalid hypothesis tests, and confidence intervals and forecasts will also be inefficient.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading