Skip to main content icon/video/no-internet

Residuals

Residual is a term with several closely related meanings that occur in the fields of business, finance, optimization, and statistics. Here, it will be only considered as a statistical term. In statistics, residuals are differences between observed values and values predicted on the basis of a statistical model.

This entry presents what residuals are in the examples of continuous (regression analysis) and categorical (contingency tables analysis) data. It is important to understand what residuals stand for, and what their properties are, as they are part of every statistical model used in educational research. Moreover, a wide repertoire of statistical models has specific assumptions about distribution of residuals that has to be met, so that the model represents unbiased relationships between variables. Besides the diagnostic of a statistical model, residuals can be used to identify unusual observations (anomalies) or to indicate which category occurs less often or more often than expected.

Residuals should not be confused with statistical error, which is the amount by which observations are different from their expected value based on the whole population (this quantity cannot be observed directly). Residuals refer to the amount by which observations are different from the sample mean. Therefore, residuals usually are treated as the estimates of statistical error.

Common Types of Residuals

Ordinary residuals are expressed on the scale of the variable for which they are being computed. Let us assume that a person’s height (195 cm) residual is to be computed. The expected value of the height for each person in the sample is equal to the mean height of the sample (170 cm, standard deviation equals 8). That means that the residual is equal to 25 cm (195 – 170 = 25). Frequently, it is more convenient to use another type of residual, depending on the purpose of analysis.

Standardized residuals are raw residuals transformed to the so-called standard score (also called z score) and are useful in identifying observations that are not typical. Whenever a value of the variable for which residuals are being computed comes from normal distribution, standardized residuals inform whether the observation is usual or not. If the value of a standardized residual is below −2.58 or above 2.58, such an observation is treated as an anomaly. It means that observations with such a residual represent not more than 1% of the population. A standardized residual of a person’s height is equal to 3.12 (25 cm divided by 8 cm—the length of a standard deviation). A person with the height of 195 cm appears rarely in the population (assuming that values of height come from a normal distribution).

Studentized residuals are especially useful in cases of multiple linear regression. In contrast to standardized residuals, they are more robust for anomalies.

Residuals in Linear Regression

Let us see what residuals look like in the case of a simple linear regression.

Figure 1 is a typical example of a positive correlation between two variables. Values of variable Y can be expressed in a regression model as a linear combination of intercept (which in this case is equal to zero) and slope times variable X and residuals (Y = intercept + slope × X + E, residuals are usually marked by E).

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading