Skip to main content icon/video/no-internet

Ordinary Least Squares

Two basic approaches to statistical analysis included in many statistical packages are ordinary least squares (OLS) and maximum likelihood estimation. This entry considers the implications and practices involved in OLS. OLS employs a procedure most often associated with typical statistical procedures and corresponds to many common techniques in use (correlation, t-test, mean). The term least squares corresponds to the idea that the best value of estimation involves a parameter that minimizes the value of the sum of squared error or deviation. For example, one definition of the mean involves the value for a set of data where the sum of squared deviation from that value is the smallest. The term or estimate often becomes evaluated as the best fit or most accurate representation available in the analysis.

Least Square Approach to Analysis

The least square approach provides a set of assumptions for the estimation of statistical parameters, like the estimation of a mean. The mean, also known as expected value, constitutes the value where the sum of squared error, Ʃ(XM)2, is the smallest for that set of data, where X is the raw score and M is the arithmetic mean, Ʃ(X/N), where N is the number of scores in the analysis. The impact of this set of assumptions and mathematical operations provides the basis for most of the assumptions and practices employed in statistical analysis.

For example, the estimations involved in the independent groups t-test involve a term in the denominator that is the weighted average of the variances:

t = M 1 M 2 ( s 1 2 [ n 1 1 ] + s 2 2 [ n 2 1 ] ) / ( n 1 + n 2 2 ) ( n 1 + n 2 ) / ( n 1 × n 2 ) ,

where s indicates the variance for each variable and each sample size is designated by the term n, the means for each group are indicated by M (M1 and M2) and the variance or sum of squared error is represented by the term s ( s12 and s22, one calculated for each group). The sample size for each of the two groups is represented by n (n1 and n2). The estimate essentially examines the difference between groups, indicated in the numerator, versus the within-group variability, weighted by sample size.

The same can be applied to the correlation coefficient implied by the formula used by Pearson. The formula for the correlation coefficient provides for a numerator that generates the covariance between two variables compared with the level of variability for both

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading