Skip to main content icon/video/no-internet

Chi-Square Test

The chi-square test refers to a family of statistical tests that have been utilized to determine whether the observed (sampling) distribution or outcome differs significantly from an a priori or theoretically anticipated outcome or distribution. More simply stated, the test is formulated to determine whether the difference observed was due to a chance occurrence. This entry further describes the chi-square test and looks at its basic principles, applications, and limitations.

Although the most common chi-square test statistic is Pearson’s chi-square test, there are other test statistics that exist with the same theoretical foundation including Yates’s chi-square test, Tukey’s test of additivity, Cochran–Mantel–Haenszel test, and likelihood ratio tests. Although the chi-square test has been applied to a plethora of statistical applications, the fundamental utilization has been as a goodness-of-fit statistic and difference test by comparing a hypothesized distribution to an observed distribution.

In 1900, Karl Pearson developed the chi-square test and published his work entitled “On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling,” where he outlined the limitations of related measures and the functional utility of the test he developed. While the idea of determining whether standard distributions gave acceptable fits to data sets was well established early in Pearson’s career, detailed in his 1900 paper, he was determined to derive a test procedure that further advanced the problem of goodness of fit. As a result, the formulation of the chi-square statistic stands as one of the greatest statistical achievements of the 20th century.

Basic Principles and Applications

Generally speaking, a chi-square test (also commonly referred to as χ2) refers to a bevy of statistical hypothesis tests where the objective is to compare a sample distribution to a theorized distribution to confirm (or refute) a null hypothesis. Two important conditions that must exist for the chi-square test are independence and sample size or distribution. For independence, each case that contributes to the overall count or data set must be independent of all other cases that make up the overall count. Second, each particular scenario must have a specified number of cases within the data set to perform the analysis. The literature points to a number of arbitrary cutoffs for the overall sample size.

The chi-square test has most often been utilized in two types of comparison situations: a test of goodness of fit or a test of independence. One of the most common uses of the chi-square test is to determine whether a frequency data set can be adequately represented by a specified distribution function. More clearly, a chi-square test is appropriate when you are trying to determine whether sample data are consistent with a hypothesized distribution. The test includes the following procedures: Compute the chi-square statistic, determine the degrees of freedom, select the desired confidence level or p value, compare the chi-square value to the critical value in a chi-square distribution table, and decide to accept or reject the null hypothesis on the basis that the observed distribution differs from the theoretical distribution based upon whether the chi-square value exceeds or is less than the critical value.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading