Skip to main content icon/video/no-internet

Fisher Exact Test

The Fisher exact test is an inferential statistical procedure to compare the number of people or things falling into different categories. It is applicable in two situations. The first one is when a sample is drawn from a population and two categorical variables are recorded for each element in the sample, for example, political affiliation (Democrat/Republican/Other) and potential vote on a certain proposal (in favor/opposed/abstained). In this case, researchers would be testing whether there is an association between these two variables (or, putting it more rigorously, whether the two variables are independent). The second situation arises when two or more samples are drawn from independent populations and measurements for one categorical variable are recorded for each sampled element. In this instance, the hypothesis of interest is whether proportions for each level of the categorical variable are equal across the samples. To illustrate, a sample of freshmen and a sample of seniors are drawn and students’ employment status (unemployed/part-time/full-time) is recorded. Investigators would be interested in testing whether proportions of students in each category of the employment status differ between freshmen and seniors. After this entry further explores the fundamental attributes of the Fisher exact test, it examines how statistical hypotheses are formulated and the procedure for conducting the test. Next, examples of a Fisher exact test for independence and test for equality of proportions are provided. Finally, limitations of the Fisher exact test are discussed.

In preparation for conducting the Fisher exact test, observations are arranged in an r by c table called a contingency table. It may also be called a two-way table or cross tabulation or, simply, cross tab. In the former case, when a single sample is drawn from one population and two categorical variables with r and c levels, respectively, are observed for each unit in the sample, the r rows of the contingency table correspond to the levels of the first variable, whereas the c columns represent the levels of the second variable. In the latter situation, when r samples are drawn from independent populations and a categorical variable with c levels is observed for each sample element, in the contingency table, the r rows represent the samples, and the c columns contain frequencies of the c levels of the observed variable.

Each cell in the contingency table contains the frequency of observations in the corresponding level–level combination of the two observed variables (in the former situation) and in the corresponding sample at the certain level of the observed variable (in the latter situation). These frequencies are commonly referred to as observed counts.

In order to prepare the data for analysis, the marginal totals must be computed and added to the contingency table. They are defined as the total for each row and column. The row totals are put in an additional column on the right of the table, whereas the column totals go into the row added to the bottom of the table. As the name suggests, these totals are placed on the “margins” of the table. Next, the grand total is calculated and written below the column with row totals (or to the right of the row of column totals). The grand total is defined as the sum of all observed counts. It is also equal to the sum of row totals and, likewise, to the sum of column totals.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading