Skip to main content icon/video/no-internet

Test Bias

Test bias is one of the most important issues in the development of measures. However, it is often confused with fairness. On one hand, fairness is a social concept that is concerned with whether one views test scores as being used in an appropriate manner. There is no right way to examine whether test scores are used appropriately, as this is based on an individual’s subjective perception. On the other hand, bias is viewed as a statistical issue and is concerned about whether there is systematic error in measuring a trait or attribute across groups. If systematic differences due to group membership on a test exist, this would suggest that bias is present in the test. This article outlines several methods used to examine test bias.

Differential Item Functioning (DIF)

DIF is a method to determine whether a measure (e.g., a personality, intelligence, or academic achievement measure) is equivalent across groups. DIF occurs when an item on a measure is responded to differently by individuals in different groups, such as different gender, age, or ethnic groups, who have the same amount of an attribute or a latent trait. The latent trait or attribute could be an ability, skill, or personality characteristic. If DIF exists on an item, then the item may be biased. DIF is important to investigate because group comparisons, such as age, gender, or ethnic differences, on a measure cannot be made unless the items are found to be equivalent across the groups of interest. Equivalence of items across groups should be examined when one develops new instruments or it can be examined with measures already existing in the field.

There are two types of DIF: uniform DIF and nonuniform DIF. Uniform DIF is when the probability of a specific response to an item (e.g., the probability of endorsing a yes response on an item) is higher for one group than for another group (e.g., females than males) at each level of the attribute (e.g., anxiety) that is being measured. In contrast, nonuniform DIF is when the probability of a specific response to an item differs at different levels of the attribute. For example, males may be more likely to endorse a no response on an item at lower levels of an attribute but are more likely to endorse a yes response on the item at higher levels of the attribute.

Different procedures exist for detecting DIF. Some of these approaches are nonparametric and others are parametric. One of the most common nonparametric approaches for detecting DIF is the Mantel-Haenszel method. The Mantel-Haenszel is a contingency table-based approach that uses odds ratios to determine whether one group outperforms the other group on each of the items. If a common odds ratio indicates that one group outperforms the other group across all levels of the trait or attribute for a specific item, then DIF is said to be present for that item. Besides the nonparametric approach, there are two common parametric methods to detect DIF: the logistic regression and item response theory (IRT) approaches. Hariharan Swaminathan and H. Jane Rogers indicate that nested models can be compared in the logistic regression approach. One model, referred to as the augmented model, that includes the group (e.g., gender), the trait (e.g., anxiety), and the interaction between the group and the trait variable is tested against another model, referred to as the compact model, that includes the group and the trait variable, but not the interaction term. Jeanne A. Teresi and John A. Fleishman assert that when the augmented and compact models are estimated using the maximum likelihood parameter estimator, a likelihood value is obtained for the augmented and the compact models, and the difference in the log-likelihood values between these models is examined using a chi-square test. If the chi-square test is significant, indicating a significant group by trait interaction, then nonuniform DIF is present. If the chi-square test is not significant, then the compact model is compared to a model where no group effect is assumed. If the chi-square test is significant, indicating a group effect exists, then uniform DIF is present. This procedure is repeated for items on the measure. However, it should be noted that variations do exist in the logistic regression approach to detect DIF. IRT is another method that can be used to detect DIF, and there are variations in this approach too. In IRT, the item characteristic curve, an S-shaped curve, represents the graphic relationship between the probability of giving a certain response on an item on a measure and an individual’s position on the latent trait continuum. The shape of the curve is determined by its parameters (discrimination, difficulty, and if applicable, guessing). IRT models are derived from these parameters, including one-, two-, and three-parameter models. To detect DIF, the item characteristic curves of two groups are compared and if one of the parameters is different for the two groups, then DIF is likely to be present. Different statistical tests, such as a likelihood ratio test, or magnitude measures are used to determine the salience of DIF.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading