Skip to main content icon/video/no-internet

Item Response Theory

Item response theory (IRT) is a measurement framework for the development of tests and the scoring of item responses on tests. Key aspects of the IRT framework include the focus on items as the units of observed measurement, the fitting of parametric statistical models to categorical item response data, the estimation of a latent trait variable, and the conditional nature of reliability and the standard error of measurement. IRT is relevant to the field of educational measurement in that it is widely used by measurement practitioners for developing and scoring standardized tests, such as the SAT, ACT, and statewide K–12 achievement tests in mathematics, reading, and other academic domains. IRT is also widely used for the scoring of scale data, such as attitude scales consisting of a series of Likert-type items. In educational research, IRT is most often used for developing measurement tools, evaluating reliability of test data, and estimating latent abilities that can then be used as variables in research studies. The remainder of this entry provides details for the conceptual understanding of some selected statistical models in the IRT family, reliability and error under the IRT framework, and the relationship of IRT to other measurement theories.

Statistical Models in the Item Response Theory Family

IRT encompasses a family of parametric statistical models that are fit to item response data to estimate scores on a latent trait variable. Many models within the family fall into one of the two types: those that are built for dichotomous item responses (i.e., the data for each item can take on only two values) and those that are built for polytomous item responses (i.e., the data for each item can take on three or more values). Fitting an IRT model to a data set of item responses produces, among other things, a series of item characteristic curves (ICCs) that relate the underlying latent trait variable to the probability of scoring in each of the item response categories. Figure 1 shows examples of ICCs for several dichotomously and polytomously scored items.

Figure 1 Item Characteristic Curves

Figure

Conceptually, the ICC is the statistically estimated answer to the measurement question: “What is the probability of an examinee providing a particular response to this test item, given that the examinee has a particular trait level?” The ICC reflects the nature of all IRT model specifications in that the latent trait variable is treated as an independent variable that has a functional relationship to the multiple dependent variables of observed item responses. For dichotomously scored items, the ICC is displayed as an S-shaped curve representing the conditional probability of scoring a 1 on the item as opposed to a 0. The curve representing the conditional probability of scoring a 0 is not shown because it is redundant information (i.e., the probability of a 0 response is one minus the probability of a 1 response). For polytomously scored items, the ICC is displayed as a series of conditional probability curves, one for each response category on the item. At any point on the latent trait variable, the sum of the probabilities of scoring in each of the possible response categories is 1.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading