Skip to main content icon/video/no-internet

Instructional Sensitivity

A test item is deemed instructionally sensitive if, when controlling for other factors, students who receive high-quality instruction on the content of the item do better than students who have not received high-quality instruction. Although some specialized achievement tests, such as the National Assessment of Educational Progress, Trends in International Mathematics and Science Study, and Progress in International Reading Literacy Study, are designed to estimate the distribution of a state or nation’s scores rather than the scores of individuals, in general, academic achievement tests are designed to measure the knowledge of the individual students to whom the tests are administered. This entry discusses the increasing focus on instructional sensitivity as part of teacher evaluation systems that incorporate student test scores, data collection designs and data analysis approaches for determining instructional sensitivity, and research on instructional sensitivity.

In recent years, student test scores have been used as a measure of teacher effectiveness, in some cases as part of state accountability systems. Many have argued that this latter use requires an assumption that tests are sensitive to the instruction. That is, to attribute student success on achievement tests to teacher quality requires a belief that high-quality instruction reliably leads to higher student test scores than low-quality instruction and even more so when compared to test scores of students who did not receive in-school instruction at all. There is limited evidence for this claim.

Counterarguments have been made that the items most sensitive to instruction are likely to be those at low levels of cognitive complexity, for example, items measuring factual knowledge or low-level comprehension (per Bloom’s taxonomy). Items measuring higher levels of cognition, such as analysis, evaluation, and synthesis, are conjectured to be less instructionally sensitive due to their greater complexity. There is no empirical evidence to support this.

Although the term instructional sensitivity was first used in the 1970s, the concept that the quality of achievement test items should be judged by their ability to reflect improvement in student achievement following instruction stems from the origins of objective student achievement testing circa 1920. At that time, item quality was typically judged by comparing the item scores of students in the grade in which related curriculum was taught to the scores of students in the previous grade on the same item. That is, item quality was defined by the increase in the percentage of students who answered correctly in the grade at which instruction on the topic occurred to the percentage of students in the previous grade who responded correctly.

Data Collection Designs for Determining Instructional Sensitivity

Methods of detecting the instructional sensitivity of test items can be divided into four broad data collection designs based on (1) item data from two representative groups, one that was exposed to the content and one that was not, (2) pretest–posttest administration to the same group, (3) item data from a single group where some had been exposed to the content and some had not, and (4) expert judgment.

As mentioned earlier, the use of item data from two representative groups has been practiced since the 1920s, before the use of the term instructional sensitivity or the use of test scores for evaluating educator quality. In such studies, item data are collected from all or a random sample of students in the grade for which the item is intended as well as in a random sample of students in the previous grade. Because the groups are representative of entire grade levels, it is usually reasonable to assume that most of the variability in item difficulty is due to instruction. Typically, item sensitivity is measured as the difference in percent correct between the two groups. However, percent correct has certain undesirable statistical characteristics. For example, the standard error of percent correct is dependent on the value of p. If p values are analyzed using least squares regression, this violates the assumption of heteroscedasticity.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading