Skip to main content icon/video/no-internet

Generalizability Theory

Educational researchers are often interested in making inferences from what may be considered observable to that which is unobserved. Responses to items on a multiple-choice test, an argumentative essay, and other overt behaviors (e.g., number of times a child raises her hand in a classroom) are observable. Unobserved variables on the other hand are used by educational researchers to explain patterns in observations. Intelligence, personality, aptitude, and critical thinking cannot, strictly speaking, be directly observed. Such variables refer to theoretical attributes that are at best indirectly investigated. For example, a researcher may hypothesize that differences in critical thinking (i.e., unobserved) account for why some students have higher scores than others on an assignment (i.e., observable). The extent to which it is reasonable to conclude that observed scores reflect critical thinking is a validity issue. Measurement error constrains the validity of score-based interpretations.

Measurement may be defined as the systematic assignment of numerals according to a set of rules. Measurements may distinguish mutually exclusive categories (e.g., ethnic groups), rank-order observations (e.g., high school class rank), or indicate differences in magnitude (e.g., degrees in Fahrenheit). Reliability assessment, traditionally conceived, aims to quantify the consistency of scores in a population whereas measurement error reflects random inconsistencies. Generalizability theory—hereafter referred to as G theory—provides a framework for investigating the extent to which distinct sources of error influence the precision of scores obtained from a measurement procedure.

The basic concepts of G theory, such as variance decomposition, universe scores, and facets of measurement, are introduced in this entry using a hypothetical example. This is followed by discussing simple extensions in measurement design employed within G theory, such as whether a facet is treated as fixed or random. Finally, the entry concludes by summarizing the strengths and limitations of G theory when compared to traditional approaches for assessing reliability.

Universe Scores and Facets of Measurement Error

Assume a researcher sampled thirteen students to assess their critical thinking. Each student has submitted two assignments with each assignment scored by the same two raters. Possible scores range from 0 to 4 with higher values indicating greater critical thinking (see Table 1). Students are considered the object of measurement because the researcher aims to use this procedure to differentiate students according to their level of critical thinking. Raters and assignments are sources of imprecision or error. For example, it is unlikely that each rater will provide the same score to a student for a single assignment. Even if raters perfectly agreed about scores for one assignment, it is unlikely that the student would receive the same critical thinking score across multiple assignments. Given such possibilities, what would be the best estimate of a student’s critical thinking?

Table 1 Student Critical Thinking Scores Assigned by the Same Two Raters Across Two Assignments

Student

Assignment 1

Assignment 2

Rater 1

Rater 2

Rater 1

Rater 2

1

0

1

1

2

2

3

4

1

2

3

2

2

1

1

4

2

2

0

1

5

1

2

2

1

6

4

4

3

4

7

1

1

2

1

8

3

3

1

1

9

1

1

3

4

10

1

1

1

0

11

1

2

1

1

12

1

2

1

1

13

2

1

1

0

Rater mean for each assignment (µra)

R1 at A1 1.6

R2 at A1 2.0

R1 at A2 1.6

R2 at A2 1.6

Assignment mean (µa)

A1 1.81

A2 1.77

Note: Critical thinking ranges from 0 to 4 with higher scores indicating more critical thinking. Values indicate ith person’s score provided by rth rater on the ath Assignment. R1 at A1 = mean of rater 1 on Assignment 1; R2 at A2 = mean of rater 2 on Assignment 2; R1 at A2 = mean of rater 1 on Assignment 2; R2 at A2 = mean of rater 2 at Assignment 2. A1 = overall mean for Assignment 1; A2 = overall mean for Assignment 2.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading