Skip to main content icon/video/no-internet

Kappa Coefficient of Agreement

Kappa, one of several coefficients used to estimate inter-rater and similar types of reliability, was developed in 1960 by Jacob Cohen. In its original conception, kappa, denoted κ, was an index used to measure the level of consistency between two raters who use rubrics or other instruments to place subjects (i.e., people) into one of κ nominal categories. For example, two evaluators might use a rubric to classify the instructional strategies used in a classroom as one of two nominal categories, such as effective or ineffective, or two psychologists might use an instrument to identify a person’s depression as major depression, bipolar disorder, persistent depression, or psychotic depression. In both of these cases, subjectivity might lead to disagreements about the category assigned and raise concerns about the use and interpretation of the instrument. For this reason, coefficients of agreement, such as κ, have been developed.

Since its inception, κ has been one of the most widely utilized coefficients of agreement and has been used in such fields as education, medicine, and the social sciences. In these and other fields, κ can be used not only to estimate reliability but also to quantify the variance that can be attributed to the rating process. This entry provides a description of the development of κ and some extensions. It will also include a discussion of considerations that must be attended to when using this coefficient of agreement.

Development of κ

A simple, logical coefficient of agreement between two raters is the observed proportion of subjects the raters placed into the same category. This is called the observed proportion of agreement, denoted pa. Table 1 is a contingency table containing hypothetical data that depicts the proportion of subjects and the categories in which each rater placed them. For example, 4% of subjects were placed in Category C by Rater 1 and in Category A by Rater 2. Based on Table 1, the proportion of agreement, which is the sum of the proportions along the main diagonal given by pa=i=1kpii, is .70. Therefore, the raters agreed and placed 70% of subjects into the same category.

Table 1 Hypothetical Data Depicting the Proportion of Subjects Placed by Raters in Each Category

Rater 2

Total 1

Category A

Category B

Category C

Category D

Rater 1

Category A

.21

.03

.02

.02

.28

Category B

.01

.16

.03

.02

.22

Category C

.04

.02

.15

.03

.24

Category D

.00

.02

.06

.18

.26

Total 2

.26

.23

.26

.25

1.00

Although the proportion of agreement is a simple and logical coefficient of agreement, it has been criticized by Cohen and others as being insufficient. The inadequacy stems from the idea that raters will have a certain level of agreement by chance alone, and pa does not take that into consideration. Others before Cohen developed their own corrections for chance agreement; however, Cohen’s correction is one of the few that involves the use of marginal proportions in its calculation with the assumption that each rater’s marginal proportions are specific to that rater and not common across raters. To be exact, the expected proportion of chance agreement, denoted pc, is given by pc=i=1kpi.p.i, where κ is the number of categories, pi. is the proportion of subjects Rater 1 put into Category i, and p.i is the proportion of subjects Rater 2 put into Category i. Thus, Cohen’s expected proportion of agreement by chance is the sum of the product of marginal proportions for each category. Using the data in Table 1, pc = 0.2508. This is interpreted to mean that by chance alone, it is expected that the raters will agree on approximately 25% of the ratings. Cohen’s κ uses both pa and pc. More specifically, the formula for κ is given by κ=papc1pc. Proper use of κ requires the following assumptions set forth by Cohen: (a) independence of subjects; (b) nominal, independent, mutually exclusive, and exhaustive categories; and (c) independence of raters. Assuming that these assumptions have been satisfied for the data in Table 1, κ = 0.60. This value represents the proportion of agreement between the raters after the removal of chance agreement.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading