Understanding Cohen's Kappa

Understanding Cohen's Kappa

The Paradox of High Agreement

Your LLM judge and a human reviewer label the same 100 responses as Pass or Fail. They agree on 90. A 90% match rate sounds like the validation can end early.

The catch: 85 of those responses were obviously fine, and both raters passed them. They're agreeing on the easy cases. On the 15 borderline responses, they only agreed on 5.

Raw agreement conflates skill with task difficulty. If 95% of your data falls into one category, two raters who match that base rate while guessing will "agree" about 90% of the time by accident.

Enter Cohen's Kappa

Jacob Cohen's insight (1960) was to ask: how much better is our agreement than random chance?

κ=po−pe1−pe\kappa = \frac{p_o - p_e}{1 - p_e}

Symbol guide:

  • κ (kappa): the agreement coefficient we're computing (range: -1 to 1)
  • p_o: observed agreement, the actual proportion of items where both raters agreed
  • p_e: expected agreement, the agreement we'd expect by pure chance
  • 1 - p_e: the maximum possible improvement over chance

Think of it as measuring the "improvement over guessing":

  • κ=1\kappa = 1: Perfect agreement. The raters always agree.
  • κ=0\kappa = 0: Agreement equals chance. No better than random.
  • κ<0\kappa < 0: Worse than chance, which means systematic disagreement.

The 2×2 Contingency Table

Kappa is computed from a simple table counting agreements and disagreements:

Rater B: YesRater B: No
Rater A: Yesaa (both yes)bb (A yes, B no)
Rater A: Nocc (A no, B yes)dd (both no)

Observed agreement: po=a+dnp_o = \frac{a + d}{n}

Expected agreement: If raters were independent:

pe=P(both yes)+P(both no)=(a+b)(a+c)n2+(c+d)(b+d)n2p_e = P(\text{both yes}) + P(\text{both no}) = \frac{(a+b)(a+c)}{n^2} + \frac{(c+d)(b+d)}{n^2}

This accounts for the marginal distributions. If Rater A says "Yes" 80% of the time and Rater B says "Yes" 70% of the time, we'd expect them to both say "Yes" about 56% of the time by chance.

A Worked Example

You're validating an LLM judge against a human reviewer. Both label the same 50 responses as Pass or Fail:

Judge: PassJudge: Fail
Human: Pass184
Human: Fail622

Step 1: observed agreement. They agree on 18+22=4018 + 22 = 40 responses: po=4050=80%p_o = \frac{40}{50} = 80\%.

Step 2: the marginals. The human says Pass 18+4=2218+4 = 22 times; the judge, 18+6=2418+6 = 24 times.

Step 3: expected agreement.

pe=22⋅24+28⋅26502=528+7282500=12562500=50.2%p_e = \frac{22 \cdot 24 + 28 \cdot 26}{50^2} = \frac{528 + 728}{2500} = \frac{1256}{2500} = 50.2\%

Two raters with these base rates would agree half the time with zero skill.

Step 4: kappa.

κ=0.800−0.5021−0.502=0.2980.498≈0.60\kappa = \frac{0.800 - 0.502}{1 - 0.502} = \frac{0.298}{0.498} \approx 0.60

The 80% headline becomes "moderate, brushing substantial." The judge captured about 60% of the headroom that chance left available: useful, but not the near-perfect proxy the raw number implied.

Try It Yourself

Two simulated raters label the same items yes or no. Each has a lean (how readily they say yes) and their own independent noise. Every item appears as a dot in the quadrant for its pair of answers. On the diagonal, pale dots are agreement that chance would have produced anyway at these yes-rates, and solid dots are agreement beyond it; κ is the solid share of the room chance left over.

Interactive

Two raters, one yes-or-no question

Half the items are yeses. Chance agreement sits near 50%, so most of what the raters agree on is earned.

Base rate50%
Items200

Rater A

Lean0.00
Noise0.3

Rater B

Lean0.00
Noise0.3

Every item, placed by what the two raters said

B says yesB says noA says yes
both yes91 · chance 55
A yes, B no12 · chance 48
A says no
A no, B yes15 · chance 51
both no82 · chance 46
  • agreement chance would give anyway
  • agreement beyond chance
  • disagreement
  • what chance predicted but didn't happen

Agreement, with chance's share marked off

0%chance 50%100%

The raters agree on 87% of items. Chance alone would get them to 50%, leaving 50 points of headroom. They earned 36 of those points: κ = 0.73.

Observed po
87%
(a + d) / n
Chance pe
50%
from the margins
Cohen's κ
0.73
substantial
κ = 0.73
0
0.40
0.60
0.80
1

Start with Balanced classes, then switch to Prevalence paradox: raw agreement hardly moves, but the diagonal turns pale. Same lean and Opposite leans give the raters identical noise and differ only in which way they lean, and κ drops from about 0.65 to about 0.4. That is the bias problem described below.

Your Turn

Fresh counts each time, same four moves as the worked example: agreements, marginals, the chance floor, then κ. The hints, if you need them, come in three sizes.

Practice problem

Moderation Queue

A senior moderator and a trainee moderator each labeled the same 50 posts as Approve or Remove. The counts are below. Work down to κ.
Trainee moderator: ApproveTrainee moderator: Remove
Senior moderator: Approve217
Senior moderator: Remove517

n = 50 posts

Step 1 of 40/4

Observed Agreement (pₒ)

Of the 50 posts, how many did the two raters call the same way--and what share is that?

%

Why Kappa Can Be Misleading

The Prevalence Problem

When one category dominates, Kappa can be counterintuitively low despite high agreement. With 95% base rate:

  • Raw agreement might be 94%
  • But chance agreement is ~90%
  • So κ=0.94−0.901−0.90=0.40\kappa = \frac{0.94 - 0.90}{1 - 0.90} = 0.40, which Landis and Koch would file under merely "fair."

Kappa is doing its job here: most of the "agreement" was coming from the easy majority class, and the low score says so.

The Bias Problem

If both raters have the same bias (e.g., both tend to say "Yes"), Kappa will be higher than if they have opposite biases, even when their accuracy is identical.

Interpretation Guidelines

Landis & Koch (1977) proposed these benchmarks:

κ\kappaInterpretation
< 0.00Poor (worse than chance)
0.00–0.20Slight
0.21–0.40Fair
0.41–0.60Moderate
0.61–0.80Substantial
0.81–1.00Almost perfect

Treat the benchmarks as folklore rather than law. A Kappa of 0.60 might be fine for a subjective task like rating "creativity" and disqualifying for a safety-critical classifier.

Kappa vs Krippendorff's Alpha

FeatureCohen's KappaKrippendorff's Alpha
Number of raters2 onlyAny number
Missing dataNot allowedHandles gracefully
Data typesNominal (extensions exist)Nominal, ordinal, interval, ratio
Use caseSimple pairwise agreementComplex annotation projects

Rule of thumb: Use Kappa for quick two-rater sanity checks. Use Alpha for serious annotation quality measurement.

Code

from sklearn.metrics import cohen_kappa_score
 
rater_a = [1, 1, 0, 1, 0, 0, 1, 1, 0, 1]
rater_b = [1, 0, 0, 1, 0, 1, 1, 1, 0, 1]
 
kappa = cohen_kappa_score(rater_a, rater_b)
print(f"Cohen's Kappa: {kappa:.3f}")
library(irr)
kappa2(data.frame(rater_a, rater_b))

Further Reading

  • Cohen, J. (1960). "A coefficient of agreement for nominal scales." Educational and Psychological Measurement
  • Landis & Koch (1977). "The measurement of observer agreement for categorical data"
  • Wikipedia: Cohen's kappa -- good worked examples