Statistical Methods Atlas
Atlas › Measure Reliability › Cohen's Kappa

Cohen's Kappa

Chance-corrected agreement between two raters assigning categories — e.g., two coders classifying interview excerpts into themes.

Measure ReliabilityBivariate also known as: κ, Inter-rater agreement (categorical)

✓ When to use

  • Two raters, nominal categories, same set of objects rated by both.
  • Content-analysis coding reliability; diagnostic agreement; audit consistency.
  • Weighted kappa for ordinal categories where near-misses deserve partial credit.

✗ When NOT to use

  • More than two raters — use Fleiss' kappa or Krippendorff's alpha.
  • Continuous ratings — ICC.
  • Raters classify different subsets of objects.
  • Extremely skewed category distributions — kappa becomes paradoxical (high raw agreement, low κ); report percent agreement and prevalence-adjusted variants alongside.

Data requirements

Dependent / outcome variableTwo parallel categorical ratings of N objects (a k×k agreement table).
Independent / grouping variable—
DesignBoth raters classify every object independently.
Sample size guidanceN ≥ 30 objects for a usable estimate; more for many categories.

Assumptions

Hypotheses

H₀ — Agreement equals chance level (κ = 0).
H₁ — Agreement exceeds chance (κ > 0). (Usually reported as an estimate with CI rather than a test.)

The concept

Raw percent agreement flatters raters because some agreement happens by chance. Kappa subtracts the chance expectation: κ = (Pₒ − Pₑ)/(1 − Pₑ), where Pₒ is observed agreement and Pₑ the agreement expected from the raters' marginal category frequencies. κ = 1 is perfect agreement; 0 is chance-level; negative values mean worse than chance.

Landis–Koch benchmarks: .21–.40 fair, .41–.60 moderate, .61–.80 substantial, .81–1.00 almost perfect — though many content-analysis standards require ≥ .70–.80 for publishable coding. For ordinal codes, weighted kappa (linear or quadratic weights) credits near-agreements; quadratic-weighted kappa equals the ICC under common conditions.

Worked example

Two coders independently assign 120 open-ended survey comments to five theme categories; they agree on 94 (Pₒ = .78), with chance agreement Pₑ = .24.

Result: κ = .71, 95% CI [.62, .80] — substantial agreement; disagreements are reconciled by discussion before analysis.

How to run it

library(irr)
kappa2(df[, c("coder1", "coder2")])            # unweighted
kappa2(df[, c("coder1", "coder2")], weight = "squared")  # weighted (ordinal)

# CI via psych:
psych::cohen.kappa(table(df$coder1, df$coder2))

Interpreting the output

APA-style reporting

Inter-coder reliability for theme assignment was substantial, Cohen's κ = .71, 95% CI [.62, .80], with 78% raw agreement across 120 comments; disagreements were resolved through discussion.

Common mistakes

Related methods

Intraclass Correlation Coefficient (ICC)Continuous ratingsChi-Square Test of IndependenceAssociation ≠ agreementCronbach's AlphaItem consistency instead
← Composite Reliability (CR)Intraclass Correlation Coefficient (ICC) →