Weighted kappa for ordinal categories where near-misses deserve partial credit.
✗ When NOT to use
More than two raters — use Fleiss' kappa or Krippendorff's alpha.
Continuous ratings — ICC.
Raters classify different subsets of objects.
Extremely skewed category distributions — kappa becomes paradoxical (high raw agreement, low κ); report percent agreement and prevalence-adjusted variants alongside.
Data requirements
Dependent / outcome variable
Two parallel categorical ratings of N objects (a k×k agreement table).
Independent / grouping variable
—
Design
Both raters classify every object independently.
Sample size guidance
N ≥ 30 objects for a usable estimate; more for many categories.
Assumptions
Raters judge independently.
Categories mutually exclusive and exhaustive.
The rated objects are independent.
Raters use the same category definitions (a coding manual).
Hypotheses
H₀ — Agreement equals chance level (κ = 0).
H₁ — Agreement exceeds chance (κ > 0). (Usually reported as an estimate with CI rather than a test.)
The concept
Raw percent agreement flatters raters because some agreement happens by chance. Kappa subtracts the chance expectation: κ = (Pₒ − Pₑ)/(1 − Pₑ), where Pₒ is observed agreement and Pₑ the agreement expected from the raters' marginal category frequencies. κ = 1 is perfect agreement; 0 is chance-level; negative values mean worse than chance.
Landis–Koch benchmarks: .21–.40 fair, .41–.60 moderate, .61–.80 substantial, .81–1.00 almost perfect — though many content-analysis standards require ≥ .70–.80 for publishable coding. For ordinal codes, weighted kappa (linear or quadratic weights) credits near-agreements; quadratic-weighted kappa equals the ICC under common conditions.
Worked example
Two coders independently assign 120 open-ended survey comments to five theme categories; they agree on 94 (Pₒ = .78), with chance agreement Pₑ = .24.
Result: κ = .71, 95% CI [.62, .80] — substantial agreement; disagreements are reconciled by discussion before analysis.
Report κ, its ASE-based CI, N, and the observed agreement percentage; for ordinal codes use weighted kappa (Analyze → Scale → Weighted Kappa in SPSS 26+).
κ with CI, plus raw percent agreement for context.
Benchmarks: ≥ .61 substantial; many journals want ≥ .70 for coding studies.
Inspect the agreement table: which categories get confused informs codebook revision.
With skewed marginals, note the kappa paradox and report PABAK or Gwet's AC1 as sensitivity checks.
APA-style reporting
Inter-coder reliability for theme assignment was substantial, Cohen's κ = .71, 95% CI [.62, .80], with 78% raw agreement across 120 comments; disagreements were resolved through discussion.
Common mistakes
Reporting only percent agreement.
Using Cohen's kappa for 3+ raters.
Ignoring the paradox under skewed category prevalence.
Letting raters discuss during 'independent' coding.