Statistical Methods Atlas
Atlas › Measure Reliability › Intraclass Correlation Coefficient (ICC)

Intraclass Correlation Coefficient (ICC)

Reliability of continuous ratings — consistency or agreement among raters, or of repeated measurements; also the aggregation statistic in multilevel research.

Measure ReliabilityMultivariate also known as: ICC(1), ICC(2,k), Inter-rater reliability (continuous)

✓ When to use

  • Several raters score the same targets on continuous scales (performance ratings, essay marks).
  • Test–retest reliability of continuous measures.
  • Multilevel research: ICC(1) as the share of variance between groups; ICC(2) as reliability of group means (justifying aggregation of individual responses to team/firm level).

✗ When NOT to use

  • Categorical ratings — kappa family.
  • Only two variables' linear association — Pearson r ignores systematic rater differences (a rater scoring everyone 1 point higher preserves r = 1 but hurts agreement).
  • Choosing an ICC form blindly — consistency vs absolute agreement, single vs average measures answer different questions.

Data requirements

Dependent / outcome variableAn objects × raters (or occasions) matrix of continuous scores.
Independent / grouping variable—
DesignFully crossed preferred (all raters rate all targets); one-way designs when raters differ per target.
Sample size guidance≥ 30 targets recommended; more raters sharpen average-measure ICCs.

Assumptions

Hypotheses

H₀ — ICC = 0 (no between-target consistency beyond chance).
H₁ — ICC > 0. (Reported as estimate with CI.)

The concept

The ICC is a variance ratio: between-target variance divided by total variance (between-target + rater + error, depending on the form). Intuitively, reliability is high when true differences between the rated targets dominate the noise raters add.

The form matters: ICC(2,1) (two-way random, single rater, absolute agreement) asks how trustworthy ONE rater's score is; ICC(2,k) asks how trustworthy the AVERAGE of k raters is — always higher. 'Consistency' forgives constant rater leniency; 'absolute agreement' punishes it. In multilevel HR research, ICC(1) values as small as .05–.20 are common and meaningful (5–20% of variance lies between teams), while ICC(2) ≥ .70 supports using team means. Benchmarks for rater reliability: < .50 poor, .50–.75 moderate, .75–.90 good, > .90 excellent (Koo & Li, 2016).

Worked example

Three assessors rate 40 candidates' interview performance (0–100) in an assessment center; all rate all candidates.

Result: ICC(2,1) absolute agreement = .62 (single assessor: moderate) but ICC(2,3) = .83 — the panel average is reliable enough for selection decisions.

How to run it

library(psych)
ICC(df[, c("rater1", "rater2", "rater3")])
# read the row matching your design, e.g. ICC2 (single) or ICC2k (average)

# multilevel ICC(1) from a null model:
library(lme4); library(performance)
m0 <- lmer(engagement ~ 1 + (1 | team), data = dfl)
icc(m0)

Interpreting the output

APA-style reporting

Inter-rater reliability for interview scores was moderate for a single assessor, ICC(2,1) = .62, 95% CI [.48, .74], but good for the three-assessor average, ICC(2,3) = .83, 95% CI [.73, .90] (two-way random effects, absolute agreement).

Common mistakes

Related methods

Cohen's KappaCategorical ratingsCronbach's AlphaItems instead of ratersRepeated-Measures ANOVAUnderlying variance machinery
← Cohen's KappaK-Means Clustering →