Reliability of continuous ratings — consistency or agreement among raters, or of repeated measurements; also the aggregation statistic in multilevel research.
Measure ReliabilityMultivariatealso known as: ICC(1), ICC(2,k), Inter-rater reliability (continuous)
✓ When to use
Several raters score the same targets on continuous scales (performance ratings, essay marks).
Test–retest reliability of continuous measures.
Multilevel research: ICC(1) as the share of variance between groups; ICC(2) as reliability of group means (justifying aggregation of individual responses to team/firm level).
✗ When NOT to use
Categorical ratings — kappa family.
Only two variables' linear association — Pearson r ignores systematic rater differences (a rater scoring everyone 1 point higher preserves r = 1 but hurts agreement).
Choosing an ICC form blindly — consistency vs absolute agreement, single vs average measures answer different questions.
Data requirements
Dependent / outcome variable
An objects × raters (or occasions) matrix of continuous scores.
Independent / grouping variable
—
Design
Fully crossed preferred (all raters rate all targets); one-way designs when raters differ per target.
Sample size guidance
≥ 30 targets recommended; more raters sharpen average-measure ICCs.
Assumptions
Continuous, roughly normal ratings.
Independence of targets.
The chosen ICC form matches the design (one-way vs two-way; random vs fixed raters; single vs average; consistency vs agreement — the Shrout & Fleiss / McGraw & Wong taxonomy).
Hypotheses
H₀ — ICC = 0 (no between-target consistency beyond chance).
H₁ — ICC > 0. (Reported as estimate with CI.)
The concept
The ICC is a variance ratio: between-target variance divided by total variance (between-target + rater + error, depending on the form). Intuitively, reliability is high when true differences between the rated targets dominate the noise raters add.
The form matters: ICC(2,1) (two-way random, single rater, absolute agreement) asks how trustworthy ONE rater's score is; ICC(2,k) asks how trustworthy the AVERAGE of k raters is — always higher. 'Consistency' forgives constant rater leniency; 'absolute agreement' punishes it. In multilevel HR research, ICC(1) values as small as .05–.20 are common and meaningful (5–20% of variance lies between teams), while ICC(2) ≥ .70 supports using team means. Benchmarks for rater reliability: < .50 poor, .50–.75 moderate, .75–.90 good, > .90 excellent (Koo & Li, 2016).
Worked example
Three assessors rate 40 candidates' interview performance (0–100) in an assessment center; all rate all candidates.
Result: ICC(2,1) absolute agreement = .62 (single assessor: moderate) but ICC(2,3) = .83 — the panel average is reliable enough for selection decisions.
How to run it
library(psych)
ICC(df[, c("rater1", "rater2", "rater3")])
# read the row matching your design, e.g. ICC2 (single) or ICC2k (average)
# multilevel ICC(1) from a null model:
library(lme4); library(performance)
m0 <- lmer(engagement ~ 1 + (1 | team), data = dfl)
icc(m0)
import pingouin as pg
# long format: targets, raters, scores
res = pg.intraclass_corr(data=dfl, targets="candidate", raters="rater",
ratings="score")
print(res) # ICC1..ICC3k with CIs — pick the row matching your design
Data layout: one row per target, one column per rater.
Feasible but fiddly — prefer R/Python/SPSS for CIs.
Interpreting the output
Name the exact form: e.g., 'ICC(2,1), two-way random effects, absolute agreement'.
Single-measure ICC for one-rater decisions; average-measure for panel scores.
Koo–Li benchmarks for rater work; multilevel ICC(1) interpreted as variance share, not against these benchmarks.
Report the CI — ICCs from 30–40 targets are imprecise.
APA-style reporting
Inter-rater reliability for interview scores was moderate for a single assessor, ICC(2,1) = .62, 95% CI [.48, .74], but good for the three-assessor average, ICC(2,3) = .83, 95% CI [.73, .90] (two-way random effects, absolute agreement).
Common mistakes
Reporting 'the ICC' without specifying the form.
Using Pearson r between raters as 'agreement'.
Average-measure ICC quoted for decisions made by single raters.
Consistency ICC when leniency differences matter.
In multilevel work, dismissing ICC(1) = .10 as 'too low' — it often justifies multilevel modeling.