The classic index of internal consistency — how coherently a set of items measures one construct, e.g., a 6-item commitment scale.
Measure ReliabilityMultivariatealso known as: Coefficient alpha, α
✓ When to use
Reporting internal consistency of multi-item reflective scales (the near-universal journal expectation).
Comparing your scale's consistency with values reported in prior studies.
Item analysis during scale refinement (alpha-if-item-deleted, item-total correlations).
✗ When NOT to use
Multidimensional scales scored as one total — alpha assumes unidimensionality; compute per subscale.
Items with clearly unequal loadings (congeneric items) — McDonald's omega is more accurate.
Two-item scales — report the Spearman–Brown coefficient instead.
As evidence of validity — a reliable scale can consistently measure the wrong thing.
Formative indices — internal consistency logic does not apply.
Data requirements
Dependent / outcome variable
k ≥ 3 items intended to measure one construct, on the same response scale, scored in the same direction (reverse-code first).
Independent / grouping variable
—
Design
One administration to one sample.
Sample size guidance
n ≥ 100 for a reasonably stable estimate; report a confidence interval either way.
Assumptions
Unidimensionality (verify with EFA/CFA — alpha itself does not test it).
Tau-equivalence: all items load equally on the construct; violations make alpha a LOWER bound of reliability.
Uncorrelated item errors (violated by shared wording, inflating alpha).
Items on comparable metrics.
Hypotheses
H₀ — — (an estimate with CI, not a significance test).
H₁ — —
The concept
Alpha rises with the average inter-item correlation and the number of items: α = k·r̄ /(1 + (k − 1)·r̄). Conceptually it estimates the proportion of scale-score variance attributable to the common construct rather than noise — under the assumption that all items are equally good indicators.
Benchmarks: ≥ .70 acceptable for research, ≥ .80 good, ≥ .90 may signal redundancy in long scales. Because alpha mechanically increases with item count, a long scale can post a high alpha out of sheer length while items barely correlate — always look at the mean inter-item correlation (ideal ≈ .15–.50) alongside. When item loadings differ (they usually do), alpha underestimates reliability, which is why omega is now the recommended default; report both to satisfy tradition and best practice.
Worked example
A 6-item affective-commitment scale administered to 240 employees yields α = .86, mean inter-item r = .51; alpha-if-item-deleted shows no item whose removal would raise alpha.
The scale shows good internal consistency; the composite mean score is used in subsequent regressions.
How to run it
library(psych)
items <- df[, paste0("ac", 1:6)] # reverse-code first if needed
alpha(items) # alpha, 95% CI, item-total stats, alpha-if-deleted
# with CI via bootstrap:
alpha(items, n.iter = 1000)
import pingouin as pg
res = pg.cronbach_alpha(data=df[["ac1","ac2","ac3","ac4","ac5","ac6"]])
print(res) # (alpha, 95% CI)
Reverse-code negative items first (Transform → Recode into Different Variables).
Analyze → Scale → Reliability Analysis; move the items in; Model = Alpha.
Statistics: tick 'Scale if item deleted', 'Item', 'Inter-Item Correlations'.
Report α, number of items, n, and note any item whose deletion would raise α (with the content-based justification if you drop it).
Compute each item's variance =VAR.S(item_range) and the total score's variance =VAR.S(total_range).
α = (k/(k−1)) * (1 − Σ(item variances)/variance of total).
Check inter-item correlations with =CORREL per pair (or Data Analysis → Correlation).
Interpreting the output
α with its CI, k items, and n; ≥ .70 acceptable in most research contexts.
Mean inter-item correlation ≈ .15–.50 — coherence without redundancy.
Alpha-if-item-deleted: flag items only when deletion is also justified by content.
Very high α (> .95): consider whether items are paraphrases.
Alpha is about the scale scores in THIS sample — it is not a fixed property of the instrument.
APA-style reporting
The six-item affective commitment scale demonstrated good internal consistency, Cronbach's α = .86, 95% CI [.83, .89] (mean inter-item r = .51).
Common mistakes
Computing one alpha across a multidimensional instrument.
Forgetting to reverse-code before analysis (alpha collapses).
Dropping items purely to maximize alpha.
Citing 'the scale is valid because alpha was high'.
Reporting alpha from a previous study instead of your own sample.