Tests whether the observed frequency distribution of one categorical variable matches expected proportions — e.g., do complaint types occur equally often?
Test AssociationUnivariatealso known as: One-sample chi-square
✓ When to use
One categorical variable; comparison against theoretical or benchmark proportions (equal shares, census proportions, last year's mix).
Manipulation checks: are conditions balanced? Response-mix questions: does our applicant pool match the labor market?
✗ When NOT to use
Two categorical variables — use the test of independence.
Expected counts < 5 in over 20% of categories — merge categories or use an exact multinomial test.
Ordered categories with a trend question — ordinal methods are more powerful.
Continuous data — don't bin them just to run χ².
Data requirements
Dependent / outcome variable
One categorical variable with k ≥ 2 categories; observed counts per category.
Independent / grouping variable
A set of expected proportions fixed before seeing the data.
Design
Independent observations, each in exactly one category.
Sample size guidance
n large enough that every expected count ≥ 5 (n × smallest expected proportion ≥ 5).
Assumptions
Independence of observations.
Mutually exclusive, exhaustive categories.
Expected proportions specified a priori (theory, benchmark, or equality).
Adequate expected counts.
Hypotheses
H₀ — The population proportions equal the specified values (π₁ = p₁, …, πk = pk).
H₁ — At least one population proportion differs from its specified value.
The concept
Expected counts are n × specified proportions. χ² = Σ(O − E)²/E with k − 1 degrees of freedom measures the total misfit between the observed distribution and the benchmark pattern. It is an omnibus test — a significant result means the profile deviates somewhere, and standardized residuals (O − E)/√E identify which categories are over- or under-represented.
A service center logs 240 complaints across four types and asks whether they are equally distributed (expected 60 each). Observed: billing 85, delivery 62, quality 55, other 38.
Result: χ²(3, N = 240) = 18.55, p < .001, w = .28 — billing complaints are over-represented (std. residual +3.2), 'other' under-represented (−2.8).
How to run it
obs <- c(billing = 85, delivery = 62, quality = 55, other = 38)
res <- chisq.test(obs, p = rep(1/4, 4))
res
res$stdres # which categories deviate
sqrt(res$statistic / sum(obs)) # Cohen's w
from scipy import stats
import numpy as np
obs = np.array([85, 62, 55, 38])
exp = np.repeat(obs.sum()/4, 4)
chi2, p = stats.chisquare(obs, exp)
print(chi2, p, np.sqrt(chi2/obs.sum())) # + Cohen's w
Standardized residuals per category (|res| > 2 noteworthy) show where the misfit lives.
Cohen's w for magnitude.
Substantive reporting: observed % vs expected % per category.
APA-style reporting
A chi-square goodness-of-fit test indicated that complaint types were not equally distributed, χ²(3, N = 240) = 18.55, p < .001, w = .28; billing complaints occurred more often than expected (35.4% vs 25.0%).
Common mistakes
Choosing expected proportions after seeing the data.
Testing against equality when a meaningful external benchmark exists.
Very uneven category definitions that make 'equal shares' a straw man.
Ignoring which categories deviate — the omnibus p alone is uninformative.