Statistical Methods Atlas
Atlas › Test Association › Chi-Square Goodness-of-Fit Test

Chi-Square Goodness-of-Fit Test

Tests whether the observed frequency distribution of one categorical variable matches expected proportions — e.g., do complaint types occur equally often?

Test AssociationUnivariate also known as: One-sample chi-square

✓ When to use

  • One categorical variable; comparison against theoretical or benchmark proportions (equal shares, census proportions, last year's mix).
  • Manipulation checks: are conditions balanced? Response-mix questions: does our applicant pool match the labor market?

✗ When NOT to use

  • Two categorical variables — use the test of independence.
  • Expected counts < 5 in over 20% of categories — merge categories or use an exact multinomial test.
  • Ordered categories with a trend question — ordinal methods are more powerful.
  • Continuous data — don't bin them just to run χ².

Data requirements

Dependent / outcome variableOne categorical variable with k ≥ 2 categories; observed counts per category.
Independent / grouping variableA set of expected proportions fixed before seeing the data.
DesignIndependent observations, each in exactly one category.
Sample size guidancen large enough that every expected count ≥ 5 (n × smallest expected proportion ≥ 5).

Assumptions

Hypotheses

H₀ — The population proportions equal the specified values (π₁ = p₁, …, πk = pk).
H₁ — At least one population proportion differs from its specified value.

The concept

Expected counts are n × specified proportions. χ² = Σ(O − E)²/E with k − 1 degrees of freedom measures the total misfit between the observed distribution and the benchmark pattern. It is an omnibus test — a significant result means the profile deviates somewhere, and standardized residuals (O − E)/√E identify which categories are over- or under-represented.

Effect size: Cohen's w = √(χ²/n) (.10 small, .30 medium, .50 large).

Worked example

A service center logs 240 complaints across four types and asks whether they are equally distributed (expected 60 each). Observed: billing 85, delivery 62, quality 55, other 38.

Result: χ²(3, N = 240) = 18.55, p < .001, w = .28 — billing complaints are over-represented (std. residual +3.2), 'other' under-represented (−2.8).

How to run it

obs <- c(billing = 85, delivery = 62, quality = 55, other = 38)
res <- chisq.test(obs, p = rep(1/4, 4))
res
res$stdres          # which categories deviate
sqrt(res$statistic / sum(obs))    # Cohen's w

Interpreting the output

APA-style reporting

A chi-square goodness-of-fit test indicated that complaint types were not equally distributed, χ²(3, N = 240) = 18.55, p < .001, w = .28; billing complaints occurred more often than expected (35.4% vs 25.0%).

Common mistakes

Related methods

Chi-Square Test of IndependenceTwo categorical variablesFisher's Exact TestSparse 2×2 tables
← Fisher's Exact TestSimple Linear Regression →