Tests whether a variable's distribution departs from normality — a gatekeeper check for t-tests, ANOVA, and regression residuals.
Check AssumptionsUnivariatealso known as: Normality test
✓ When to use
Checking normality of outcomes within groups (t-tests/ANOVA) or of residuals (regression), especially in small-to-moderate samples.
Alongside visual checks — histogram and Q–Q plot — never instead of them.
Small samples (n < 50) where visual judgment alone is hard; SW has the best power among normality tests there.
✗ When NOT to use
Large samples (n > ~200) — trivial departures become 'significant' while the parametric tests are already robust; rely on Q–Q plots and skew/kurtosis.
As an automatic switch to non-parametric tests — the decision needs sample size and severity context.
On the raw DV in regression — the assumption concerns residuals.
Discrete scales with few values — normality is false by construction; the question is whether it matters.
Data requirements
Dependent / outcome variable
One continuous variable (or residual vector), 3 ≤ n ≤ 5000 in most implementations.
Independent / grouping variable
—
Design
Applied per group/residual set as required by the main analysis.
Sample size guidance
Most informative at n ≈ 10–100.
Assumptions
Independent observations.
Continuous measurement (heavy ties distort the test).
Applied to the quantity the main analysis actually assumes normal (residuals, difference scores, within-group values).
Hypotheses
H₀ — The data come from a normal distribution.
H₁ — The data do not come from a normal distribution. (Note: non-rejection ≠ proof of normality, especially at small n.)
The concept
The W statistic compares the observed order statistics with those expected under normality — effectively asking how straight the Q–Q plot is. W near 1 = compatible with normality; smaller W = departure. It is generally the most powerful omnibus normality test at small n (hence its default status over Kolmogorov–Smirnov).
The paradox to manage: at small n the test is weak (real skew slips through), at large n it is hypersensitive (harmless wobbles flagged) — exactly opposite to where parametric tests need protection. Mature practice: treat SW as one input, weight the Q–Q plot and |skewness| < 2 / |kurtosis| < 7 rules of thumb, and remember the CLT protects mean-based tests as n grows.
Worked example
Before an independent t-test (n = 24 and 26 per group), satisfaction scores are checked per group: W = .96, p = .42 and W = .95, p = .28; Q–Q plots near-linear.
No evidence against normality; the t-test proceeds. (Had n been 800, mild p < .05 'violations' would have been noted and ignored with justification.)
How to run it
by(df$satisfaction, df$group, shapiro.test) # per group
# for regression: test the residuals
model <- lm(y ~ x1 + x2, data = df)
shapiro.test(resid(model))
qqnorm(resid(model)); qqline(resid(model))
from scipy import stats
import statsmodels.api as sm
for g, sub in df.groupby("group"):
print(g, stats.shapiro(sub["satisfaction"]))
# residuals + QQ plot
sm.qqplot(model.resid, line="45", fit=True)
Analyze → Descriptive Statistics → Explore.
DV into Dependent List; grouping variable into Factor List.
Plots: tick 'Normality plots with tests' → output includes Shapiro-Wilk per group plus Q-Q plots.
Report W, df, p per group together with your visual judgment.
No built-in Shapiro–Wilk. Practical checks: histogram (Insert → Statistic Chart), =SKEW(range) and =KURT(range) (rules of thumb |skew| < 2, |kurt| < 7).
Q–Q plot: sort data, compute expected normal quantiles =NORM.S.INV((rank−0.5)/n), scatter observed vs expected.
For the formal test, use R/Python/SPSS/jamovi.
Interpreting the output
p ≥ .05: no detected departure (not proof of normality).
p < .05: departure detected — judge severity via Q–Q/skew/kurtosis and sample size before switching methods.
Report W, df (n), and p per tested set, with the visual evidence.
Base the parametric/non-parametric decision on the whole picture, not the p alone.
APA-style reporting
Shapiro–Wilk tests indicated no significant departure from normality in either group (remote: W = .96, p = .42; office: W = .95, p = .28), and Q–Q plots were approximately linear; parametric analysis was retained.
Common mistakes
Testing the pooled sample instead of each group/residuals.
Abandoning t-tests over p = .03 at n = 500.
Claiming 'data are normal' from a non-significant test at n = 12.
Ignoring plots entirely.
Applying SW to 5-point single items and acting surprised.