Tests whether the mean of the same group differs across two occasions or conditions — e.g., job satisfaction before vs after a training program.
Compare GroupsBivariatealso known as: Dependent t-test, Matched-pairs t-test, Repeated-measures t-test
✓ When to use
The same participants are measured twice (pre/post, condition A/condition B), or pairs are naturally matched (twins, matched employees).
The outcome is continuous and the difference scores are approximately normal.
Classic uses: training evaluations, intervention studies, within-person comparisons of two job conditions.
✗ When NOT to use
The two measurements come from different, unrelated people — use the Independent-Samples t-Test.
There are three or more measurement occasions — use Repeated-Measures ANOVA.
Difference scores are clearly non-normal in a small sample, or the outcome is ordinal — use the Wilcoxon Signed-Rank test.
Substantial dropout between waves creates non-random missingness — address missing data first.
Data requirements
Dependent / outcome variable
One continuous variable measured twice (two columns: time 1, time 2).
Independent / grouping variable
The within-person factor with two levels (time or condition).
Design
Within-subjects or matched pairs; each row is one participant (or one matched pair).
Sample size guidance
Power depends on the correlation between occasions; paired designs are usually more powerful than independent ones. G*Power: dz = 0.50, power .80 → n ≈ 34 pairs.
Assumptions
Pairs are independent of other pairs.
The DIFFERENCE scores (time 2 − time 1) are approximately normally distributed — not necessarily the raw scores.
The outcome is continuous; both occasions use the same instrument/scale.
Hypotheses
H₀ — The mean difference in the population is zero (μd = 0).
H₁ — The mean difference is not zero (μd ≠ 0).
The concept
The paired t-test is simply a one-sample t-test on the difference scores: compute d = X₂ − X₁ for each person, then test whether the mean difference departs from zero. Because each person serves as their own control, stable individual differences cancel out, which typically shrinks the error term and boosts power relative to a between-subjects comparison.
Effect size is usually Cohen's dz (mean difference divided by the SD of the differences); some report dav or drm based on raw-score SDs — state which you used.
Worked example
A company measures job satisfaction (1–5) in 48 employees before and six months after introducing flexible work. Pre: M = 3.42 (SD = 0.68); Post: M = 3.71 (SD = 0.63); mean gain = 0.29 (SDdiff = 0.55).
Result: t(47) = 3.65, p < .001, dz = 0.53 — satisfaction rose significantly after the policy.
Or Data → Data Analysis → 't-Test: Paired Two Sample for Means'.
Cohen's dz: compute the difference column C = B − A, then =AVERAGE(C:C)/STDEV.S(C:C).
Interpreting the output
t and df (n pairs − 1).
p < α: the two occasions differ; the sign of the mean difference gives direction.
95% CI of the mean difference — practical range of the change.
dz — magnitude of change in difference-score SD units; also report both occasion means and SDs.
APA-style reporting
A paired-samples t-test showed that job satisfaction increased significantly from before (M = 3.42, SD = 0.68) to after (M = 3.71, SD = 0.63) the flexible-work policy, t(47) = 3.65, p < .001, 95% CI [0.13, 0.45], dz = 0.53.
Common mistakes
Treating paired data as independent groups (or vice versa) — this misstates the standard error.
Checking normality of the raw scores instead of the difference scores.
Attributing the change to the intervention without a control group — a pre/post design alone cannot rule out history or maturation effects.
Deleting participants listwise without examining why they dropped out between waves.