Statistical Methods Atlas
Atlas › Compare Groups › Paired-Samples t-Test

Paired-Samples t-Test

Tests whether the mean of the same group differs across two occasions or conditions — e.g., job satisfaction before vs after a training program.

Compare GroupsBivariate also known as: Dependent t-test, Matched-pairs t-test, Repeated-measures t-test

✓ When to use

  • The same participants are measured twice (pre/post, condition A/condition B), or pairs are naturally matched (twins, matched employees).
  • The outcome is continuous and the difference scores are approximately normal.
  • Classic uses: training evaluations, intervention studies, within-person comparisons of two job conditions.

✗ When NOT to use

  • The two measurements come from different, unrelated people — use the Independent-Samples t-Test.
  • There are three or more measurement occasions — use Repeated-Measures ANOVA.
  • Difference scores are clearly non-normal in a small sample, or the outcome is ordinal — use the Wilcoxon Signed-Rank test.
  • Substantial dropout between waves creates non-random missingness — address missing data first.

Data requirements

Dependent / outcome variableOne continuous variable measured twice (two columns: time 1, time 2).
Independent / grouping variableThe within-person factor with two levels (time or condition).
DesignWithin-subjects or matched pairs; each row is one participant (or one matched pair).
Sample size guidancePower depends on the correlation between occasions; paired designs are usually more powerful than independent ones. G*Power: dz = 0.50, power .80 → n ≈ 34 pairs.

Assumptions

Hypotheses

H₀ — The mean difference in the population is zero (μd = 0).
H₁ — The mean difference is not zero (μd ≠ 0).

The concept

The paired t-test is simply a one-sample t-test on the difference scores: compute d = X₂ − X₁ for each person, then test whether the mean difference departs from zero. Because each person serves as their own control, stable individual differences cancel out, which typically shrinks the error term and boosts power relative to a between-subjects comparison.

Effect size is usually Cohen's dz (mean difference divided by the SD of the differences); some report dav or drm based on raw-score SDs — state which you used.

Worked example

A company measures job satisfaction (1–5) in 48 employees before and six months after introducing flexible work. Pre: M = 3.42 (SD = 0.68); Post: M = 3.71 (SD = 0.63); mean gain = 0.29 (SDdiff = 0.55).

Result: t(47) = 3.65, p < .001, dz = 0.53 — satisfaction rose significantly after the policy.

How to run it

# wide format: pre, post
diffs <- df$post - df$pre
shapiro.test(diffs)                      # normality of differences

t.test(df$post, df$pre, paired = TRUE)

library(effectsize)
cohens_d(df$post, df$pre, paired = TRUE)  # dz

Interpreting the output

APA-style reporting

A paired-samples t-test showed that job satisfaction increased significantly from before (M = 3.42, SD = 0.68) to after (M = 3.71, SD = 0.63) the flexible-work policy, t(47) = 3.65, p < .001, 95% CI [0.13, 0.45], dz = 0.53.

Common mistakes

Related methods

Independent-Samples t-TestDifferent people in each groupWilcoxon Signed-Rank TestNon-parametric alternativeRepeated-Measures ANOVAThree or more occasions
← Welch's t-TestMann–Whitney U Test →