Statistical Methods Atlas
Atlas › Compare Groups › Welch's t-Test

Welch's t-Test

Compares two independent group means without assuming equal variances — the safer default when group sizes or spreads differ.

Compare GroupsBivariate also known as: Unequal-variances t-test, Welch's unequal variances t-test

✓ When to use

  • Two independent groups, continuous outcome — same situation as the independent t-test.
  • Levene's test is significant, the group SDs differ noticeably, or group sizes are unequal.
  • Many methodologists recommend Welch as the default two-group test because it costs almost no power when variances are equal and protects you when they are not.

✗ When NOT to use

  • Paired/repeated measurements — use the Paired t-Test.
  • Three or more groups — use Welch ANOVA or one-way ANOVA.
  • Severely non-normal outcome with small samples — use Mann–Whitney U.
  • Ordinal single-item outcomes.

Data requirements

Dependent / outcome variableOne continuous variable (interval/ratio).
Independent / grouping variableOne categorical variable with two independent levels.
DesignBetween-subjects; independent observations.
Sample size guidanceWorks with unequal ns; aim for ≥ 20–30 per group. Power analysis as for the standard t-test.

Assumptions

Hypotheses

H₀ — The two population means are equal (μ₁ = μ₂).
H₁ — The two population means differ (μ₁ ≠ μ₂).

The concept

Welch's procedure computes the same mean difference as Student's t-test but estimates the standard error from each group's own variance instead of a pooled variance, and adjusts the degrees of freedom downward using the Welch–Satterthwaite formula (df is usually a decimal number).

When variances happen to be equal, Welch and Student give nearly identical answers; when they are unequal — especially with unequal group sizes — Student's t can badly distort Type I error while Welch stays accurate. This is why journals increasingly accept, and reviewers often prefer, Welch by default.

Worked example

Comparing turnover intention between a small startup unit (n = 24, M = 3.10, SD = 1.05) and a large established unit (n = 96, M = 2.61, SD = 0.58). SDs differ clearly, so Welch is appropriate.

Result: Welch's t(27.8) = 2.21, p = .036, d = 0.68 — the startup unit reports higher turnover intention.

How to run it

# Welch is R's default for t.test
t.test(turnover ~ unit, data = df)          # var.equal = FALSE by default

library(effectsize)
cohens_d(turnover ~ unit, data = df, pooled_sd = FALSE)

Interpreting the output

APA-style reporting

Because Levene's test indicated unequal variances, Welch's t-test was used; startup employees (M = 3.10, SD = 1.05) reported significantly higher turnover intention than established-unit employees (M = 2.61, SD = 0.58), t(27.8) = 2.21, p = .036, d = 0.68.

Common mistakes

Related methods

Independent-Samples t-TestEqual-variance versionMann–Whitney U TestNon-parametric alternativeWelch ANOVA3+ groups, unequal variancesLevene's TestDiagnose unequal variances
← One-Sample t-TestPaired-Samples t-Test →