Compares two independent group means without assuming equal variances — the safer default when group sizes or spreads differ.
Compare GroupsBivariatealso known as: Unequal-variances t-test, Welch's unequal variances t-test
✓ When to use
Two independent groups, continuous outcome — same situation as the independent t-test.
Levene's test is significant, the group SDs differ noticeably, or group sizes are unequal.
Many methodologists recommend Welch as the default two-group test because it costs almost no power when variances are equal and protects you when they are not.
✗ When NOT to use
Paired/repeated measurements — use the Paired t-Test.
Three or more groups — use Welch ANOVA or one-way ANOVA.
Severely non-normal outcome with small samples — use Mann–Whitney U.
Ordinal single-item outcomes.
Data requirements
Dependent / outcome variable
One continuous variable (interval/ratio).
Independent / grouping variable
One categorical variable with two independent levels.
Design
Between-subjects; independent observations.
Sample size guidance
Works with unequal ns; aim for ≥ 20–30 per group. Power analysis as for the standard t-test.
Assumptions
Independence of observations.
Approximate normality within each group (robust with larger n).
Does NOT assume equal variances — that is the point of the correction.
Hypotheses
H₀ — The two population means are equal (μ₁ = μ₂).
H₁ — The two population means differ (μ₁ ≠ μ₂).
The concept
Welch's procedure computes the same mean difference as Student's t-test but estimates the standard error from each group's own variance instead of a pooled variance, and adjusts the degrees of freedom downward using the Welch–Satterthwaite formula (df is usually a decimal number).
When variances happen to be equal, Welch and Student give nearly identical answers; when they are unequal — especially with unequal group sizes — Student's t can badly distort Type I error while Welch stays accurate. This is why journals increasingly accept, and reviewers often prefer, Welch by default.
Worked example
Comparing turnover intention between a small startup unit (n = 24, M = 3.10, SD = 1.05) and a large established unit (n = 96, M = 2.61, SD = 0.58). SDs differ clearly, so Welch is appropriate.
Result: Welch's t(27.8) = 2.21, p = .036, d = 0.68 — the startup unit reports higher turnover intention.
How to run it
# Welch is R's default for t.test
t.test(turnover ~ unit, data = df) # var.equal = FALSE by default
library(effectsize)
cohens_d(turnover ~ unit, data = df, pooled_sd = FALSE)
from scipy import stats
import pingouin as pg
a = df.loc[df.unit == "Startup", "turnover"]
b = df.loc[df.unit == "Established", "turnover"]
print(stats.ttest_ind(a, b, equal_var=False))
print(pg.ttest(a, b, correction=True)) # Welch + effect size
Analyze → Compare Means → Independent-Samples T Test (same dialog as the standard t-test).
SPSS always prints both rows; read the 'Equal variances NOT assumed' row.
Report the Welch t, its (decimal) df, p, mean difference and CI.
Data → Data Analysis → 't-Test: Two-Sample Assuming Unequal Variances'.
Or in a cell: =T.TEST(range1, range2, 2, 3) — type 3 = two-sample unequal variance, two-tailed.
Interpreting the output
Report the decimal df — it signals to readers that Welch was used.
p and CI read exactly as in the standard t-test.
Effect size: Cohen's d computed with unpooled SDs (or report Glass's Δ when one group is a natural control).
APA-style reporting
Because Levene's test indicated unequal variances, Welch's t-test was used; startup employees (M = 3.10, SD = 1.05) reported significantly higher turnover intention than established-unit employees (M = 2.61, SD = 0.58), t(27.8) = 2.21, p = .036, d = 0.68.
Common mistakes
Running Levene's test and still reading the equal-variances row when it is significant.
Rounding Welch's decimal df to an integer in the report.
Thinking Welch is 'less valid' — it is at least as valid and usually safer.
Forgetting that Welch does not fix non-normality — only unequal variances.