Tests whether the means of two independent groups differ on a continuous outcome — e.g., do male and female employees differ in job satisfaction?
Compare GroupsBivariatealso known as: Two-sample t-test, Student's t-test, Unpaired t-test
✓ When to use
You have exactly two independent groups (different people in each group), e.g., public vs private sector employees.
The outcome is measured on a continuous (interval/ratio) scale, such as a job satisfaction score.
You want to know whether the difference between the two group means is larger than chance would produce.
Group sizes are reasonably large or the outcome is approximately normally distributed in each group.
✗ When NOT to use
The same participants are measured twice (before/after) — use the Paired t-Test instead.
You have three or more groups — use One-Way ANOVA; running multiple t-tests inflates Type I error.
The outcome is ordinal (e.g., a single 5-point item) or clearly non-normal with a small sample — use Mann–Whitney U.
Group variances are clearly unequal (Levene's test significant) — use the Welch t-Test.
The outcome is categorical (yes/no) — use a Chi-Square Test of Independence or logistic regression.
Data requirements
Dependent / outcome variable
One continuous variable (interval or ratio), e.g., a mean scale score for job satisfaction.
Independent / grouping variable
One categorical variable with exactly two independent levels, e.g., sector (public / private).
Design
Between-subjects: each participant appears in only one group, and observations are independent of each other.
Sample size guidance
Roughly 20–30 per group gives reasonable robustness to non-normality; use G*Power for an a-priori calculation (e.g., detecting d = 0.50 at power .80, α = .05 requires ≈ 64 per group).
Assumptions
Independence of observations — participants in one group are unrelated to those in the other, and no participant appears twice.
Normality — the outcome is approximately normally distributed within each group (check Shapiro–Wilk, histograms, Q–Q plots); with n ≥ 30 per group the test is robust to moderate violations.
Homogeneity of variance — the two groups have similar variances (check Levene's test); if violated, report the Welch correction.
The outcome is continuous (interval/ratio); the grouping variable is genuinely dichotomous, not an artificially split continuous variable.
Hypotheses
H₀ — The population means of the two groups are equal (μ₁ = μ₂); any observed difference is due to sampling error.
H₁ — The population means of the two groups are not equal (μ₁ ≠ μ₂). (One-tailed: μ₁ > μ₂ or μ₁ < μ₂, only when direction was predicted in advance.)
The concept
The test computes a t statistic: the difference between the two sample means divided by the standard error of that difference. Intuitively, it asks how many 'standard errors of the difference' apart the two means are. If the groups truly came from the same population, t values near zero are common and large values are rare.
The resulting t is compared against the t distribution with (n₁ + n₂ − 2) degrees of freedom to obtain a p-value — the probability of observing a difference this large (or larger) if H₀ were true. Because statistical significance says nothing about practical importance, always pair the p-value with an effect size: Cohen's d expresses the mean difference in standard-deviation units (≈ 0.20 small, 0.50 medium, 0.80 large).
Worked example
A researcher asks whether work-from-home (n = 62) and in-office (n = 58) employees differ in job satisfaction, measured with a 5-item Likert-type scale (1–5, averaged). Mean satisfaction is 3.92 (SD = 0.61) for remote employees and 3.64 (SD = 0.66) for office employees.
Levene's test is non-significant (variances equal), normality is acceptable, so the standard independent t-test applies. The result: t(118) = 2.41, p = .017, d = 0.44 — remote employees report significantly higher satisfaction, a small-to-medium effect.
How to run it
# outcome: satisfaction (numeric), group: workmode (factor: Remote/Office)
library(effectsize)
# 1. Descriptives by group
aggregate(satisfaction ~ workmode, data = df, FUN = function(x) c(M = mean(x), SD = sd(x)))
# 2. Check assumptions
shapiro.test(df$satisfaction[df$workmode == "Remote"])
shapiro.test(df$satisfaction[df$workmode == "Office"])
car::leveneTest(satisfaction ~ workmode, data = df)
# 3. The test (var.equal = TRUE for classic Student's t)
t.test(satisfaction ~ workmode, data = df, var.equal = TRUE)
# 4. Effect size
cohens_d(satisfaction ~ workmode, data = df)
import pandas as pd
from scipy import stats
import pingouin as pg # pip install pingouin
remote = df.loc[df["workmode"] == "Remote", "satisfaction"]
office = df.loc[df["workmode"] == "Office", "satisfaction"]
# 1. Descriptives
print(df.groupby("workmode")["satisfaction"].agg(["mean", "std", "count"]))
# 2. Assumptions
print(stats.shapiro(remote), stats.shapiro(office))
print(stats.levene(remote, office))
# 3. Test + Cohen's d in one call
print(pg.ttest(remote, office, correction=False)) # correction=True -> Welch
Analyze → Compare Means → Independent-Samples T Test.
Move the outcome (e.g., Satisfaction) into 'Test Variable(s)'.
Move the grouping variable (e.g., WorkMode) into 'Grouping Variable' and click 'Define Groups' (enter the two codes, e.g., 1 and 2).
Click OK. In the output, first read Levene's Test: if p > .05 use the 'Equal variances assumed' row; if p ≤ .05 use 'Equal variances not assumed' (Welch).
Report t, df, p (Sig. 2-tailed), and the mean difference with its 95% CI. SPSS 27+ also prints Cohen's d under 'Effect Sizes'.
Arrange the two groups' scores in two columns (e.g., A = Remote, B = Office).
Data → Data Analysis → 't-Test: Two-Sample Assuming Equal Variances' (or 'Unequal Variances' for Welch).
Set Variable 1/2 ranges, Hypothesized Mean Difference = 0, Alpha = 0.05 → OK.
Read 't Stat', 'P(T<=t) two-tail', and df. Quick alternative in a cell: =T.TEST(A2:A63, B2:B59, 2, 2) returns the two-tailed p-value directly.
Compute Cohen's d manually: (mean1 − mean2) / pooled SD, where pooled SD = SQRT(((n1−1)*s1^2 + (n2−1)*s2^2)/(n1+n2−2)).
Interpreting the output
t and df — the test statistic and degrees of freedom (n₁ + n₂ − 2); report both.
p-value — if p < .05 (or your chosen α), the group means differ significantly; if p ≥ .05, you failed to find a difference (which is not proof of equality).
Mean difference and its 95% CI — a CI that excludes 0 corresponds to a significant result and shows the plausible range of the true difference.
Cohen's d — the size of the difference in SD units; a significant but tiny d may be practically meaningless in large samples.
Direction — look at the group means to state which group scored higher; the sign of t depends only on coding order.
APA-style reporting
An independent-samples t-test showed that remote employees (M = 3.92, SD = 0.61) reported significantly higher job satisfaction than office employees (M = 3.64, SD = 0.66), t(118) = 2.41, p = .017, 95% CI of the difference [0.05, 0.51], d = 0.44.
Common mistakes
Running several t-tests across many groups or outcomes without correction — this inflates Type I error; use ANOVA or adjust α.
Reading the 'Equal variances assumed' row in SPSS even when Levene's test is significant.
Reporting p-values without an effect size and descriptives (M, SD) for each group.
Using a t-test on a single ordinal item; t-tests suit continuous scores such as averaged multi-item scales.
Treating p ≥ .05 as evidence that the groups are equal — absence of evidence is not evidence of absence.
Median-splitting a continuous predictor to force two groups — this throws away information; use regression instead.