The correlation between one true dichotomy and one continuous variable — mathematically equivalent to an independent t-test, e.g., gender and salary.
Test AssociationBivariatealso known as: rpb
✓ When to use
One naturally dichotomous variable (yes/no, member/non-member) and one continuous variable.
Item analysis in test construction: item (correct/incorrect) vs total score.
You want an effect-size-style summary of a two-group difference on a correlation metric.
✗ When NOT to use
The dichotomy was created by artificially splitting a continuous variable — use the biserial correlation or, better, keep the variable continuous.
Both variables continuous — Pearson.
The continuous variable is heavily non-normal — consider Mann–Whitney U and rank-biserial correlation.
More than two groups — ANOVA/η².
Data requirements
Dependent / outcome variable
One continuous variable.
Independent / grouping variable
One genuinely dichotomous variable (coded 0/1).
Design
Between-subjects; independent observations.
Sample size guidance
As for the independent t-test; very unequal group proportions (e.g., 95/5) depress the maximum attainable rpb.
Assumptions
Independence of observations.
Normality of the continuous variable within each group (for inference).
The dichotomy is a true category, not a split continuum.
Hypotheses
H₀ — No association between group membership and the continuous variable (ρpb = 0).
H₁ — An association exists (equivalently: the two group means differ).
The concept
Code the dichotomy 0/1 and compute an ordinary Pearson correlation with the continuous variable — that is the point-biserial. It is algebraically linked to the t-test: rpb = √(t²/(t² + df)), so the two analyses always agree on significance; rpb simply re-expresses the group difference as variance explained (rpb² is the η² of the two-group comparison).
Its magnitude depends on the group split: with a 50/50 split rpb can reach 1, but skewed splits cap it well below 1 — compare observed values against what the split allows.
Worked example
Relating union membership (0/1; 40% members) to monthly salary among 210 workers.
Result: rpb = .21, p = .002 — members earn somewhat more; equivalently t(208) = 3.10 with d ≈ 0.43.
How to run it
cor.test(df$union, df$salary) # union coded 0/1: Pearson = point-biserial
# equivalence with t-test:
t.test(salary ~ union, data = df, var.equal = TRUE)
from scipy import stats
r, p = stats.pointbiserialr(df["union"], df["salary"])
print(r, p)
Code the dichotomy 0/1.
Analyze → Correlate → Bivariate with the 0/1 variable and the continuous variable; Pearson output IS the point-biserial.
Report rpb, n, p (and the group means for context).
p-value via t = r*SQRT(n−2)/SQRT(1−r^2), =T.DIST.2T(ABS(t), n−2).
Interpreting the output
Sign depends on which group is coded 1 — state the coding.
rpb² = proportion of variance in the continuous variable explained by group membership.
Report group means/SDs so the correlation has concrete meaning.
Judge magnitude with the group split in mind.
APA-style reporting
Union membership was positively associated with monthly salary, rpb = .21, p = .002, n = 210; members (M = ₹42,300, SD = 8,100) earned more than non-members (M = ₹38,900, SD = 7,600).
Common mistakes
Median-splitting a continuous variable and correlating the halves.
Ignoring the coding direction when interpreting the sign.
Comparing rpb across samples with different group splits.
Running both a t-test and rpb as if they were independent pieces of evidence — they are the same test.