Measures the strength and direction of the linear relationship between two continuous variables — e.g., does job autonomy rise with job satisfaction?
Test AssociationBivariatealso known as: Pearson's r, Product-moment correlation
✓ When to use
Two continuous variables and a question about linear association.
Preliminary screening before regression or SEM (correlation matrices).
Both variables approximately normal, relationship linear on a scatterplot, no dominating outliers.
✗ When NOT to use
The relationship is curved — r understates or misses it; transform or model the curve.
Ordinal data or serious outliers — use Spearman's rho.
One variable is dichotomous — use point-biserial (a special case) or logistic regression.
You want to claim causation — correlation cannot deliver it.
Data are clustered/nested (employees within firms) — the independence assumption fails.
Data requirements
Dependent / outcome variable
Two continuous variables (no DV/IV distinction — the measure is symmetric).
Independent / grouping variable
—
Design
One sample with both variables measured per case; independent observations.
Sample size guidance
n ≥ 30 for minimally stable estimates; n ≈ 84 to detect r = .30 at power .80, α = .05. CIs shrink slowly — precision needs hundreds.
Assumptions
Independent observations (one row per unrelated case).
Linearity — inspect the scatterplot first; r only summarizes straight-line association.
Bivariate normality for accurate p-values and CIs (robust with larger n).
No extreme outliers — a single point can manufacture or destroy a correlation.
Unrestricted range — range restriction (e.g., only high performers sampled) attenuates r.
Hypotheses
H₀ — The population correlation is zero (ρ = 0).
H₁ — The population correlation differs from zero (ρ ≠ 0).
The concept
r standardizes the covariance of two variables by their SDs, yielding a pure number from −1 to +1: the sign gives direction, the magnitude gives strength of linear association. r² is the proportion of variance in one variable linearly shared with the other.
Conventional benchmarks (Cohen): .10 small, .30 medium, .50 large — though in organizational research observed correlations above .50 between distinct constructs may signal common-method variance rather than a strong causal link. Always look at the scatterplot: Anscombe's quartet shows four wildly different datasets with identical r.
Worked example
In a survey of 156 employees, job autonomy and job satisfaction (both 1–5 composites) correlate r = .42.
Result: r(154) = .42, p < .001, 95% CI [.28, .54] — a medium-to-large positive association: more autonomous employees tend to be more satisfied (r² = .18 shared variance).
How to run it
plot(df$autonomy, df$satisfaction) # always look first
cor.test(df$autonomy, df$satisfaction) # r, t-test, 95% CI
# correlation matrix with p-values
library(Hmisc)
rcorr(as.matrix(df[, c("autonomy", "satisfaction", "commitment")]))
import pingouin as pg
import seaborn as sns
sns.scatterplot(data=df, x="autonomy", y="satisfaction")
print(pg.corr(df["autonomy"], df["satisfaction"])) # r, CI, p, power
print(df[["autonomy", "satisfaction", "commitment"]].corr().round(2))
Graphs → Chart Builder → Scatter/Dot first, to check linearity.
Analyze → Correlate → Bivariate; move both variables in; tick Pearson, Two-tailed, Flag significant.
Options: Means and standard deviations.
Report r, n (or df = n − 2), p, and preferably the 95% CI (SPSS 27+ prints it via Confidence Interval option).
Scatterplot: select both columns → Insert → Scatter.
r in a cell: =PEARSON(A2:A157, B2:B157) or =CORREL(...).
t statistic: =r*SQRT(n−2)/SQRT(1−r^2); p-value: =T.DIST.2T(ABS(t), n−2).
Or Data → Data Analysis → Correlation for a matrix.
Interpreting the output
Sign = direction; magnitude = strength of linear association (benchmarks .10/.30/.50).
p tests only whether ρ = 0; with large n, tiny correlations become 'significant' — judge magnitude, not stars.
r² = shared variance; r = .30 means just 9% shared.
The 95% CI conveys precision — wide intervals in small samples are the norm.
Correlation ≠ causation: third variables and reverse causality remain live possibilities.
APA-style reporting
Job autonomy was positively correlated with job satisfaction, r(154) = .42, 95% CI [.28, .54], p < .001, indicating that employees with greater autonomy tended to report higher satisfaction.
Common mistakes
Skipping the scatterplot and missing curvature or a single influential outlier.
Interpreting significance as importance in large samples.
Causal language ('autonomy increases satisfaction') from cross-sectional correlation.
Comparing correlations across samples without a formal test (Fisher z).