Statistical Methods Atlas
Atlas › Test Association › Pearson Correlation

Pearson Correlation

Measures the strength and direction of the linear relationship between two continuous variables — e.g., does job autonomy rise with job satisfaction?

Test AssociationBivariate also known as: Pearson's r, Product-moment correlation

✓ When to use

  • Two continuous variables and a question about linear association.
  • Preliminary screening before regression or SEM (correlation matrices).
  • Both variables approximately normal, relationship linear on a scatterplot, no dominating outliers.

✗ When NOT to use

  • The relationship is curved — r understates or misses it; transform or model the curve.
  • Ordinal data or serious outliers — use Spearman's rho.
  • One variable is dichotomous — use point-biserial (a special case) or logistic regression.
  • You want to claim causation — correlation cannot deliver it.
  • Data are clustered/nested (employees within firms) — the independence assumption fails.

Data requirements

Dependent / outcome variableTwo continuous variables (no DV/IV distinction — the measure is symmetric).
Independent / grouping variable—
DesignOne sample with both variables measured per case; independent observations.
Sample size guidancen ≥ 30 for minimally stable estimates; n ≈ 84 to detect r = .30 at power .80, α = .05. CIs shrink slowly — precision needs hundreds.

Assumptions

Hypotheses

H₀ — The population correlation is zero (ρ = 0).
H₁ — The population correlation differs from zero (ρ ≠ 0).

The concept

r standardizes the covariance of two variables by their SDs, yielding a pure number from −1 to +1: the sign gives direction, the magnitude gives strength of linear association. r² is the proportion of variance in one variable linearly shared with the other.

Conventional benchmarks (Cohen): .10 small, .30 medium, .50 large — though in organizational research observed correlations above .50 between distinct constructs may signal common-method variance rather than a strong causal link. Always look at the scatterplot: Anscombe's quartet shows four wildly different datasets with identical r.

Worked example

In a survey of 156 employees, job autonomy and job satisfaction (both 1–5 composites) correlate r = .42.

Result: r(154) = .42, p < .001, 95% CI [.28, .54] — a medium-to-large positive association: more autonomous employees tend to be more satisfied (r² = .18 shared variance).

How to run it

plot(df$autonomy, df$satisfaction)          # always look first

cor.test(df$autonomy, df$satisfaction)       # r, t-test, 95% CI

# correlation matrix with p-values
library(Hmisc)
rcorr(as.matrix(df[, c("autonomy", "satisfaction", "commitment")]))

Interpreting the output

APA-style reporting

Job autonomy was positively correlated with job satisfaction, r(154) = .42, 95% CI [.28, .54], p < .001, indicating that employees with greater autonomy tended to report higher satisfaction.

Common mistakes

Related methods

Spearman Rank CorrelationOrdinal / non-normal dataPartial CorrelationControl a third variablePoint-Biserial CorrelationOne dichotomous variableSimple Linear RegressionPredict one variable from the other
← Friedman TestSpearman Rank Correlation →