Statistical Methods Atlas
Atlas › Predict & Explain › Simple Linear Regression

Simple Linear Regression

Models a continuous outcome as a straight-line function of one predictor — e.g., predicting sales performance from training hours.

Predict & ExplainBivariate also known as: Bivariate regression, OLS with one predictor

✓ When to use

  • One continuous outcome and one predictor, with a directional question: how much does Y change per unit of X?
  • You want predictions (ŷ for a new x) or an interpretable slope, not just an association.
  • A stepping stone to multiple regression.

✗ When NOT to use

  • You only need association strength — Pearson's r says it with less machinery (in simple regression, β = r).
  • The scatterplot is curved — add polynomial terms or transform.
  • The outcome is binary/ordinal/count — use the appropriate generalized model.
  • Clustered data (employees in firms) — mixed models.

Data requirements

Dependent / outcome variableOne continuous variable.
Independent / grouping variableOne continuous (or dummy-coded dichotomous) predictor.
DesignOne sample, both variables per case; independent observations.
Sample size guidancen ≥ 50 for stable estimates as a practical floor; power depends on expected R².

Assumptions

Hypotheses

H₀ — The population slope is zero (β₁ = 0) — X does not linearly predict Y.
H₁ — The slope differs from zero (β₁ ≠ 0).

The concept

Ordinary least squares picks the line ŷ = b₀ + b₁x that minimizes squared vertical distances to the points. b₁ is the expected change in Y per one-unit increase in X; b₀ is the expected Y at X = 0 (meaningful only if X = 0 is meaningful — consider centering).

R² is the fraction of Y's variance the line reproduces; in the one-predictor case R² = r². The t-test on b₁ and the model F-test are equivalent here. Prediction intervals for individuals are much wider than confidence intervals for the mean line — quote the right one.

Worked example

Predicting quarterly sales (₹ lakh) from training hours across 60 sales reps: b₁ = 0.42, meaning each extra training hour is associated with ₹42,000 more in sales; b₀ = 8.1.

Result: b₁ = 0.42, SE = 0.11, t(58) = 3.82, p < .001, R² = .20 — training hours explain 20% of the variance in sales.

How to run it

model <- lm(sales ~ training_hours, data = df)
summary(model)               # coefficients, R², F
confint(model)               # CIs for b0, b1

par(mfrow = c(2, 2)); plot(model)   # residual diagnostics

Interpreting the output

APA-style reporting

Simple linear regression indicated that training hours significantly predicted quarterly sales, b = 0.42, SE = 0.11, β = .44, t(58) = 3.82, p < .001, R² = .20, F(1, 58) = 14.6.

Common mistakes

Related methods

Pearson CorrelationSymmetric associationMultiple Linear RegressionSeveral predictorsBinary Logistic RegressionBinary outcome
← Chi-Square Goodness-of-Fit TestMultiple Linear Regression →