Models a continuous outcome as a straight-line function of one predictor — e.g., predicting sales performance from training hours.
Predict & ExplainBivariatealso known as: Bivariate regression, OLS with one predictor
✓ When to use
One continuous outcome and one predictor, with a directional question: how much does Y change per unit of X?
You want predictions (ŷ for a new x) or an interpretable slope, not just an association.
A stepping stone to multiple regression.
✗ When NOT to use
You only need association strength — Pearson's r says it with less machinery (in simple regression, β = r).
The scatterplot is curved — add polynomial terms or transform.
The outcome is binary/ordinal/count — use the appropriate generalized model.
Clustered data (employees in firms) — mixed models.
Data requirements
Dependent / outcome variable
One continuous variable.
Independent / grouping variable
One continuous (or dummy-coded dichotomous) predictor.
Design
One sample, both variables per case; independent observations.
Sample size guidance
n ≥ 50 for stable estimates as a practical floor; power depends on expected R².
Assumptions
Linearity — the X–Y relation is straight (check scatter/residual plot).
Independence of residuals.
Homoscedasticity — residual spread constant across fitted values.
Normality of residuals (not of the raw variables).
No influential outliers (Cook's distance).
Hypotheses
H₀ — The population slope is zero (β₁ = 0) — X does not linearly predict Y.
H₁ — The slope differs from zero (β₁ ≠ 0).
The concept
Ordinary least squares picks the line ŷ = b₀ + b₁x that minimizes squared vertical distances to the points. b₁ is the expected change in Y per one-unit increase in X; b₀ is the expected Y at X = 0 (meaningful only if X = 0 is meaningful — consider centering).
R² is the fraction of Y's variance the line reproduces; in the one-predictor case R² = r². The t-test on b₁ and the model F-test are equivalent here. Prediction intervals for individuals are much wider than confidence intervals for the mean line — quote the right one.
Worked example
Predicting quarterly sales (₹ lakh) from training hours across 60 sales reps: b₁ = 0.42, meaning each extra training hour is associated with ₹42,000 more in sales; b₀ = 8.1.
Result: b₁ = 0.42, SE = 0.11, t(58) = 3.82, p < .001, R² = .20 — training hours explain 20% of the variance in sales.
How to run it
model <- lm(sales ~ training_hours, data = df)
summary(model) # coefficients, R², F
confint(model) # CIs for b0, b1
par(mfrow = c(2, 2)); plot(model) # residual diagnostics
import statsmodels.formula.api as smf
model = smf.ols("sales ~ training_hours", data=df).fit()
print(model.summary()) # coefficients, CIs, R², diagnostics
# residual plot
import matplotlib.pyplot as plt
plt.scatter(model.fittedvalues, model.resid); plt.axhline(0)
Analyze → Regression → Linear.
Outcome into 'Dependent', predictor into 'Independent(s)'.
Statistics: tick Estimates, Confidence intervals, Model fit.
Plots: ZRESID vs ZPRED for homoscedasticity; tick Histogram and Normal probability plot of residuals.
Report b (with SE and CI), β, t, p, and R² with the model F.
Scatterplot with trendline: Insert → Scatter → Chart Elements → Trendline → Display Equation and R².
Formal output: Data → Data Analysis → Regression; set Y range and X range; tick Residual plots.
Slope/intercept in cells: =SLOPE(Y,X), =INTERCEPT(Y,X); =RSQ(Y,X).
Prediction: =FORECAST.LINEAR(x_new, Y, X).
Interpreting the output
b₁: expected change in Y per unit X — interpret in real units.
p and CI for b₁: does the slope credibly differ from zero?
R²: share of outcome variance explained; modest values are normal in behavioral data.
Check residual plots before trusting any of the above.
Association-based slopes are not causal effects without design support.
APA-style reporting
Simple linear regression indicated that training hours significantly predicted quarterly sales, b = 0.42, SE = 0.11, β = .44, t(58) = 3.82, p < .001, R² = .20, F(1, 58) = 14.6.
Common mistakes
Fitting a line to a curve.
Interpreting the intercept when X = 0 is impossible.
Extrapolating predictions beyond the observed X range.
Confusing the CI of the mean with the (much wider) prediction interval for individuals.