Statistical Methods Atlas
Atlas › Analyze Time Series › ARIMA Forecasting

ARIMA Forecasting

Models a series' own past values and shocks to forecast its future — the classic statistical workhorse for univariate forecasting.

Analyze Time SeriesMultivariate also known as: Box–Jenkins method, ARIMA(p,d,q), SARIMA

✓ When to use

  • Forecasting a single series (demand, headcount, revenue) from its own history.
  • Series with autocorrelation left after trend/seasonality handling.
  • Benchmarking: ARIMA (or ETS) is the baseline any fancier forecast must beat.
  • Seasonal variants (SARIMA) for calendar-patterned business data.

✗ When NOT to use

  • Fewer than ~50 observations — estimates are fragile; use simple ETS or naive benchmarks.
  • Known future drivers matter (promotions, price) — ARIMAX/regression with ARIMA errors or ML approaches.
  • Series dominated by structural breaks — model segments or use intervention terms.
  • Long-horizon forecasts — ARIMA reverts to the mean/trend; uncertainty balloons.

Data requirements

Dependent / outcome variableOne numeric series, regular intervals, ≥ 50 points (more for seasonal models).
Independent / grouping variableIts own lags and lagged forecast errors (plus seasonal lags in SARIMA).
DesignLongitudinal; hold out the final segment for forecast validation.
Sample size guidance3+ seasonal cycles for SARIMA; more history → better identified models.

Assumptions

Hypotheses

H₀ — Diagnostic nulls: unit root present (ADF); residuals independent (Ljung–Box — here you want NOT to reject).
H₁ — Series stationary (ADF rejection); residual autocorrelation remains (Ljung–Box rejection = model inadequate).

The concept

ARIMA(p,d,q) says: difference the series d times until stationary, then explain each value by p of its own past values (AR — momentum/mean reversion) and q past forecast errors (MA — shock absorption). SARIMA adds the same trio at the seasonal lag. The Box–Jenkins cycle is identify (ACF/PACF, unit-root tests) → estimate → check diagnostics → forecast; modern practice automates order selection by AICc (auto.arima) but keeps the human on diagnostics.

Honest evaluation is out-of-sample: hold out the last 6–12 periods, compare MAE/MAPE/RMSE against naive and seasonal-naive benchmarks. Always report prediction intervals — a point forecast without its widening uncertainty cone invites false confidence.

Worked example

Forecasting monthly customer-support tickets from 60 months of history. Log transform stabilizes variance; auto-selection lands on SARIMA(1,1,1)(0,1,1)₁₂.

Diagnostics clean (Ljung–Box p = .48). Holdout (last 12 months): MAPE 6.8% vs seasonal-naive 11.2%. Twelve-month forecast with 95% intervals feeds the staffing plan.

How to run it

library(forecast)
y <- ts(df$tickets, frequency = 12, start = c(2021, 1))

fit <- auto.arima(log(y))       # AICc-based order selection
summary(fit)
checkresiduals(fit)              # Ljung-Box + plots

fc <- forecast(fit, h = 12)
plot(fc)
accuracy(fc)                     # or split train/test for holdout MAPE

Interpreting the output

APA-style reporting

A SARIMA(1,1,1)(0,1,1)₁₂ model on log-transformed ticket counts (selected by AICc) showed adequate diagnostics, Ljung–Box Q(18) = 17.6, p = .48. Out-of-sample accuracy over the final 12 months (MAPE = 6.8%) outperformed a seasonal-naive benchmark (11.2%). Twelve-month forecasts with 95% prediction intervals are reported in Figure 2.

Common mistakes

Related methods

Trend Analysis & DecompositionDescribe before forecastingMultiple Linear RegressionAdd external drivers (ARIMAX)Simple Linear RegressionNaive trend baseline
← Trend Analysis & DecompositionShapiro–Wilk Test →