Models a series' own past values and shocks to forecast its future — the classic statistical workhorse for univariate forecasting.
Analyze Time SeriesMultivariatealso known as: Box–Jenkins method, ARIMA(p,d,q), SARIMA
✓ When to use
Forecasting a single series (demand, headcount, revenue) from its own history.
Series with autocorrelation left after trend/seasonality handling.
Benchmarking: ARIMA (or ETS) is the baseline any fancier forecast must beat.
Seasonal variants (SARIMA) for calendar-patterned business data.
✗ When NOT to use
Fewer than ~50 observations — estimates are fragile; use simple ETS or naive benchmarks.
Known future drivers matter (promotions, price) — ARIMAX/regression with ARIMA errors or ML approaches.
Series dominated by structural breaks — model segments or use intervention terms.
Long-horizon forecasts — ARIMA reverts to the mean/trend; uncertainty balloons.
Data requirements
Dependent / outcome variable
One numeric series, regular intervals, ≥ 50 points (more for seasonal models).
Independent / grouping variable
Its own lags and lagged forecast errors (plus seasonal lags in SARIMA).
Design
Longitudinal; hold out the final segment for forecast validation.
Sample size guidance
3+ seasonal cycles for SARIMA; more history → better identified models.
Assumptions
Stationarity after differencing (constant mean/variance; check ADF/KPSS tests).
Residuals behave as white noise: no leftover autocorrelation (Ljung–Box), roughly normal, constant variance.
No unmodeled structural breaks.
Variance stabilization (log/Box–Cox) applied when swings grow with level.
Hypotheses
H₀ — Diagnostic nulls: unit root present (ADF); residuals independent (Ljung–Box — here you want NOT to reject).
H₁ — Series stationary (ADF rejection); residual autocorrelation remains (Ljung–Box rejection = model inadequate).
The concept
ARIMA(p,d,q) says: difference the series d times until stationary, then explain each value by p of its own past values (AR — momentum/mean reversion) and q past forecast errors (MA — shock absorption). SARIMA adds the same trio at the seasonal lag. The Box–Jenkins cycle is identify (ACF/PACF, unit-root tests) → estimate → check diagnostics → forecast; modern practice automates order selection by AICc (auto.arima) but keeps the human on diagnostics.
Honest evaluation is out-of-sample: hold out the last 6–12 periods, compare MAE/MAPE/RMSE against naive and seasonal-naive benchmarks. Always report prediction intervals — a point forecast without its widening uncertainty cone invites false confidence.
Worked example
Forecasting monthly customer-support tickets from 60 months of history. Log transform stabilizes variance; auto-selection lands on SARIMA(1,1,1)(0,1,1)₁₂.
Diagnostics clean (Ljung–Box p = .48). Holdout (last 12 months): MAPE 6.8% vs seasonal-naive 11.2%. Twelve-month forecast with 95% intervals feeds the staffing plan.
How to run it
library(forecast)
y <- ts(df$tickets, frequency = 12, start = c(2021, 1))
fit <- auto.arima(log(y)) # AICc-based order selection
summary(fit)
checkresiduals(fit) # Ljung-Box + plots
fc <- forecast(fit, h = 12)
plot(fc)
accuracy(fc) # or split train/test for holdout MAPE
import pandas as pd
from statsmodels.tsa.statespace.sarimax import SARIMAX
import pmdarima as pm # pip install pmdarima
y = pd.Series(df["tickets"].values,
index=pd.date_range("2021-01", periods=60, freq="MS"))
auto = pm.auto_arima(y, seasonal=True, m=12, trace=True)
print(auto.summary())
model = SARIMAX(y, order=auto.order, seasonal_order=auto.seasonal_order).fit()
print(model.summary())
fc = model.get_forecast(12)
print(fc.summary_frame()) # mean + CI
Define dates (Data → Define date and time).
Analyze → Forecasting → Create Traditional Models; Method: Expert Modeler (considers ARIMA and ETS) or ARIMA with chosen orders.
Criteria: set forecast horizon; Statistics: Parameter estimates, Ljung-Box; Plots: fit and forecasts with CIs.
Report the selected model, Ljung-Box Q (want p > .05), fit stats (RMSE/MAPE), and forecasts with intervals.
Native option: =FORECAST.ETS (exponential smoothing with seasonality — not ARIMA but a respectable business alternative), plus FORECAST.ETS.CONFINT for intervals; or Data → Forecast Sheet.
True ARIMA is not built in; use R/Python/SPSS when ARIMA specifically is required.
Always keep a holdout: compare forecast vs actual with =ABS(A−F)/A averaged (MAPE).
Interpreting the output
Model orders and any transformation, with the selection criterion (AICc).
Out-of-sample MAE/MAPE vs naive benchmarks — the credibility test.
Forecasts WITH prediction intervals; emphasize widening uncertainty.
Refit as new data arrive; ARIMA models age.
APA-style reporting
A SARIMA(1,1,1)(0,1,1)₁₂ model on log-transformed ticket counts (selected by AICc) showed adequate diagnostics, Ljung–Box Q(18) = 17.6, p = .48. Out-of-sample accuracy over the final 12 months (MAPE = 6.8%) outperformed a seasonal-naive benchmark (11.2%). Twelve-month forecasts with 95% prediction intervals are reported in Figure 2.
Common mistakes
Skipping stationarity checks and differencing blindly (or twice too often).
Judging the model on in-sample fit only.
Ignoring Ljung–Box failure.
Reporting point forecasts without intervals.
Feeding a series with a known structural break into one homogeneous model.