Forecasting Fundamentals

Module 6 · Point forecasts, prediction intervals, and loss functions

Gary Cornwall

Econ 6376 · The George Washington University

Where we are: five modules, one chain

Module 1Time series have memoryThe past is informative about the present. ϕ\phi measures how long a shock lasts.
→
Modules 1–2StationarityA verdict built from evidence, not a measurement: the plot, the ACF, ADF with KPSS, the rule of thumb.
→
Module 3The ACF and PACF are the keys to the processEvery ARMA class leaves a fingerprint: which one tails off, which one cuts off.
→
Module 4Model selection criteriaWhen both tail off, information criteria score the candidates: fit minus penalty.
→
Modules 4–5Box-Jenkins general-to-specific ARMA modelingA procedure, not a pick. Start one size bigger, check the residuals, then trim.
→
Module 5Serial correlationA property of the fitted model, not of the data. Residuals ete_t should behave like innovations ϵt\epsilon_t.

yt=α+δt+∑j=1pϕjyt−j+∑l=1qθlϵt−l+ϵty_t=\alpha+\delta t+\sum_{j=1}^{p}\phi_j y_{t-j}+\sum_{l=1}^{q}\theta_l\epsilon_{t-l}+\epsilon_t

Every knob on the mixing board is on. Module 6 adds no new slider: it uses the fitted equation to project forward.

Before we forecast anything: the whole chain, run once, on one real series.

Know your data: real GDP per capita

Real gross domestic product per capitaFRED A939RX0Q048SBEA · chained (2017) dollars per person · as published September 30, 2026
real GDP per capita=real GDP, chained (2017) dollarsmidperiod population\text{real GDP per capita}=\dfrac{\text{real GDP, chained (2017) dollars}}{\text{midperiod population}}
Who publishes itBureau of Economic AnalysisNIPA Table 7.1, line 10.Numerator: Table 1.1.6, line 1. Denominator: Table 7.1, line 18.
FrequencyQuarterlyQuarterly from 1947 Q1; annual from 1929.Here: 1990 Q1 to 2006 Q4, 68 quarters.
PopulationCensus Bureau estimatesTotal population, including Armed Forces overseas and the institutionalized population.Quarterly figure: the average of the months.
DeflatorNo single deflatorBuilt bottom-up. Most detailed components: BLS price indexes (CPI, PPI, International Price Indexes).Aggregated with a Fisher chain-type index. The implicit price deflator is a by-product.
RevisionsAdvance, second, thirdAn annual update each September; a comprehensive update about every five years.Population revises on its own calendar.
Seasonal adjustmentSeasonally adjusted at annual ratesIndirect: components are adjusted, then aggregated.Population is not seasonally adjusted.
AccountingAn aggregateGDP=C+I+G+(X−M)\text{GDP}=C+I+G+\text{(}X-M\text{)}GDP=GDI+statistical discrepancy\text{GDP}=\text{GDI}+\text{statistical discrepancy}Chained-dollar components are not additive: the tables carry a “Residual” line.
Guiding documentsSNA 2008 · NIPA HandbookThe NIPAs are “generally consistent with the SNA”.BEA: Concepts and Methods of the U.S. National Income and Product Accounts.

A ratio of two estimates. Both get revised; only one is seasonally adjusted.

The chain starts: is it stationary?

yty_t is the natural log of real GDP per capita. Why logs: constant growth is a straight line, and Δyt\Delta y_t is the quarterly growth rate. The fitted trend grows 2.17% a year.

ADF · trendH0H_0: unit root2 lags, chosen by BIC
Alternative: stationary around a trend
τ\tau against the 5% critical value−3.11 vs −3.45
Not below the critical value: fail to reject a unit root
KPSS · trendH0H_0: stationary around a trend
Statistic, short truncation0.150 vs 0.146
Statistic, long truncation0.085 vs 0.146
Above by a hair with one truncation, below with the other: no clear rejection
Rule of thumbNo null: a persistence flag
R=σΔy/σyR=\sigma_{\Delta y}/\sigma_y against the threshold0.050 vs 0.5
Far below 0.5, but the trend itself inflates σy\sigma_y: this row cannot separate the two readings by itself

Neither null is clearly rejected. Can 68 quarters tell a trend from a unit root?

Trend-stationary, or a unit root?

DeviationsHow persistent are they?
Rule of thumb on the deviations, σΔy/σû\sigma_{\Delta y}/\sigma_{\hat u} vs 0.50.359
AR(1) coefficient of ût\hat u_t (descriptive)0.879
Half-life of a deviation (descriptive)5.4 quarters
RobustnessRejected at 5%? Lags 0, 1, 2, 4, 8
Unit root, ADF: τ\tau from −3.11 to −1.95 vs −3.45no
Unit root, DF-GLS: from −2.11 to −1.39 vs −3.03no
Trend-stationarity, KPSS: 0.150 short, 0.085 long vs 0.146split

No lag choice rejects a unit root. With 68 quarters that is low power (Module 2: near the boundary these tests rarely reject), not evidence of a unit root. KPSS sits on its critical value. The deviations are persistent, yet keep returning to the line: the borderline case.

A choice, not a finding. 68 quarters cannot separate the two readings. The course’s rule under doubt is to difference: treat yty_t as I(1)I\text{(}1\text{)}, model the growth rate Δyt\Delta y_t, then check that fit for over-differencing.

Why it matters for forecasting: under a deterministic trend the forecast returns to the line and the interval stops widening; under a unit root it does not.

Was the first difference enough?

ADF · driftH0H_0: unit root1 lag, chosen by BIC
τ\tau against the 5% critical value−3.83 vs −2.89
Below the critical value: reject a unit root
One lag choice disagrees. At 0, 1, 2 and 4 lags ADF and DF-GLS both reject; at 8 lags neither does (ADF −2.16 vs −2.89).
KPSS · levelH0H_0: stationary around a constant mean
Statistic, short truncation0.214 vs 0.463
Statistic, long truncation0.166 vs 0.463
Below the critical value with both: fail to reject stationarity
Rule of thumbNo null: a persistence flag
R=σΔ2y/σΔyR=\sigma_{\Delta^2 y}/\sigma_{\Delta y} against the threshold1.186 vs 0.5
Above 0.5: no second difference needed. Keep Δyt\Delta y_t

One difference is enough; the only dissent is the 8-lag test. Δyt\Delta y_t is the series we model.

No sign of over-differencing so far: ρ̂1=0.30\hat\rho_1 = 0.30, where an MA(1) with θ=−1\theta=-1 would give −0.5-0.5. The real check is θ̂\hat\theta near −1-1 once we fit.

The series we model: ACF and PACF of Δyt\Delta y_t

Both tail off. The ACF clears the band at lags 1 and 2; the PACF at lag 1, with lag 2 (0.22) just inside; then both fade. The ACF decays no faster than the PACF.

Module 3’s table: ACF tails off, PACF cuts off, means AR; both tail off means ARMA. Nothing here cuts off cleanly, so AR or ARMA, not MA.

The general choice is ARMA: ARMA(p,qp,q) nests AR(pp) by setting the θ\theta’s to zero, so it costs parameters, not possibilities.

Start one size bigger, then trim (Module 5’s rule): fit the larger model and let the criteria and the residuals cut it down.

Starting model: ARIMA(2,1,2) with drift

ARMA(2,2) on Δyt\Delta y_t with a constant. The constant is the drift: growth averages 0.47% a quarter, and a model without it would forecast zero growth.

Information criteria: one scoreboard, four tariffs

IC=−2logL(𝛝̂)+penalty(k,T)\text{IC} = -2\log L\text{(}\hat{\boldsymbol{\vartheta}}\text{)} + \text{penalty(}k, T\text{)}

Fit first: a bigger log-likelihood is a smaller deviance. Then the tariff for complexity. Smaller IC is better.

AIC (Akaike, 1974)2k2k
AICc (Hurvich and Tsai, 1989)2k+2k(k+1)T−k−12k + \dfrac{2k\text{(}k+1\text{)}}{T-k-1}
BIC (Schwarz, 1978)klogTk\log T
HQIC (Hannan and Quinn, 1979)2kloglogT2k\log\log T

kk counts every estimated parameter: the AR and MA coefficients, the drift, and σ2\sigma^2. TT is the number of observations the fit uses.

Per parameter at T=75T = 75: BIC log75=4.32\log 75 = 4.32, HQIC 2loglog75=2.932\log\log 75 = 2.93, AIC 22; AICc adds a small-sample term. BIC punishes parameters hardest at this TT, so it favours the smallest models.

The candidates, scored

library(forecast)
source("helpers/model_selection.R")     # hqic()
y   <- log(gdp_pc)                      # 1990 Q1 to 2006 Q4
fit <- function(p, q)
  Arima(y, order = c(p, 1, q), include.drift = TRUE)

fits <- list(m212 = fit(2, 2), m211 = fit(2, 1),
             m111 = fit(1, 1))
sapply(fits, function(f)
  c(k = length(coef(f)) + 1, logLik = f$loglik,
    AIC = f$aic, AICc = f$aicc, BIC = f$bic,
    HQIC = hqic(f)))

for (ic in c("aicc", "aic", "bic"))    # each minimiser
  print(auto.arima(y, d = 1, seasonal = FALSE,
                   stepwise = FALSE,
                   approximation = FALSE, ic = ic))

grid <- expand.grid(p = 0:5, q = 0:5)
grid <- grid[grid$p + grid$q <= 5, ]   # its grid
grid$HQIC <- mapply(function(p, q) hqic(fit(p, q)),
                    grid$p, grid$q)
grid[which.min(grid$HQIC), ]           # HQIC by hand
Model kk logL\log L AIC AICc BIC HQIC
ARIMA(2,1,2) 6 261.62 −511.25 −509.85 −498.02 −506.01
ARIMA(2,1,1) 5 261.28 −512.56 −511.58 −501.54 −508.20
ARIMA(1,1,1) 4 260.28 −512.56 −511.92 −503.74 −509.07
ARIMA(2,1,0) † 4 260.82 −513.63 −512.99 −504.81 −510.14
ARIMA(1,1,0) † 3 259.00 −512.00 −511.61 −505.38 −509.38
Picks. AIC, AICc and HQIC: ARIMA(2,1,0). BIC: ARIMA(1,1,0). Neither is one of the three we wrote down: † marks the rows auto.arima added. All five carry drift.
ARIMA(2,1,2): θ̂=(−1.65,0.65)\hat\theta = (-1.65,\ 0.65), an MA root on the unit circle, Module 3’s over-differencing warning light. auto.arima’s root screen drops it.

Four criteria, one candidate set: which model, and are its residuals clean?

You will be wrong

the question is by how much, and whether you said so in advance

Nine stopped clocks A three-by-three grid of analog clock faces, each stopped at a different time. One face is cracked, one has lost its minute hand, one has both hands drooping toward six-thirty, and one is tilted.
Even a stopped clock is right twice a day.proverb; earliest form Addison, The Spectator no. 129, 28 July 1711
“The moment you forecast you know you’re going to be wrong, you just don’t know when and in which direction.”from Fiedler’s forecasting rules, Across the Board, June 1977
“All models are wrong but some are useful”George E. P. Box, 1979, p. 202
“Wall Street indexes predicted nine out of the last five recessions!”Paul Samuelson, Newsweek, 19 September 1966
Our sample ends in 2006 Q4. What came next?
“At this juncture, however, the impact on the broader economy and financial markets of the problems in the subprime market seems likely to be contained.”Ben Bernanke, testimony before the Joint Economic Committee, 28 March 2007
“We simply do not know. Nevertheless, the necessity for action and for decision compels us as practical men to do our best to overlook this awkward fact.”Keynes, Quarterly Journal of Economics, February 1937, p. 214

A forecast has two parts

Origin ttwhere you stand
Horizon hhhow far ahead you look
Information set ℱt={yt,yt−1,…}\mathcal{F}_t=\{y_t, y_{t-1}, \ldots\}everything you know at tt
The point forecast. ŷt+h|t=𝔼[yt+h∣ℱt]\hat y_{t+h|t}=\mathbb{E}\!\left[y_{t+h}\mid\mathcal{F}_t\right]: the best guess under squared-error loss, because the conditional mean minimises expected squared error.
The forecast error. et+h|t=yt+h−ŷt+h|te_{t+h|t}=y_{t+h}-\hat y_{t+h|t}: unknowable at tt, which is the point. It is observable after the fact, so it is an ee, not an ϵ\epsilon.
The prediction interval. ŷt+h|t±ch\hat y_{t+h|t}\pm c_h with a stated coverage, 80% or 95%: the fine print, where the honesty lives.
“Between 3% and 12%”, or “4.1%, give or take 0.3”: which forecast is more likely to be right, and which is more useful?

The headline number is the point forecast; the interval is what keeps it honest. Both are wrong, in different ways: the point in the third decimal, the interval one time in twenty.

The AR(1) point forecast is a recursion

yt+1=α+ϕyt+ϵt+1y_{t+1}=\alpha+\phi\,y_t+\epsilon_{t+1}

ŷt+1|t=α+ϕytbecause𝔼[ϵt+1∣ℱt]=0\hat y_{t+1|t}=\alpha+\phi\,y_t \qquad\text{because}\qquad \mathbb{E}\!\left[\epsilon_{t+1}\mid\mathcal{F}_t\right]=0

ŷt+h|t=α+ϕŷt+h−1|t\hat y_{t+h|t}=\alpha+\phi\,\hat y_{t+h-1|t}

ŷt+h|t−μ=ϕh(yt−μ)withμ=α1−ϕ\hat y_{t+h|t}-\mu=\phi^h\,\text{(}y_t-\mu\text{)}\qquad\text{with}\qquad \mu=\frac{\alpha}{1-\phi}

Two rules. Replace future innovations with zero. Replace future yy’s with their forecasts. Each step feeds the previous forecast forward: that is why it is called recursive.
R’s intercept is the mean μ̂\hat\mu, not α̂\hat\alpha. Recover α̂=μ̂(1−ϕ̂)\hat\alpha=\hat\mu\,\text{(}1-\hat\phi\text{)} before you run the recursion, or run it in deviations from μ̂\hat\mu.
As h→∞h\to\infty, where does the forecast go?

hh 1 2 3 5 20
ŷT+h|T=0.7hyT\hat y_{T+h|T}=0.7^h\,y_T 1.086 0.760 0.532 0.261 0.001

Known parameters: ϕ=0.7\phi=0.7, σ2=1\sigma^2=1, α=μ=0\alpha=\mu=0; T=60T=60, origin yT=1.552y_T=1.552. A nonzero μ\mu only shifts everything: the forecast reverts to μ\mu instead of 00.

The AR(1) prediction interval

et+h|t=∑i=0h−1ϕiϵt+h−ie_{t+h|t}=\sum_{i=0}^{h-1}\phi^i\,\epsilon_{t+h-i}

The shocks you have not seen yet, each weighted by how long the model remembers it. In general et+h|t=∑i=0h−1ψiϵt+h−ie_{t+h|t}=\sum_{i=0}^{h-1}\psi_i\,\epsilon_{t+h-i} with Module 3’s ψ\psi-weights; for an AR(1), ψi=ϕi\psi_i=\phi^i.

Var(et+h|t)=σ2∑i=0h−1ϕ2i=σ21−ϕ2h1−ϕ2\text{Var(}e_{t+h|t}\text{)}=\sigma^2\sum_{i=0}^{h-1}\phi^{2i}=\sigma^2\,\frac{1-\phi^{2h}}{1-\phi^2}

At h=1h=1 it is σ2\sigma^2. As h→∞h\to\infty it is γ0=σ2/(1−ϕ2)\gamma_0=\sigma^2/\text{(}1-\phi^2\text{)}: as uncertain as knowing nothing.

ŷt+h|t±zα/2σ1−ϕ2h1−ϕ2\hat y_{t+h|t}\;\pm\;z_{\alpha/2}\,\sigma\sqrt{\frac{1-\phi^{2h}}{1-\phi^2}}

Gaussian innovations make the error Gaussian, so zα/2=1.96z_{\alpha/2}=1.96 for a 95% interval and 1.281.28 for 80%.

The 95% rests on ϵt∼N(0,σ2)\epsilon_t\sim N\text{(}0,\sigma^2\text{)}, the right model, and a known σ2\sigma^2. Break any of them and the interval is nominal, not delivered. The break comes later.

Var(eT+1|T)\text{Var(}e_{T+1|T}\text{)}1.000
Var(eT+5|T)\text{Var(}e_{T+5|T}\text{)}1.905
Var(eT+20|T)\text{Var(}e_{T+20|T}\text{)} against γ0\gamma_01.961 vs 1.961

Now move ϕ\phi and σ2\sigma^2 and watch the fan.

Play it out: an AR(1) forecast with your ϕ\phi and σ2\sigma^2

Stationary forecasts are short-run forecasts

Same series as slides 13 to 15, ϕ=0.7\phi=0.7 and μ=0\mu=0 known. One-step: ŷt+1|t=ϕyt\hat y_{t+1|t}=\phi\,y_t from each origin t=60,…,79t=60,\ldots,79. Two-step: ŷt+2|t=ϕ2yt\hat y_{t+2|t}=\phi^2\,y_t from t=59,…,78t=59,\ldots,78. Grey: the single path from T=60T=60.

Why. ŷt+h|t−μ=ϕh(yt−μ)\hat y_{t+h|t}-\mu=\phi^h\,\text{(}y_t-\mu\text{)}: the information in yty_t decays geometrically, so after a few steps the forecast is μ\mu and the interval is the unconditional one; here the single path is within 0.2 of μ\mu by h=6h=6. The forecast explains the share Var(ŷt+h|t)/γ0=ϕ2h\mathrm{Var}\text{(}\hat y_{t+h|t}\text{)}/\gamma_0=\phi^{2h} of the variance; the error carries the rest.
at ϕ=0.7\phi=0.7 h=1h=1 h=2h=2 h=4h=4 h=8h=8
explained, ϕ2h\phi^{2h} 49% 24% 6% 0.3%
left over, 1−ϕ2h1-\phi^{2h} 51% 76% 94% 99.7%
Persistence is the dial. Half the variance is still explained at h=ln0.5/(2lnϕ)h=\ln 0.5\,/\,\text{(}2\ln\phi\text{)}: about 1 step at ϕ=0.7\phi=0.7, about 7 at ϕ=0.95\phi=0.95. ϕ\phi is memory (Module 1).
Reading the picture. One-step forecasts track the series because each uses the latest observation; two-step forecasts track it less; the single path stops tracking after a few steps. Module 7 evaluates forecasts this way: rolling origin, not fixed origin (notes 6.3).

A stationary model forecasts the next few steps; after that it forecasts the mean. ϕ\phi sets how many steps “a few” is.

The example meets its test set: 2007 to 2011

How do we judge ourselves if we are always wrong?

A loss function is the cost of being wrong. g(e)g\text{(}e\text{)}, with e=yt+h−ŷt+h|te=y_{t+h}-\hat y_{t+h|t} the forecast error from “A forecast has two parts”. It is gg, not LL: LL is the lag operator.
Zero when the forecast is perfectg(0)=0g\text{(}0\text{)}=0
Larger as the miss growsa bigger |e||e| costs more
Chosen before you see the results, not after. The loss is a decision about what kind of errors hurt most; the choice comes first and the formula follows.

Compare within a loss, never between: same loss, same series, the same NN test-set observations.

The Texas grid. You forecast energy demand and a cold snap hits. Under-forecast, and there is not enough supply: people freeze. Over-forecast, and you ran extra generators: wasted fuel. Both are errors; they are not the same error. A symmetric loss says 500 MW under and 500 MW over are equally bad. They are not.
A hospital forecasts daily ER admissions to set its staffing. Too few staff is far more costly than too many. What should the loss look like, and what does the best forecast then represent?

Match the loss to the cost of being wrong

The marginal cost of one more unit of error picks the shape. Ask what each additional unit of miss does to you, then write the loss that charges for it.
Cost of being wrong Loss Formula
Explosive per-unit costeach unit of error does more damage than the last MSE 1N∑t=1N(yt−ŷt)2\frac{1}{N}\sum_{t=1}^{N}\text{(}y_t-\hat y_t\text{)}^2
Constant per-unit costan error of 10 is exactly twice an error of 5 MAE 1N∑t=1N|yt−ŷt|\frac{1}{N}\sum_{t=1}^{N}\lvert y_t-\hat y_t\rvert
Quadratic for small errors, linear for largeprecision in the normal range, robust to outliers Huber gδ(e)=12e2g_\delta\text{(}e\text{)}=\tfrac{1}{2}e^2 if |e|≤δ\lvert e\rvert\le\delta
δ(|e|−12δ)\delta\text{(}\lvert e\rvert-\tfrac{1}{2}\delta\text{)} if |e|>δ\lvert e\rvert>\delta
Directional costone direction costs more than the other Asymmetric α|e|\alpha\lvert e\rvert if e>0e>0 (under-prediction)
(1−α)|e|\text{(}1-\alpha\text{)}\lvert e\rvert if e≤0e\le 0 (over-prediction)

NN counts test-set observations, not the sample size TT. δ>0\delta>0 sets where Huber switches from quadratic to linear; α∈(0,1)\alpha\in\text{(}0,1\text{)} is the weight on under-prediction, and α=0.5\alpha=0.5 is MAE.

The height of each line is what the next unit of error costs. MSE’s rises without bound; MAE’s is flat; Huber’s rises, then stops at δ=2\delta=2; the asymmetric loss charges α=0.75\alpha=0.75 per unit of under-prediction and 0.250.25 per unit over.

The shapes

The quadratic explodes. MSE is the cost curve where tail events are catastrophic: off by 10 is four times as bad as off by 5, not twice.

The linear grows steadily. MAE: every unit of error costs the same.

Huber caps marginal influence. A quadratic near zero, a straight line beyond δ=2\delta=2: large errors are not costless, but no single outlier can dominate the evaluation.

The asymmetric loss charges more on one side. α=0.75\alpha=0.75 per unit of under-prediction, 0.250.25 per unit of over-prediction.

MAPE, 100N∑t=1N|yt−ŷtyt|\dfrac{100}{N}\sum_{t=1}^{N}\left\lvert\dfrac{y_t-\hat y_t}{y_t}\right\rvert, is scale-free but undefined at yt=0y_t=0 and asymmetric by construction. Rarely the right choice for a macro series.

The loss picks the target

Under MSE, the best point forecast is theconditional mean
Under MAEconditional median
Under asymmetric linear loss with weight α\alphathe α\alpha-quantile

The Texas operator at α=0.9\alpha=0.9 forecasts the 90th percentile of demand, not the mean. Symmetric shocks hide the difference; skew them and the loss decides:

set.seed(2020); eps <- rexp(600) - 1   # skewed right
y <- numeric(600)
for (t in 2:600) y[t] <- 0.6 * y[t - 1] + eps[t]
y_train <- window(ts(y), end = 350)
fit  <- Arima(y_train, order = c(1, 0, 0))
full <- Arima(ts(y), model = fit)      # no refit
fc_mean   <- window(fitted(full), start = 351)
fc_median <- fc_mean + median(residuals(fit)) # -0.39
Which forecast wins?
mean forecast median forecast
MSE 1.063 1.207
MAE 0.753 0.727
Each wins under the loss it was built for. Not a paradox: the loss doing its job.

Most shocks are small and negative, a few are large and positive. The mean track leans toward the expensive jumps; the median track sits where most observations land, 0.390.39 below it.

“A forecast has two parts” called the conditional mean the best guess under squared-error loss. This is why the qualifier was there.

Write your own; compare within

The standard losses are starting points, not laws. A loss function is any function of the actual and the forecast that is zero when the forecast is perfect and grows as the forecast gets worse. Beyond that, the shape is yours to choose.
A call centre: under-staffing by one costs $500 in dropped calls and escalates, over-staffing by one costs $100 in idle wages. Write the cost, not a textbook metric:
call_center_loss <- function(actual, forecast) {
  e <- actual - forecast
  cost <- ifelse(e > 0,
                 500 * e + 50 * e^2,   # under-staffed
                 100 * abs(e))         # over-staffed
  return(mean(cost))
}
Score the competing models with call_center_loss(actuals, fc_A$mean) and call_center_loss(actuals, fc_B$mean): the winner minimises your actual cost.
Within, never between. Two models are compared under one loss, on one series, over the same NN test-set observations. “A wins on MSE, B wins on MAE” is two questions with two answers, and the rule for the next slide.

Next: the GDP forecasts you just watched, scored under each loss. Module 7 asks whether the differences are significant.

The same forecasts, scored under each loss