Module 2: Testing for Stationarity

Econ 6376 — Applied Time Series Econometrics

How to Use These Notes

This chapter has one main reading path and three optional layers.

  • Core material is the main text. It contains the concepts, notation, interpretations, and applied skills expected of everyone.
  • Deeper Dive sections explain why a result works or develop it more fully. They are useful, but they can be skipped on a first reading.
  • Technical Note sections state qualifications or formal details that matter for precise reasoning.
  • Looking Ahead sections introduce an idea that will be taught formally in a later module.

If you missed lecture, read the main text, run the Core code, and complete the Core Practice problems. Then return to the optional sections that address your questions or interests.

Core Learning Objectives

By the end of this module, you should be able to:

  1. Apply the difference operator \(\Delta\) and compute first and second differences both algebraically and in R.
  2. Conduct visual inspection of a time series and its ACF to form a preliminary judgment about stationarity.
  3. Build the Dickey-Fuller test regression from the AR(1) equation and explain why it takes that specific form.
  4. Distinguish the three ADF specifications (no constant, drift only, drift + trend) and choose the appropriate one for a given series.
  5. Explain why the DF \(t\)-statistic does not follow a \(t\)-distribution under the null and demonstrate this via Monte Carlo.
  6. Use MacKinnon critical values to make a test decision and recognize the role of sample size.
  7. Recognize KPSS as a reverse-null test and use it as a complement to ADF (confirmatory analysis).
  8. Apply the rule-of-thumb ratio \(\sigma_{\Delta y} / \sigma_y\) and understand its underlying logic.
  9. Define the order of integration and recognize the over-differencing problem when a trend-stationary series is differenced.
  10. Synthesize multiple pieces of evidence into a stationarity judgment, erring toward differencing when in doubt.

2.1 Setup: Why We Are Doing This

Last module we ended on spurious regression — two completely independent random walks producing a beautiful \(R^2\) and a fake significant coefficient. Granger and Newbold’s (1974) result is the entire reason this module exists. If you ignore non-stationarity, your regression output will lie to you with confidence. The first defense against that is detecting non-stationarity before you run any regression that depends on the time series properties of your data.

Recall the boundary from Module 1:

  • \(|\phi| < 1\): stationary, finite variance, and eligible for the standard stationary modeling toolkit (subject, as always, to the assumptions of the model you fit).
  • \(\phi = 1\): random walk, time-dependent and unbounded variance, and a recipe for spurious results in unrelated, non-cointegrated levels regressions.
  • \(\phi = 0.95\) vs \(\phi = 1.0\): nearly indistinguishable in the small samples we typically work with.

Where we are on the mixing board. Module 2 turns on no new sliders in the master equation:

\[y_t = \alpha + \delta t + \sum_{j=1}^{p} \phi_j y_{t-j} + \sum_{l=1}^{q} \theta_l \epsilon_{t-l} + \epsilon_t.\]

The DGP is still Module 1’s AR(1) — \(\alpha\) and \(\phi_1\) are the only active channels. What changes is the question we ask of it: is \(\phi_1 = 1\)? Along the way, the deterministic-trend slider \(\delta t\) makes its first cameo, because deciding whether a trending series is trend-stationary or difference-stationary turns out to be part of the same testing problem.

Today’s question is concrete: given a time series, which side of the boundary are we on? We will develop several different sources of evidence — visual inspection, the autocorrelation function, formal statistical tests, and a rule of thumb — and combine them into a workflow that produces a defensible verdict. None of these tools is conclusive on its own, especially in the borderline cases. The discipline of stationarity analysis is using them together and being honest about what you don’t know.


2.2 The Difference Operator

Before we test for stationarity, we need the tool that will appear repeatedly in this module and throughout the course: the difference operator.

First Difference

The first difference of \(y_t\) is the change from one period to the next: \[\Delta y_t = y_t - y_{t-1}\]

In the lag operator notation from Module 1, \(\Delta = (1 - L)\). The difference operator is the discrete-time analog of a derivative: it measures the change over one observation interval. When we say a series is “growing at 2% per year,” we are making a statement about its first difference (or, more precisely, its log difference).

Second Difference

The second difference is the difference of the first difference: \[\Delta^2 y_t = \Delta(\Delta y_t)\]

Apply the inner \(\Delta\) first: \[\Delta^2 y_t = \Delta(y_t - y_{t-1}) = (y_t - y_{t-1}) - (y_{t-1} - y_{t-2})\]

Combining terms: \[\Delta^2 y_t = y_t - 2y_{t-1} + y_{t-2}\]

The lag operator gives the same result more cleanly: \[\Delta^2 y_t = (1 - L)^2 y_t = (1 - 2L + L^2) y_t = y_t - 2y_{t-1} + y_{t-2}\]

The analogy to calculus is useful rather than exact: the first difference is like a first derivative (change over one interval), and the second difference is like a second derivative (change in that change). For most economic time series you encounter, you almost never need beyond a second difference. We will return to this in Section 2.10 when we define the order of integration.

Why This Matters Right Now

The difference operator does double duty in this module:

  1. It is the dependent variable in the Dickey-Fuller test regression we will build in Section 2.5.
  2. It is the tool we use to make a non-stationary series stationary, once we have detected non-stationarity (Section 2.10).

The algebra needs to be comfortable before we go further.


2.3 Building Evidence I: Visual Inspection

The first source of evidence is the cheapest: look at the data. Visual inspection cannot give you a definitive answer about stationarity, but it can rule out the easy cases and flag the hard ones.

What Stationarity Looks Like

What should a stationary series look like? Recall the three conditions from Module 1:

  1. Constant mean across time.
  2. Constant variance across time.
  3. Autocovariance that depends only on the lag, not on when you measure it.

Visually, this means: if you take any two non-overlapping windows of the series, the means should be similar, and the spreads (standard deviations) should be similar. The series will have its ups and downs — a stationary series can be highly volatile — but the envelope of its movement should be roughly the same throughout the sample.

library(forecast); library(ggplot2)
set.seed(123)
y_stat <- arima.sim(n = 1000, list(ar = 0.5))
autoplot(y_stat) + theme_bw(base_size = 14) +
  labs(title = "Stationary AR(1), phi = 0.5", y = "")

Compare this to a random walk:

set.seed(123)
y_rw <- cumsum(rnorm(1000))
autoplot(ts(y_rw)) + theme_bw(base_size = 14) +
  labs(title = "Random Walk", y = "")

The stationary series oscillates around a fixed mean (zero) with a roughly consistent envelope. The random walk wanders — there is no level the series returns to, and the spread of values you see in the second half of the sample looks different from the first half. This is the visual fingerprint of a unit root: the series goes places and stays there.

The Trap: Trend-Stationary Series

Visual inspection alone is not enough. Consider a series with a deterministic trend on top of stationary AR(1) deviations: \[y_t = 0.05\, t + u_t, \qquad u_t = 0.5\, u_{t-1} + \epsilon_t\]

set.seed(456)
n <- 300

# Build the DGP literally: stationary AR(1) deviations u_t ...
u <- numeric(n)
u[1] <- 0
for (t in 2:n) {
  u[t] <- 0.5 * u[t-1] + rnorm(1)   # u_t = 0.5 u_{t-1} + epsilon_t
}

# ... riding on a deterministic trend: y_t = 0.05 t + u_t
y_trend <- 0.05 * (1:n) + u
autoplot(ts(y_trend)) + theme_bw(base_size = 14)

This series clearly trends upward. A naive viewer would say “non-stationary” and reach for the difference operator. But this series is trend-stationary (the concept introduced in Module 1, Section 1.4): the deviations from the trend are stationary. The right treatment is to detrend (subtract \(\hat{\alpha} + \hat{\delta} t\)), not to difference. Visual inspection cannot distinguish trend-stationary from difference-stationary — both will look “trended” in a level plot.

This is why we need more tools.

Three Things to Look For

When you plot a series for the first time, ask yourself:

  1. Does the series wander without returning to a level? This suggests possible non-stationarity (a unit root or a structural break).
  2. Is the variance changing over time? Periods of high volatility followed by calm periods suggest variance non-stationarity. For now we will set this aside; we revisit it in Module 14 when we discuss ARCH/GARCH.
  3. Is there an obvious trend? If yes, the question becomes “trend-stationary or difference-stationary?” — and visual inspection alone cannot answer it.

2.4 Building Evidence II: The Autocorrelation Function

The second source of evidence is the ACF, which we introduced in Module 1. Recall that for a stationary AR(1), \(\rho_k = \phi^k\) — the autocorrelation decays geometrically at rate \(\phi\). This decay pattern is what we look for in the correlogram.

Patterns to Recognize

  • Fast geometric decay: the ACF drops quickly toward zero. The bars are inside the confidence bands within a handful of lags. Suggests stationarity with short-to-moderate memory.
  • Slow but visible decay: the bars decline gradually but are still inside the bands by 15-20 lags. Suggests stationarity with high persistence (large \(\phi\), perhaps 0.85-0.95).
  • Persistent / no decay: the bars stay near 1.0 for many lags, often dozens. The decay looks linear rather than exponential, or there is barely any decay at all. Strong evidence of a unit root.

The visual difference between “very slow decay” and “no decay” is exactly the \(\phi = 0.95\) vs \(\phi = 1.0\) problem from Module 1. In small samples, the ACF often cannot tell them apart.

(A note on packages: patchwork sits outside the core course package stack. It does no modeling — it only arranges several ggplots into one figure, here and in the Module 2 application below — so we use it as a stated display-only convenience.)

library(forecast); library(patchwork)
set.seed(123)
phis <- c(0.3, 0.7, 0.95, 1.0)
plots <- lapply(phis, function(p) {
  if (p < 1) {
    y <- arima.sim(n = 500, list(ar = p))
  } else {
    y <- cumsum(rnorm(500))
  }
  ggAcf(y, lag.max = 30) + theme_bw(base_size = 12) +
    ggtitle(paste("phi =", p))
})
wrap_plots(plots, ncol = 2)

The population pattern is the point: for \(\phi = 0.3\) the ACF decays quickly, for \(\phi = 0.7\) it decays more gradually, and for \(\phi = 0.95\) it remains persistent across many lags. The random walk has the most persistent sample ACF. In the rendered finite sample, individual bars can cross a confidence band well after the dominant decay has faded, so do not treat the last significant bar as an exact cutoff. Read the overall shape.

“Build a Case” Framing

We now have two sources of evidence — visual inspection of the series and visual inspection of the ACF. Neither is conclusive. Both can be misled by trend-stationary series, near-unit-root processes, and small samples.

The framing for the rest of the module is this: stationarity is a verdict, not a measurement. No single test gives you the answer. We collect multiple lines of evidence and weigh them together. When the evidence is consistent, we have a clear verdict. When it conflicts, we have to make a judgment call, and we will lean on the cost asymmetry (Section 2.10): spurious regression is more dangerous than over-differencing.


2.5 The Dickey-Fuller and Augmented Dickey-Fuller Tests

The third source of evidence is the formal statistical test. The most widely used test for a unit root is the Dickey-Fuller (DF) test and its augmented version, the ADF.

What We Are Testing

Return to the AR(1) data generating process: \[y_t = \alpha + \phi y_{t-1} + \epsilon_t\]

We want to test:

  • \(H_0: \phi = 1\) (unit root, non-stationary)
  • \(H_1: \phi < 1\) (stationary under the maintained restriction \(\phi > -1\))

This is a one-sided test. In the usual applied setting we maintain \(\phi > -1\), so the left-sided alternative corresponds to \(|\phi| < 1\). The standard ADF workflow is not designed to diagnose a negative unit root or an explosively oscillating process; if the plot suggests either, investigate that behavior separately.

Why Standard OLS Inference Fails

The naive approach is: estimate the AR(1) by OLS, get \(\hat{\phi}\), compute its standard error, form a \(t\)-statistic for \(H_0: \phi = 1\), and compare to standard \(t\)-tables. This fails. Under \(H_0\), the regressor \(y_{t-1}\) is itself a random walk — non-stationary. Standard OLS asymptotics require the regressors to be well-behaved (loosely: they have finite second moments and the regressor moment matrix converges to a positive-definite limit). A random walk’s variance grows linearly with \(t\), so this assumption is violated. As a result:

  • \(\hat{\phi}\) is still consistent for \(\phi\), but the limit distribution is no longer normal.
  • The standard error formula does not give a valid Wald-type statistic.
  • The \(t\)-statistic exists, but the standard critical values do not apply.

This is the same conceptual problem as spurious regression in Module 1. Whenever non-stationary regressors enter an OLS regression, standard inference breaks down. The DF test is essentially a fix for this in the unit-root testing context, and the MacKinnon critical values we will introduce below are the empirical quantiles of the resulting non-standard distribution.

Building the DF Regression

The trick is to rearrange the AR(1) so that the dependent variable is stationary under the null. Subtract \(y_{t-1}\) from both sides: \[y_t - y_{t-1} = \alpha + \phi y_{t-1} - y_{t-1} + \epsilon_t\] \[\Delta y_t = \alpha + (\phi - 1) y_{t-1} + \epsilon_t\]

Define \(\gamma = \phi - 1\). The DF regression is: \[\Delta y_t = \alpha + \gamma y_{t-1} + \epsilon_t\]

The hypothesis test becomes:

  • \(H_0: \gamma = 0\), equivalently \(\phi = 1\) (unit root)
  • \(H_1: \gamma < 0\), equivalently \(\phi < 1\) (stationary under the maintained restriction \(\phi > -1\))

Technical Note — Two different \(\gamma\)’s

The \(\gamma\) in the DF regression is a regression coefficient, \(\gamma = \phi - 1\). It is unrelated to the autocovariances \(\gamma_k\) from Module 1 (which always carry a lag subscript). The unit-root literature’s use of \(\gamma\) here is standard, so we keep it — just read the symbol in context.

Why this rearrangement? Two reasons:

  1. The dependent variable is stationary under \(H_0\). Under the null, \(\Delta y_t = \alpha + \epsilon_t\) in the drift specification (and \(\Delta y_t=\epsilon_t\) in the no-constant specification). The left-hand side is therefore stationary regardless of whether the null or alternative holds, which is at least somewhat better-behaved than having a non-stationary dependent variable. The regression is still subject to the non-standard asymptotics caused by the non-stationary regressor \(y_{t-1}\) on the right, but the left side is not the source of the problem.

  2. The hypothesis maps directly to the parameter. The test statistic of interest is the \(t\)-stat on \(\hat{\gamma}\). If \(\hat{\gamma}\) is significantly less than zero, we reject the unit root. The interpretation is direct.

This \(t\)-statistic has its own conventional symbol: \(\tau\) (tau), the Dickey-Fuller test statistic. It is computed like an ordinary \(t\)-ratio but compared against DF critical values, not \(t\)-tables (the next two subsections show why). The name survives in software: urca::ur.df() labels its critical values tau2 for the drift specification and tau3 for the trend specification. Throughout this module, “the ADF statistic” and \(\tau\) are the same object.

The Dickey-Fuller test was designed in 1979 and the choice of regression form was deliberate — it gave Dickey and Fuller a tractable way to derive the limiting distribution of the test statistic, even though that distribution turned out to be non-standard.

The Augmented Version

The DF test as stated above is the simple version, and it has a problem: it assumes the errors \(\epsilon_t\) are white noise. If the true DGP is an AR(\(p\)) with \(p > 1\), the errors of the simple DF regression will have serial correlation, distorting the test. The fix is to include lagged differences as additional regressors: \[\Delta y_t = \alpha + \gamma y_{t-1} + \sum_{i=1}^{k} \beta_i \Delta y_{t-i} + \epsilon_t\]

Here \(p\) is the order of the underlying AR process and \(k\) is the number of augmentation lags in the ADF regression. An exact rearrangement of an AR(\(p\)) uses \(k=p-1\) lagged differences. In applied work we usually do not know \(p\), so we choose \(k\) to absorb the short-run dynamics and leave approximately white-noise residuals; the selected \(k\) need not equal the unknown \(p-1\). The test of interest is still \(H_0: \gamma = 0\). The augmented version is what people actually use in practice; the “ADF test” is the default. We will discuss how to choose \(k\) later in this section.

The Three ADF Specifications

This is where most applied mistakes happen. There are three versions of the ADF regression depending on what deterministic terms you include. Choosing the wrong specification can flip your conclusion.

Specification 1: No constant, no trend (“none”) \[\Delta y_t = \gamma y_{t-1} + \sum_{i=1}^{k} \beta_i \Delta y_{t-i} + \epsilon_t\] Use when the series has zero mean and no trend. This is rare in practice — almost all economic series have a non-zero mean. Even if they do have a non-zero mean the inclusion of a mean term is a generalization of this (though subject to estimation error).

Specification 2: Constant, no trend (“drift”) \[\Delta y_t = \alpha + \gamma y_{t-1} + \sum_{i=1}^{k} \beta_i \Delta y_{t-i} + \epsilon_t\] Use when the series has a non-zero mean but no obvious trend. This is the default for most stationary economic series — interest rates, unemployment rates, inflation rates, debt-to-GDP ratios.

Specification 3: Constant and trend (“trend”) \[\Delta y_t = \alpha + \delta t + \gamma y_{t-1} + \sum_{i=1}^{k} \beta_i \Delta y_{t-i} + \epsilon_t\] Use when the series has a clear deterministic trend. This specification tests whether the series is stationary around the trend (trend-stationary) or has a unit root with drift.

The critical values differ across specifications because including more deterministic terms changes the asymptotic distribution of the test statistic. Using critical values from the wrong specification is a real and common error.

The general guidance:

  • Look at a plot of the series. Does it have a non-zero mean? A trend?
  • Match the deterministic terms in the test to what you see.
  • If unsure, err toward including more deterministic terms — but be aware that this reduces test power.

If you want a formal, defensible procedure for making this choice — and an explanation of the extra phi1/phi2/phi3 statistics that urca::ur.df() prints in its output — the two Deeper Dive sections below develop both. On a first reading, you can move straight to the Monte Carlo demonstration.

Deeper Dive — A General-to-Specific Deterministic-Term Strategy

The “match the specification to the plot” guidance above is informal and practical. Enders (Ch. 4, pp. 235–238) also discusses a more systematic general-to-specific strategy for the deterministic terms: begin with a trend when one is plausible, then consider simpler specifications only when the excluded term is substantively and statistically defensible.

This is sometimes loosely called the “Pantula principle,” but that name belongs to a different problem. The Dickey-Pantula procedure tests the number of unit roots from the highest plausible order downward — for example, testing \(I(2)\) against \(I(1)\) before testing \(I(1)\) against \(I(0)\). We use that downward logic in Section 2.10. Choosing among no constant, drift, and trend specifications is instead a deterministic-term decision within a given unit-root test.

An applied sequence is:

  1. Start from the plot and the economic setting. If a deterministic trend is plausible, begin with \[\Delta y_t = \alpha + \delta t + \gamma y_{t-1} + \sum_{i=1}^{k} \beta_i \Delta y_{t-i} + \epsilon_t.\]

  2. Read the marginal ADF statistic \(\tau_3\) using its Dickey-Fuller critical values. Rejection is evidence against a unit root conditional on the trend specification. It does not by itself tell you whether the fitted trend is needed; that is a separate modeling decision.

  3. If \(\tau_3\) fails to reject, do not declare a unit root and do not automatically delete the trend. The trend specification has lower power, but deleting a real trend also misspecifies the test. Use the joint statistics below, the individual deterministic-term estimates, the plot, and the economic context to decide whether a simpler drift specification is credible.

  4. If a simpler specification is credible, refit and use that specification’s critical values. Moving from trend to drift, or from drift to no constant, can improve power only when the omitted deterministic term is genuinely unnecessary.

This sequence produces a working specification, not a proof. The attraction is practical: it prevents a trend term from being included mechanically while also preventing a joint rejection from being overinterpreted. For the formal distinction between deterministic-term selection and Dickey-Pantula testing of the integration order, see Enders Ch. 4 and Dickey and Pantula (1987).

Deeper Dive — The Joint Tests: \(\Phi_1\), \(\Phi_2\), \(\Phi_3\)

When you run urca::ur.df(...) and look at the summary output, you will see joint test statistics in addition to the main DF \(t\)-statistic on z.lag.1. With type = "drift", you get phi1; with type = "trend", you get phi2 and phi3. These are joint \(F\)-tests from Dickey & Fuller’s original (1981) paper that test combined hypotheses about the unit root and the deterministic terms together.

The statistics correspond to the drift or trend specification as indicated: \[\Delta y_t = \alpha + \delta t + \gamma y_{t-1} + \sum_{i=1}^{k} \beta_i \Delta y_{t-i} + \epsilon_t\]

Statistic Specification Composite null hypothesis
\(\Phi_1\) Drift \(\alpha = 0, \gamma = 0\)
\(\Phi_2\) Trend \(\alpha = 0, \delta = 0, \gamma = 0\)
\(\Phi_3\) Trend \(\delta = 0, \gamma = 0\)

These are computed as standard \(F\)-statistics from the regression sum of squares, but — like the marginal DF \(t\)-statistic — they do not follow standard \(F\)-distributions under their nulls. The non-stationarity of \(y_{t-1}\) under the unit root null means the joint distribution of \((\hat{\alpha}, \hat{\delta}, \hat{\gamma})\) is non-standard. Dickey and Fuller (1981) tabulated the critical values; modern software (including urca) reports both the test statistic and the critical values automatically.

How to interpret a rejection. A joint rejection says that at least one restriction in the composite null is inconsistent with the data. It does not identify which restriction failed. In particular, if \(\Phi_3\) rejects while the marginal \(\tau_3\) statistic fails to reject \(\gamma=0\), you cannot conclude from that pattern alone that the rejection was “driven by the trend” or that the process is trend-stationary. Inspect the component estimates, use the appropriate critical values, and refit a simpler specification only when the substantive case for dropping the trend is credible.

Practical advice. For most applied work, you do not need to compute \(\Phi_1\), \(\Phi_2\), or \(\Phi_3\) by hand. When you run summary(ur.df(y, type = "trend")), recognize the additional output as composite tests, compare each statistic with its own reported critical values, and resist turning one joint rejection into a unique diagnosis. Enders Ch. 4 provides the more nuanced reading when the deterministic specification is consequential.

References: Dickey, D.A. and Fuller, W.A. (1981), “Likelihood Ratio Statistics for Autoregressive Time Series with a Unit Root,” Econometrica, 49, 1057-1072; Enders, Ch. 4, pp. 207–209.

The Non-Standard Distribution Under H₀: A Monte Carlo Demonstration

We have asserted that the DF \(t\)-statistic does not follow a \(t\)-distribution under the null. Rather than derive this (the derivation involves functional Brownian motion arguments — see Phillips, 1987, or Hamilton Ch. 17 if you are curious), let us show it with a Monte Carlo simulation. This is this module’s central assumption-break: we take the machinery a standard \(t\)-test trusts and watch it fail under the unit-root null.

The strategy: generate many independent random walks (so \(H_0\) is true by construction), run the DF regression on each, and collect the resulting \(t\)-statistics. The empirical distribution of these statistics is the Dickey-Fuller distribution.

set.seed(42)
n_sims <- 5000
T <- 200

df_stats <- replicate(n_sims, {
  # Generate a random walk under H_0: phi = 1
  y <- cumsum(rnorm(T))
  dy <- diff(y)
  y_lag <- y[-T]
  # Run the DF regression (drift / intercept specification)
  reg <- lm(dy ~ y_lag)
  # Extract the t-statistic on y_lag
  summary(reg)$coefficients["y_lag", "t value"]
})

# Summary statistics
mean(df_stats)              # Should be around -1.5, not 0
[1] -1.546135
sd(df_stats)                # Should be around 0.86, not 1
[1] 0.8525777
quantile(df_stats, 0.05)    # Should be around -2.86, not -1.65
       5% 
-2.892739 

Now plot the resulting distribution alongside the standard normal density:

library(ggplot2)
ggplot(data.frame(stat = df_stats), aes(x = stat)) +
  geom_histogram(aes(y = after_stat(density)), bins = 60,
                 fill = "steelblue", alpha = 0.6) +
  stat_function(fun = dnorm, color = "red", linewidth = 1.2) +
  geom_vline(xintercept = -1.645, linetype = "dashed",
             color = "red", linewidth = 0.8) +
  geom_vline(xintercept = quantile(df_stats, 0.05),
             linetype = "dashed", color = "steelblue", linewidth = 0.8) +
  theme_bw(base_size = 14) +
  labs(title = "DF distribution (blue) vs standard normal (red)",
       subtitle = paste0("5% critical values: DF = ",
                         round(quantile(df_stats, 0.05), 2),
                         " | Normal = -1.65"),
       x = "t-statistic under H_0: phi = 1",
       y = "Density")

What you see: The DF distribution is shifted to the left — its mean is about \(-1.5\), not zero. It is not symmetric. The 5% left-tail critical value is about \(-2.86\) for a sample of size \(T = 200\) with the drift specification, not the \(-1.65\) you would use for a standard normal one-sided test.

The implication. If you used the wrong critical value — say, \(-1.65\) from a standard normal — you would reject the unit root null roughly 46% of the time when it is actually true, instead of the nominal 5%. Your test would be massively oversized. The MacKinnon critical values we present below are not arbitrary numbers; they are the empirical quantiles of exactly this distribution, computed once carefully so we do not have to redo the simulation every time.

This is the same problem as spurious regression: standard OLS inference breaks down when regressors are non-stationary. The DF test does not solve that problem; it works around it by using the correct critical values for the resulting non-standard distribution.

MacKinnon Critical Values

MacKinnon’s 2010 response-surface update provides convenient finite-sample critical-value formulas for the DF distribution. For the drift specification (Specification 2), the coefficients reported in its Table 2 give the approximations:

\[ \begin{aligned} \text{CV}_{0.01} &= -3.43035 - \frac{6.5393}{T} - \frac{16.786}{T^2} - \frac{79.433}{T^3} \\ \text{CV}_{0.05} &= -2.86154 - \frac{2.8903}{T} - \frac{4.234}{T^2} - \frac{40.04}{T^3} \\ \text{CV}_{0.10} &= -2.56677 - \frac{1.5384}{T} - \frac{2.809}{T^2} \end{aligned} \]

(The subscript is the significance level; we write \(\text{CV}\) rather than \(\alpha\) to avoid a collision with the intercept \(\alpha\) in the test regression.) Notice that the critical values depend on \(T\). For small samples they are more extreme; as \(T \to \infty\) they converge to their asymptotic values. The corrections shrink rapidly as \(T\) grows — for \(T > 500\) they are negligible.

Reference: MacKinnon, J.G. (2010), “Critical Values for Cointegration Tests,” Queen’s Economics Department Working Paper No. 1227, Table 2. MacKinnon (1996) develops the numerical distribution-function approach, while the displayed coefficients come from the later update. Enders (Ch. 4, pp. 207–208) also tabulates commonly used values.

A simple R function:

mackinnon_cv <- function(T, level = 0.05, spec = "drift") {
  # Returns the MacKinnon critical value for the ADF test
  # spec = "drift" gives the constant-only specification
  if (spec != "drift") {
    stop("Only drift specification implemented in this teaching example.
          Use urca::ur.df() for other specifications.")
  }
  if (level == 0.01) {
    return(-3.43035 - 6.5393/T - 16.786/T^2 - 79.433/T^3)
  }
  if (level == 0.05) {
    return(-2.86154 - 2.8903/T - 4.234/T^2 - 40.04/T^3)
  }
  if (level == 0.10) {
    return(-2.56677 - 1.5384/T - 2.809/T^2)
  }
  stop("Level must be 0.01, 0.05, or 0.10")
}

# Sanity check against the Monte Carlo result from above
mackinnon_cv(T = 200)        # Returns ~-2.876
[1] -2.876102
quantile(df_stats, 0.05)     # The empirical 5% quantile from the MC
       5% 
-2.892739 

The two values should be very close — MacKinnon’s formula is a smoothed and refined version of exactly the kind of Monte Carlo we just ran, with vastly more replications and across many sample sizes.

For the trend specification (Spec 3), the critical values are more negative still (around \(-3.41\) at the 5% level for large \(T\)). For the no-constant specification (Spec 1), they are less negative (around \(-1.95\) at the 5% level). The full set of formulas is in MacKinnon’s paper. In practice, the urca::ur.df() function returns the critical values automatically for whichever specification you chose.

(This function also lives in helpers/testing.R as the canonical course copy, so later modules can source() it instead of re-defining it.)

Lag Selection in the ADF

How many lagged differences (\(k\)) should you include in the ADF regression?

  • Too few: residuals will still have serial correlation, biasing the test.
  • Too many: you lose degrees of freedom and reduce power.

Common approaches:

  1. Information criteria: minimize AIC or BIC over a range of candidate \(k\) values. BIC tends to choose smaller models and is the more parsimonious choice. In urca::ur.df(), set selectlags = "BIC" or "AIC". One software detail matters: when you supply a positive maximum in lags, ur.df() searches positive augmentation orders from 1 through that maximum; it does not include \(k=0\) in that internal search. If zero augmentation lags is a plausible candidate, fit it separately and compare it manually using the same aligned estimation sample and information criterion. We will talk more about information criteria later.

  2. Sequential testing (top-down): start with a maximum augmentation order \(k_{\max}\) and drop the highest insignificant lag. Schwert (1989) suggested: \[k_{\max} = \left\lfloor 12 \cdot \left(\frac{T}{100}\right)^{1/4} \right\rfloor\] For \(T = 100\), this gives \(k_{\max} = 12\). For \(T = 200\), \(k_{\max} \approx 14\). Then test the lagged-difference coefficients sequentially and drop the top one if insignificant, refitting until the highest remaining lag is significant.

  3. Course default: For most applied work, BIC selection up to a modest \(k_{\max}\) is fine. We will use BIC as the default in the problem sets, checking the \(k=0\) candidate separately when it is plausible.

Building the ADF in R: From Scratch and From the Package

Live coding the test “from scratch” demystifies what urca::ur.df() is doing under the hood. Recall the embed() function from Module 1.

library(urca); library(forecast)

set.seed(8675309)
y <- arima.sim(n = 200, list(ar = 0.7))   # Truly stationary
T <- length(y)
k_lags <- 4

# Construct the regression matrix manually
dy <- diff(y)
n  <- length(dy)
X  <- embed(dy, k_lags + 1)               # [dy_t, dy_{t-1}, ..., dy_{t-k}]
dy_t    <- X[, 1]
dy_lags <- X[, -1, drop = FALSE]
y_lag   <- y[(k_lags + 1):(T - 1)]        # y_{t-1} aligned with the truncated dy

# Run the regression
adf_reg <- lm(dy_t ~ y_lag + dy_lags)
summary(adf_reg)

Call:
lm(formula = dy_t ~ y_lag + dy_lags)

Residuals:
     Min       1Q   Median       3Q      Max 
-2.76518 -0.58199  0.06084  0.58090  2.93186 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)    
(Intercept) -0.002613   0.069252  -0.038    0.970    
y_lag       -0.324343   0.065915  -4.921 1.87e-06 ***
dy_lags1     0.066426   0.080404   0.826    0.410    
dy_lags2     0.010452   0.078874   0.133    0.895    
dy_lags3     0.086795   0.074705   1.162    0.247    
dy_lags4     0.080070   0.072493   1.105    0.271    
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.9666 on 189 degrees of freedom
Multiple R-squared:  0.1501,    Adjusted R-squared:  0.1276 
F-statistic: 6.677 on 5 and 189 DF,  p-value: 9.432e-06
# The test statistic is the t-stat on y_lag
adf_stat <- summary(adf_reg)$coefficients["y_lag", "t value"]
critical <- mackinnon_cv(T = T, level = 0.05)

cat("ADF statistic:", round(adf_stat, 3), "\n")
ADF statistic: -4.921 
cat("5% critical value:", round(critical, 3), "\n")
5% critical value: -2.876 
cat("Reject H_0:", adf_stat < critical, "\n")
Reject H_0: TRUE 

For this stationary AR(1) with \(T = 200\), you get an ADF statistic around \(-5\) (the exact value depends on the simulated draw), well below the critical value of \(-2.88\). We comfortably reject the unit root null in favor of stationarity.

The packaged version:

adf_test <- ur.df(y, type = "drift", lags = 4)
summary(adf_test)

############################################### 
# Augmented Dickey-Fuller Test Unit Root Test # 
############################################### 

Test regression drift 


Call:
lm(formula = z.diff ~ z.lag.1 + 1 + z.diff.lag)

Residuals:
     Min       1Q   Median       3Q      Max 
-2.76518 -0.58199  0.06084  0.58090  2.93186 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)    
(Intercept) -0.002613   0.069252  -0.038    0.970    
z.lag.1     -0.324343   0.065915  -4.921 1.87e-06 ***
z.diff.lag1  0.066426   0.080404   0.826    0.410    
z.diff.lag2  0.010452   0.078874   0.133    0.895    
z.diff.lag3  0.086795   0.074705   1.162    0.247    
z.diff.lag4  0.080070   0.072493   1.105    0.271    
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.9666 on 189 degrees of freedom
Multiple R-squared:  0.1501,    Adjusted R-squared:  0.1276 
F-statistic: 6.677 on 5 and 189 DF,  p-value: 9.432e-06


Value of test-statistic is: -4.9206 12.1109 

Critical values for test statistics: 
      1pct  5pct 10pct
tau2 -3.46 -2.88 -2.57
phi1  6.52  4.63  3.81

The output shows the same regression and test statistic, together with the 1%, 5%, and 10% critical values for the drift specification.

For a real workflow, you would use the packaged version with selectlags = "BIC":

adf_test <- ur.df(y, type = "drift", lags = 12, selectlags = "BIC")
summary(adf_test)

############################################### 
# Augmented Dickey-Fuller Test Unit Root Test # 
############################################### 

Test regression drift 


Call:
lm(formula = z.diff ~ z.lag.1 + 1 + z.diff.lag)

Residuals:
     Min       1Q   Median       3Q      Max 
-2.57956 -0.64418  0.03028  0.67596  2.88927 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept) -0.00902    0.07172  -0.126    0.900    
z.lag.1     -0.29040    0.05569  -5.215 4.93e-07 ***
z.diff.lag   0.02940    0.07417   0.396    0.692    
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.9802 on 184 degrees of freedom
Multiple R-squared:  0.1411,    Adjusted R-squared:  0.1317 
F-statistic: 15.11 on 2 and 184 DF,  p-value: 8.396e-07


Value of test-statistic is: -5.2145 13.5965 

Critical values for test statistics: 
      1pct  5pct 10pct
tau2 -3.46 -2.88 -2.57
phi1  6.52  4.63  3.81

This searches the positive augmentation orders \(k=1,\ldots,12\) using BIC, runs the test at the selected order, and reports the results. It does not compare \(k=0\) internally; run ur.df(y, type = "drift", lags = 0) separately if that candidate belongs in your comparison set.


2.6 KPSS: The Reverse-Null Test

The Dickey-Fuller test has a particular structural feature: its null hypothesis is a unit root. In statistical testing, we can only reject the null, not accept it. So if ADF fails to reject the unit root, the correct interpretation is “we lack evidence against the unit root,” not “the series has a unit root.” Combined with the low power problem (Section 2.8), this creates an asymmetry — the test is biased toward concluding non-stationarity when the data are uninformative.

The KPSS test (Kwiatkowski, Phillips, Schmidt, and Shin, 1992) is the conceptual mirror: its null hypothesis is stationarity, and the alternative is a unit root.

Test \(H_0\) \(H_1\)
ADF unit root stationary
KPSS stationary unit root

Why Use Both

Because both tests have low power against close alternatives, neither one alone is fully informative. Used together they enable confirmatory analysis:

ADF says KPSS says Conclusion
Reject unit root Fail to reject stationarity Strong evidence of stationarity
Fail to reject unit root Reject stationarity Strong evidence of a unit root
Reject unit root Reject stationarity Contradiction — possible structural break or misspecification
Fail to reject unit root Fail to reject stationarity Inconclusive — borderline case, defer to other evidence

The diagonal cases (when both tests agree) give you high confidence. The off-diagonal cases tell you something is unusual and you need to think harder. The “inconclusive” case (both fail to reject) is itself useful information: it tells you honestly that the data are too noisy or the sample is too small to discriminate between stationarity and a unit root. In that case, you fall back on the cost asymmetry from Section 2.10.

Using KPSS: What Matters

You can use KPSS correctly knowing just three things:

  1. The null is stationarity, opposite of ADF.
  2. A large test statistic is evidence against stationarity (the opposite of how ADF works — this is a common point of confusion).
  3. Critical values come from the KPSS asymptotic distribution, tabulated in their paper and built into R packages.

The mechanics behind the statistic are in the Deeper Dive below; they are not essential for using the test.

Deeper Dive — How KPSS Works

The KPSS test decomposes the series into a deterministic component, a random walk component, and a stationary error: \[y_t = \xi_t + r_t + u_t, \quad r_t = r_{t-1} + v_t, \quad v_t \sim N(0, \sigma_v^2)\]

Under \(H_0\) (stationarity), \(\sigma_v^2 = 0\) — the random walk component does not exist. The KPSS statistic measures how much the partial sums of the residuals from regressing \(y_t\) on the deterministic component “wander.” Partial sums of stationary residuals do not literally stay bounded; they typically grow at the stationary rate \(O_p(\sqrt{T})\). After the KPSS scaling, that is the benchmark order under the null. A stochastic-trend component makes the partial sums grow much faster, producing a large statistic and rejection.

KPSS in R

library(urca)
# Level stationarity (no deterministic trend)
kpss_level <- ur.kpss(y, type = "mu")
summary(kpss_level)

####################### 
# KPSS Unit Root Test # 
####################### 

Test is of type: mu with 4 lags. 

Value of test-statistic is: 0.08 

Critical value for a significance level of: 
                10pct  5pct 2.5pct  1pct
critical values 0.347 0.463  0.574 0.739
# Trend stationarity (allow for a deterministic trend)
kpss_trend <- ur.kpss(y, type = "tau")
summary(kpss_trend)

####################### 
# KPSS Unit Root Test # 
####################### 

Test is of type: tau with 4 lags. 

Value of test-statistic is: 0.0795 

Critical value for a significance level of: 
                10pct  5pct 2.5pct  1pct
critical values 0.119 0.146  0.176 0.216

The output shows the test statistic and the critical values at 10%, 5%, 2.5%, and 1% levels. If the statistic exceeds the critical value, you reject stationarity. Note again: this is the opposite direction of comparison from ADF, where you reject the null when the statistic is below the critical value. Get this confused at your peril.

Choosing Between mu and tau

  • Use type = "mu" when testing whether the series is stationary around a constant level (no trend in the deterministic component). This pairs naturally with the drift specification of ADF.
  • Use type = "tau" when testing whether the series is stationary around a deterministic trend. This pairs naturally with the trend specification of ADF.

For the FRED unemployment series in Section 2.12, we will use type = "mu" because there is no obvious deterministic trend in unemployment.


2.7 Other Tests: Phillips-Perron and ERS

There is a small zoo of unit root tests. Two more deserve a brief mention — you need to recognize the names, not master the mechanics.

Phillips-Perron (PP): Like ADF, but uses a non-parametric correction for serial correlation in the errors instead of including lagged differences. Same null, similar critical values, generally similar conclusions. Available as urca::ur.pp(). Sometimes more powerful than ADF when the error structure is too complex for a small number of lags to capture, but in most cases the two give similar answers.

Elliott-Rothenberg-Stock (ERS / DF-GLS): A more powerful variant of the ADF test that uses generalized least squares (GLS) detrending instead of OLS detrending. ERS has higher power than ADF, particularly when the series has a non-zero mean or trend. Available as urca::ur.ers(). If you only want to remember one alternative to ADF as a “more powerful version,” this is the one.

There are other tests in the literature — Ng-Perron, Bierens, Im-Pesaran-Shin (for panels) — and active research continues. The point is not that you need to know them all. The point is that these are different tools for the same job, none of them solves the fundamental low-power problem, and they can be used as complements to ADF and KPSS when you want additional evidence.


2.8 Power, the Borderline Case, and Why We Use Multiple Tests

The single most important practical fact about unit root tests is that they have low power against highly persistent stationary alternatives.

To see what this means concretely, consider a stationary AR(1) with \(\phi = 0.95\) — clearly stationary, but very close to the boundary. How often does the ADF test correctly reject the unit root null?

set.seed(2024)
n_sims <- 1000

# Power at phi = 0.95 for various sample sizes
power_results <- sapply(c(50, 100, 200, 500), function(T) {
  rejections <- replicate(n_sims, {
    y <- arima.sim(n = T, list(ar = 0.95))
    test <- ur.df(y, type = "drift", lags = 4)
    test_stat <- test@teststat[1]
    crit <- mackinnon_cv(T = T, level = 0.05)
    test_stat < crit  # TRUE = reject H_0 = correct decision
  })
  mean(rejections)
})

names(power_results) <- c("T=50", "T=100", "T=200", "T=500")
power_results
 T=50 T=100 T=200 T=500 
0.059 0.085 0.267 0.922 

Typical results (from the run above; your exact numbers will wobble with the seed):

  • \(T = 50\): power ≈ 0.06 (barely above the nominal 5% level)
  • \(T = 100\): power ≈ 0.09
  • \(T = 200\): power ≈ 0.27
  • \(T = 500\): power ≈ 0.92

Read these numbers carefully. With 100 observations and \(\phi = 0.95\), the ADF test correctly rejects the unit root less than 10% of the time. The truly stationary series is misclassified as a unit root roughly 90% of the time. This is not a defect of the ADF test specifically; KPSS, PP, ERS all suffer from analogous power problems near the boundary.

This is why we cannot rely on a single test. The borderline cases are exactly the cases where stationarity testing is most useful and most fragile. The cure is the “build a case” framework: collect multiple lines of evidence and make a defensible judgment. It is also why the cost asymmetry in Section 2.10 matters: when the evidence is genuinely uninformative, we fall back on the cost of being wrong in each direction.


2.9 The Rule of Thumb

The rule of thumb is the fourth and last source of evidence — fast, informal, and surprisingly informative.

The Rule

Taught to me by Jeff Mills:

  1. Compute \(\sigma_y\), the standard deviation of the series in levels.
  2. Compute \(\sigma_{\Delta y}\), the standard deviation of the first differences.
  3. Compute the ratio: \[R = \frac{\sigma_{\Delta y}}{\sigma_y}\]
  4. If \(R < 0.5\): differencing reduces the variance sharply. Flag the series as highly persistent; under this course’s cost-based working rule, difference it unless the other evidence points clearly toward a deterministic trend.
  5. If \(R \geq 0.5\): differencing does not reduce variance much. Under the working rule, keep the levels unless the other evidence points clearly toward a unit root.

Why It Works

The rule captures a simple intuition. For a stationary AR(1) with \(|\phi| < 1\), the variance of the levels is \(\sigma_\epsilon^2 / (1 - \phi^2)\), while the variance of the first differences is \(\text{Var}(\Delta y_t) = 2(\gamma_0 - \gamma_1) = 2\gamma_0(1 - \phi) = 2\sigma_\epsilon^2 / (1 + \phi)\). Take the ratio: \[\frac{\text{Var}(\Delta y_t)}{\text{Var}(y_t)} = \frac{2\sigma_\epsilon^2 / (1 + \phi)}{\sigma_\epsilon^2 / (1 - \phi^2)} = 2(1 - \phi)\]

Take square roots to get the ratio of standard deviations and evaluate at a few values:

  • \(\phi = 0\) (white noise): \(R = \sqrt{2} \approx 1.41\)
  • \(\phi = 0.5\): \(R = \sqrt{1.0} = 1.00\)
  • \(\phi = 0.7\): \(R = \sqrt{0.6} \approx 0.77\)
  • \(\phi = 0.9\): \(R = \sqrt{0.2} \approx 0.45\)
  • \(\phi = 0.95\): \(R = \sqrt{0.1} \approx 0.32\)
  • \(\phi = 1.0\) (random walk): \(R \to 0\)

Two observations:

  1. The cutoff \(R < 0.5\) corresponds roughly to \(\phi > 0.875\). Series with very high persistence get flagged for differencing.
  2. The rule cannot tell the difference between \(\phi = 0.95\) and \(\phi = 1.0\) — both produce small \(R\). A value below 0.5 is therefore a high-persistence flag, not evidence of a unit root. The instruction to difference is a conservative modeling action motivated by the course’s cost asymmetry.

When to Use It

The rule of thumb is:

  • Fast: one line of R code.
  • Easy to interpret: no critical values, no specifications, no choices.
  • Not a formal test: you should not put it in a paper as your main evidence. Use it as a sanity check, especially when ADF and KPSS disagree.
ratio <- sd(diff(y)) / sd(y)
ratio
[1] 0.7599642

Below 0.5? The course’s conservative working action is to difference. Above 0.5? The working action is to keep the levels. Treat this as one piece of evidence among several, not as a unit-root test.


2.10 Differencing, Order of Integration, and the Over-Differencing Trap

We have spent most of this module detecting non-stationarity. Now we address the resolution.

Differencing as the Resolution

If a series is non-stationary because of a unit root, the most common remedy is differencing. Recall that for a random walk \(y_t = y_{t-1} + \epsilon_t\): \[\Delta y_t = y_t - y_{t-1} = \epsilon_t\]

The first difference of a random walk is white noise — manifestly stationary. More generally, if a series has exactly one unit root, taking the first difference produces a stationary series. If it has two unit roots (rare), you need to difference twice. And so on.

Order of Integration

The order of integration of a series is the number of times you need to difference it to make it stationary. Notation: \(y_t \sim I(d)\) if the series requires \(d\) differences to become stationary.

  • \(I(0)\): Already stationary. No differencing needed.
  • \(I(1)\): Stationary after one difference. The most common case for non-stationary economic series — GDP, prices, asset prices, exchange rates in levels are all typically \(I(1)\).
  • \(I(2)\): Stationary after two differences. Rare but occasionally encountered. The aggregate price level can be \(I(2)\) when both the level and the inflation rate are non-stationary — that is, when inflation itself has a unit root and needs to be differenced.
  • \(I(d)\) for \(d > 2\): Almost never seen in practice. If your testing suggests you need \(d > 2\), something is probably wrong with your specification — perhaps you have a structural break being mistaken for high-order integration.

A useful analogy: how many licks does it take to get to the center of a Tootsie Pop? The order of integration is how many differences it takes to reach stationarity. We almost never need more than two licks.

The Practical Workflow

  1. Test the levels with ADF (and KPSS if you want confirmatory analysis).
  2. If you fail to reject the unit root, take a first difference and test again.
  3. In most applied economics work, we stop there unless we have reason to think the series could be \(I(2)\).
  4. If \(I(2)\) is plausible, do not just keep differencing upward. Start from the highest order you think is possible and test downward, which is the logic behind Dickey-Pantula style procedures.
# Working helper for the common I(0)/I(1) case when no deterministic
# trend is present. It is not a general integration-order detector.
find_order_i01 <- function(y, level = 0.05) {
  cv_col <- if (level <= 0.01) {
    1
  } else if (level <= 0.05) {
    2
  } else {
    3
  }

  test_level <- ur.df(y, type = "drift", lags = 4)
  if (test_level@teststat[1] < test_level@cval[1, cv_col]) {
    return(0)
  }

  test_diff <- ur.df(diff(y), type = "drift", lags = 4)
  if (test_diff@teststat[1] < test_diff@cval[1, cv_col]) {
    return(1)
  }

  return(NA)  # Inconclusive; if I(2) is plausible, use downward Dickey-Pantula-style testing
}

For a series with no visible deterministic trend, that simple two-step \(I(0)\)/\(I(1)\) check is often a useful first pass. It always uses a drift-only ADF, so it can incorrectly label a trend-stationary series as \(I(1)\). Do not use it as a general integration-order classifier: choose the deterministic specification first, then use the helper only inside its stated scope. The find_order() convenience function in helpers/testing.R extends the same upward screen to a chosen maximum order, but it carries the same deterministic-specification limitation and is not a replacement for downward Dickey-Pantula testing when \(I(2)\) is plausible.

The Over-Differencing Trap

What happens if you difference a series that was actually trend-stationary, not difference-stationary? Suppose the true DGP is: \[y_t = \alpha + \delta t + u_t, \quad u_t \sim N(0, \sigma^2) \text{ i.i.d.}\]

This is a deterministic trend with white noise around it — the simplest possible trend-stationary process. The appropriate treatment under this maintained DGP is to detrend by regressing \(y_t\) on a constant and \(t\). Detrending removes the fitted deterministic component and, in this displayed DGP, recovers the stationary deviation up to estimation error. The resulting residuals are not stationary merely by construction; diagnose them as you would any other proposed stationary series.

Suppose instead we difference. Compute \(\Delta y_t\): \[\Delta y_t = y_t - y_{t-1}\] \[\Delta y_t = (\alpha + \delta t + u_t) - (\alpha + \delta (t-1) + u_{t-1})\] \[\Delta y_t = \delta + (u_t - u_{t-1})\]

The constant \(\alpha\) cancels (good — that’s not the problem). The linear trend \(\delta t - \delta(t-1) = \delta\) collapses to a constant (also good — the differenced series has the constant drift \(\delta\)). But the noise component is now \(u_t - u_{t-1}\), and this is not white noise.

To see what it is, write it in the master-equation form: \[\Delta y_t = \delta + u_t - u_{t-1}\]

If we relabel \(u_t\) as the new innovation \(\epsilon_t\), this is exactly an MA(1) process: \[\Delta y_t = \delta + \epsilon_t + \theta \epsilon_{t-1}, \quad \theta = -1\]

In lag operator form, the MA polynomial is: \[\Theta(L) = 1 + \theta L = 1 - L\]

The root of \(\Theta(L) = 0\) is \(L = 1\) — exactly on the unit circle. The MA polynomial has a unit root.

Why this matters. A moving average process is invertible if all roots of its MA polynomial lie outside the unit circle. Invertibility gives the usual unique, stable one-sided recovery of innovations from past observations. At a unit-circle MA root, that representation is on the boundary. In applied fitting this can show up as a boundary estimate such as \(\hat{\theta}\approx-1\), unstable innovation recovery, an ill-conditioned likelihood, convergence warnings, or nonstandard inference. These symptoms are possible rather than guaranteed: a state-space program can still produce forecasts, but the usual invertible ARMA interpretation and standard large-sample approximations need care.

In this white-noise special case, differencing a trend-stationary series replaces a deterministic trend with a non-invertible MA(1). The differenced series is stationary, but the usual invertible innovation representation sits on the boundary.

The broader result is just as important. If the stationary deviation \(u_t\) has ARMA dynamics rather than being white noise, differencing multiplies those dynamics by the extra MA factor \((1-L)\); it does not erase the original dynamics. For example, if \[u_t = \phi u_{t-1} + \epsilon_t,\] then \[(1-\phi L)\Delta u_t = (1-L)\epsilon_t.\] Thus \(\Delta u_t\) is an ARMA(1,1) with the original AR coefficient \(\phi\) and a boundary MA factor \(\theta=-1\), not a pure MA(1). In general, over-differencing adds a non-invertible \((1-L)\) factor while retaining the stationary ARMA structure already present.

For now, the lesson is: trend-stationary and difference-stationary series require different treatments. Use the trend-augmented ADF specification as evidence when you suspect a deterministic trend, and treat the final choice as a working modeling decision to be checked in diagnostics.

Looking Ahead — Invertibility and diagnostics

You have just met the vocabulary “unit root in the MA polynomial” ahead of schedule. Module 3 develops invertibility formally as the MA counterpart of AR stationarity — and demonstrates experimentally what a nearly non-invertible MA(1) does to a sample ACF and to arima() estimates. Module 5 (diagnostics) then returns to the practical symptom: a fitted ARMA model whose MA coefficient sits suspiciously close to \(-1\) is often telling you the series was over-differenced.

Difference vs. Detrend: Choosing Correctly

The correct treatment depends on what the series actually is:

  • Difference-stationary (unit root, possibly with drift): difference. Removes the stochastic trend without introducing artifacts.
  • Trend-stationary (deterministic trend with stationary deviations): detrend. Removes the deterministic trend without introducing a unit root in the MA polynomial.

How do you tell which? Use the trend-augmented ADF specification (Specification 3 from Section 2.5). It tests for a unit root while allowing a deterministic trend:

  • If you reject the unit root with the trend specification, that is evidence consistent with trend stationarity. The working action is to detrend, then check that the detrended residuals behave as a stationary series.
  • If you fail to reject the unit root even with the trend included, a unit root remains plausible; it has not been proved. Under the course’s cost-based rule, the working action is to difference, then diagnose the resulting model for over-differencing.

The test result is evidence feeding a modeling decision, not a proof of the DGP. That distinction lets us keep a firm applied workflow without pretending the borderline case has disappeared.

The Cost Asymmetry, Revisited

We told you earlier that when in doubt, err toward differencing. That is still good advice — but now you know there is a real cost to over-differencing, too. The right framing is:

  • Spurious regression (the cost of failing to difference a unit-root series) gives you confidently wrong answers with significant t-stats and high \(R^2\). You publish results that do not replicate.
  • Over-differencing (the cost of differencing a trend-stationary series) introduces a non-invertible MA factor. Its effect on coefficient estimates, standard errors, and the estimand depends on the model you fit; there is no general promise that only the standard errors suffer. It is usually easier to detect in ARMA diagnostics than a convincing spurious levels regression, but it is not harmless.

The asymmetry remains: spurious regression is the bigger inference risk. But you are now aware of the over-differencing tax, and you can use the trend-augmented ADF to make the choice deliberately rather than defaulting to differencing in all cases.


2.11 Synthesis: Building a Case

We now have four sources of evidence:

  1. Visual inspection of the series: Does it wander? Is the variance changing? Is there a trend?
  2. Visual inspection of the ACF: Does it decay quickly (stationary) or slowly/not at all (unit root)?
  3. Statistical tests: ADF (and PP, ERS) test $H_0 = $ unit root. KPSS tests $H_0 = $ stationary. Use them together when possible.
  4. Rule of thumb: \(\sigma_{\Delta y} / \sigma_y\). Quick informal check.

Combine them. When all four point the same way, the answer is clear. When they conflict, you have to make a judgment call.

The Decision Framework

  1. Plot the series. What is the gross structure? Trend? Drift? Volatility changes?
  2. Plot the ACF. Fast decay or slow decay?
  3. Choose the appropriate ADF specification based on what you saw in step 1. Run it.
  4. Run KPSS with the matching specification (mu if you ran ADF with drift, tau if you ran ADF with trend).
  5. Compute the rule-of-thumb ratio as a sanity check.
  6. Synthesize. Do the four sources agree? If yes, you have a strong working verdict. If no, apply the cost asymmetry: the course’s conservative working action is to difference unless you have strong reason not to.
  7. Remember the over-differencing trap. If you suspect a deterministic trend, use the trend-augmented ADF and matching KPSS as evidence, then choose between detrending and differencing and diagnose the result. The tests guide the decision; they do not prove which DGP generated the series.

Why Multiple Pieces of Evidence?

The recurring theme: every test we have, every visual diagnostic, every rule of thumb has limited power against the cases that matter most — the borderline ones where \(\phi\) is close to 1 but not equal. No single tool can give you certainty in those cases. The discipline is to use multiple tools and accept that sometimes the honest answer is “I don’t know, but I’m going to difference anyway because the cost asymmetry favors that choice.”

This is not a weakness of time series analysis. It is the truth: extracting causal claims from passive observation of one realization of a stochastic process is hard, and pretending otherwise is how you get spurious regressions.

Practicing the Workflow

The decision framework above only becomes useful once it comes naturally — and that takes repetition on series where you don’t already know the answer. Two ways to get that practice before the problem set:

  1. Generate your own mystery series. Use the simulators you built in Module 1: draw \(\phi\) somewhere in \(\{0.5, 0.8, 0.9, 0.95, 1.0\}\) without looking, simulate, and run the full workflow — plot, ACF, the right ADF specification, matching KPSS, rule of thumb, verdict. Then check against the \(\phi\) you drew. The borderline draws are the ones that teach you the most.
  2. Work the Core Practice problems at the end of this module. They are built around exactly this workflow, and the power study in particular previews the kind of reasoning the problem sets reward.

2.12 Core Application: US Unemployment Rate (continued from Module 1)

This application completes the Module 2 sequence: Concept \(\rightarrow\) Math \(\rightarrow\) Simulate \(\rightarrow\) Real Data.

In Module 1 we pulled the US civilian unemployment rate from FRED and plotted it. We noted that the ACF decays slowly and that the series has high persistence. Now let us run the full stationarity workflow.

To pull the series live, store your FRED key in the environment (never hardcode it in course files) and run:

library(fredr); library(forecast); library(urca); library(patchwork)
fredr_set_key(Sys.getenv("FRED_KEY"))

# Pull the available series, then retain the initial complete monthly run.
# This matches the cached workflow below if FRED contains a release gap.
unrate <- fredr(series_id = "UNRATE",
                observation_start = as.Date("1948-01-01"))
unrate <- unrate[order(unrate$date), ]
first_na <- which(is.na(unrate$value))
if (length(first_na) > 0) unrate <- unrate[seq_len(min(first_na) - 1), ]

y <- ts(unrate$value, start = c(1948, 1), frequency = 12)
T <- length(y)
sample_start <- min(unrate$date)
sample_end <- max(unrate$date)

So that these notes render identically for everyone — with or without an API key or an internet connection — the code below loads the same series from the course’s cached copy in data/UNRATE.csv. The workflow from here on is identical either way.

library(forecast); library(urca); library(patchwork)

unrate <- read.csv("../data/UNRATE.csv")  # cached FRED pull; see data/README.md
unrate$observation_date <- as.Date(unrate$observation_date)
unrate <- unrate[order(unrate$observation_date), ]

# The cache has one missing month (2025-10, a federal data-release gap).
# Keep the complete run from 1948-01 through the last month before the gap
# so the monthly ts index stays aligned.
first_na <- which(is.na(unrate$UNRATE))
if (length(first_na) > 0) unrate <- unrate[seq_len(min(first_na) - 1), ]

y <- ts(unrate$UNRATE, start = c(1948, 1), frequency = 12)
T <- length(y)
sample_start <- min(unrate$observation_date)
sample_end <- max(unrate$observation_date)
T
[1] 933

The rendered snapshot contains 933 consecutive monthly observations from January 1948 through September 2025. Truncating both the live and cached paths at the first missing observation keeps the monthly index aligned and ensures that the downstream ADF and KPSS calls receive no NA values.

Step 1: Visual Inspection

p1 <- autoplot(y) + theme_bw(base_size = 12) +
  labs(title = paste0("US Unemployment Rate, ",
                      format(sample_start, "%b %Y"), "–",
                      format(sample_end, "%b %Y")),
       y = "Percent")
p2 <- ggAcf(y, lag.max = 60) + theme_bw(base_size = 12) +
  labs(title = "ACF of unemployment rate")
p1 / p2

What you see:

  • Levels: The series oscillates roughly between 3% and 11% (with a spike near 15% at the 2020 pandemic shock), with no obvious deterministic trend. There are clear cyclical movements driven by recessions. No obvious variance changes (perhaps slightly more volatile in the post-1970 period, but not dramatically).
  • ACF: Slow decay. The autocorrelation at lag 12 is still around 0.7. At lag 36 it is still positive and likely significant. This is the visual fingerprint of high persistence — possibly a unit root, or possibly \(\phi\) very close to 1.

Verdict from visual inspection: ambiguous. Looks persistent. Could go either way.

Step 2: ADF

The series has a non-zero mean but no obvious trend, so we use the drift specification.

adf_result <- ur.df(y, type = "drift", lags = 12, selectlags = "BIC")
summary(adf_result)

############################################### 
# Augmented Dickey-Fuller Test Unit Root Test # 
############################################### 

Test regression drift 


Call:
lm(formula = z.diff ~ z.lag.1 + 1 + z.diff.lag)

Residuals:
    Min      1Q  Median      3Q     Max 
-1.8826 -0.1323 -0.0195  0.1015 10.3134 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)    
(Intercept)  0.179413   0.047927   3.743 0.000193 ***
z.lag.1     -0.031484   0.008067  -3.903 0.000102 ***
z.diff.lag   0.050759   0.032968   1.540 0.123995    
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.4137 on 917 degrees of freedom
Multiple R-squared:  0.01755,   Adjusted R-squared:  0.01541 
F-statistic: 8.191 on 2 and 917 DF,  p-value: 0.0002979


Value of test-statistic is: -3.9027 7.6157 

Critical values for test statistics: 
      1pct  5pct 10pct
tau2 -3.43 -2.86 -2.57
phi1  6.43  4.59  3.78
mackinnon_cv(T = T, level = 0.05)
[1] -2.864643

The statistic printed above is \(\tau = -3.9027\), below the 5% critical value of approximately \(-2.86\): this is a strong rejection of the unit-root null conditional on the chosen drift specification, sample window, and lag rule. Robustness to alternative samples, structural breaks, or deterministic and lag specifications is a separate question, and — as the next step shows — the other formal test is about to disagree. This workflow has not isolated which episode drives that sensitivity, so we do not assign the rejection to the 2020 spike. One test statistic is one piece of evidence, not a verdict.

Step 3: KPSS

kpss_result <- ur.kpss(y, type = "mu")
summary(kpss_result)

####################### 
# KPSS Unit Root Test # 
####################### 

Test is of type: mu with 6 lags. 

Value of test-statistic is: 0.9703 

Critical value for a significance level of: 
                10pct  5pct 2.5pct  1pct
critical values 0.347 0.463  0.574 0.739

The KPSS statistic for the unemployment rate is typically large enough to reject stationarity at standard levels.

Step 4: Rule of Thumb

sd(diff(y)) / sd(y)
[1] 0.2431225

The ratio is about \(0.24\) in this sample (the 2020 spike inflates the variance of the differences; pre-2020 samples give values closer to \(0.13\)) — either way, well below the 0.5 threshold. The rule of thumb says: difference.

Step 5: Synthesis

Lining up the evidence:

  • Visual / ACF: Suggests high persistence. Inconclusive between \(\phi = 0.95\) and \(\phi = 1.0\).
  • ADF: Rejects the unit root — but the rejection is sample- and specification-sensitive.
  • KPSS: Rejects stationarity.
  • Rule of thumb: Well below 0.5 — difference.

Notice where this lands us in the confirmatory table from Section 2.6: ADF and KPSS both reject — the contradiction row, which can flag a structural break, an outlier, or another form of misspecification. The workflow establishes conflict; it does not identify the cause. The visual evidence says “very persistent,” and the rule of thumb says “difference.” When the evidence conflicts, we fall back on the cost asymmetry. The working decision for our purposes is to treat the unemployment rate as \(I(1)\) and use \(\Delta y_t\) when we get to modeling. That is a transparent applied choice, not a claim that the tests have proved the true DGP contains a unit root.

Deeper Dive — A Note on the Macro Debate

This is real data being honest about its borderline nature. There is genuine and unresolved disagreement among macroeconomists about whether the unemployment rate is stationary. The arguments:

  • Stationary camp: The unemployment rate is bounded — it cannot go below 0% or much above 25% without the economy ceasing to function. A bounded series cannot have a true unit root in the strict sense, because a random walk has unbounded variance. So the unemployment rate must, at long horizons, be stationary. The persistence we see in 75 years of data is just \(\phi\) very close to 1, not \(\phi = 1\) exactly.

  • Unit root camp: The persistence is so strong that for any practical sample size, the unemployment rate is operationally indistinguishable from a unit root process. Whether the “true” \(\phi\) is 0.998 or 1.000 makes no difference for inference at horizons we care about. For modeling purposes, treat it as \(I(1)\).

Both arguments have merit. The data themselves cannot resolve the question — even with \(T \approx 930\) observations, the test power is not enough to distinguish \(\phi = 0.998\) from \(\phi = 1.000\). This is exactly the borderline case the module has been preparing you for. The cost asymmetry resolves the dilemma: when in doubt, treat as \(I(1)\), because spurious regression is more costly than over-differencing.

This sets up Problem Set 1, where you will run the same workflow on a different FRED series and write up your verdict and reasoning.


Common Pitfalls and Misconceptions

  1. “ADF failing to reject means the series has a unit root.” No. Failing to reject means the data are insufficient to rule out the unit root. This could be because the series truly has a unit root, or because the test has low power against a near-unit-root stationary alternative. Don’t confuse “I cannot reject” with “I have proven.”

  2. “The DF \(t\)-statistic follows a \(t\)-distribution.” No. It follows the Dickey-Fuller distribution, which is not symmetric and has a heavier left tail. Use MacKinnon critical values, not standard \(t\)-tables. The reason is the same as for spurious regression: standard OLS asymptotics break down with non-stationary regressors.

  3. “Always use the trend specification to be safe.” No. Including unnecessary deterministic terms reduces test power. Match the specification to what you see in the plot. Use the trend specification only when there is an obvious trend.

  4. “ADF and KPSS test the same thing.” No. They have opposite nulls. Use them as complements: confirmatory evidence when they agree, honest “don’t know” when both fail to reject.

  5. “Over-differencing is harmless.” No. Differencing a trend-stationary series adds a non-invertible \((1-L)\) MA factor. It is a pure MA(1) with \(\theta=-1\) only when the stationary deviation was white noise; otherwise the original ARMA dynamics remain.

  6. “The rule of thumb is just folklore.” It is a rule of thumb, but it has a real basis in the variance ratio of stationary AR(1) processes. It is not a substitute for formal testing, but it captures the same intuition more quickly.

  7. “More tests are always better.” Up to a point. Running ADF, KPSS, PP, and ERS on the same series gives you four pieces of evidence, but if they all use the same underlying data and similar testing principles, they are not fully independent. The best complement to ADF is KPSS (because the nulls are reversed); adding PP or ERS gives you incremental information at most.


Connection to Enders

  • The difference operator and difference equations: Enders Ch. 1, pp. 1-46
  • Stochastic difference equations: Enders Ch. 2, pp. 47-52
  • The Dickey-Fuller test, all three specifications, and critical values: Enders Ch. 4, pp. 206–209
  • Sequential testing and choosing deterministic terms: Enders Ch. 4, pp. 235–238
  • Phillips-Perron: Enders Ch. 4 gives a brief pointer on p. 221; see the Supplementary Manual, §4.6, for the fuller treatment
  • KPSS and the reverse-null approach: Enders Ch. 4, brief discussion on p. 214
  • Order of integration and trends: Enders Ch. 4, pp. 181-189
  • Trend-stationary vs difference-stationary: Enders Ch. 4, pp. 189-200

A convention warning when you cross-reference. Enders writes the AR(1) as \(y_t = a_0 + a_1 y_{t-1} + \epsilon_t\) (his \(a_1\) is our \(\phi\)), and he states stationarity conditions in terms of characteristic roots lying inside the unit circle. This course states the equivalent condition as roots of the lag polynomial \(\Phi(L) = 0\) lying outside the unit circle. Both are correct — they describe reciprocal roots — but keep the framing straight when reading Ch. 4. Hamilton (Ch. 17) is the rigorous reference for the asymptotics behind the DF distribution; Hyndman & Athanasopoulos discuss unit-root testing from a forecasting-workflow angle (and use the KPSS test as their default, the reverse of our ADF-first habit).

References:

  • Dickey, D.A. and Fuller, W.A. (1979), “Distribution of the Estimators for Autoregressive Time Series with a Unit Root,” Journal of the American Statistical Association, 74, 427-431.
  • Dickey, D.A. and Fuller, W.A. (1981), “Likelihood Ratio Statistics for Autoregressive Time Series with a Unit Root,” Econometrica, 49, 1057-1072. (Joint \(F\)-tests \(\Phi_1\), \(\Phi_2\), \(\Phi_3\).)
  • Dickey, D.A. and Pantula, S.G. (1987), “Determining the Order of Differencing in Autoregressive Processes,” Journal of Business & Economic Statistics, 5, 455-461.
  • Pantula, S.G. (1989), “Testing for Unit Roots in Time Series Data,” Econometric Theory, 5, 256-271.
  • MacKinnon, J.G. (1996), “Numerical Distribution Functions for Unit Root and Cointegration Tests,” Journal of Applied Econometrics, 11, 601-618.
  • MacKinnon, J.G. (2010), “Critical Values for Cointegration Tests,” Queen’s Economics Department Working Paper No. 1227. (The displayed response-surface coefficients are from Table 2.)
  • Kwiatkowski, D., Phillips, P.C.B., Schmidt, P., and Shin, Y. (1992), “Testing the Null Hypothesis of Stationarity Against the Alternative of a Unit Root,” Journal of Econometrics, 54, 159-178.
  • Phillips, P.C.B. (1987), “Time Series Regression with a Unit Root,” Econometrica, 55, 277-301.
  • Schwert, G.W. (1989), “Tests for Unit Roots: A Monte Carlo Investigation,” Journal of Business and Economic Statistics, 7, 147-159.
  • Elliott, G., Rothenberg, T.J., and Stock, J.H. (1996), “Efficient Tests for an Autoregressive Unit Root,” Econometrica, 64, 813-836.
  • Granger, C.W.J. and Newbold, P. (1974), “Spurious Regressions in Econometrics,” Journal of Econometrics, 2, 111-120.

Practice Problems

Core Practice

  1. Variance ratio computation. Compute the theoretical ratio \(\sigma_{\Delta y} / \sigma_y\) analytically for an AR(1) with \(\phi = 0.3, 0.7, 0.95\). Then verify by simulating long series (\(T = 5000\)) and computing the empirical ratio. Do they match?

  2. Power of the ADF test. Conduct a Monte Carlo study to compute the power of the ADF test against a stationary AR(1) with \(\phi = 0.95\), for sample sizes \(T \in \{50, 100, 200, 500, 1000\}\). Plot power as a function of \(T\). At what sample size does the test achieve at least 50% power? At least 80%? Repeat the exercise for \(\phi = 0.90\) and \(\phi = 0.80\) and overlay the curves on a single plot. (This style of power study reappears on the problem sets — it is worth doing carefully now.)

  3. The over-differencing trap in practice. Generate a trend-stationary series with \(T = 300\): \[y_t = 5 + 0.05\, t + u_t, \qquad u_t = 0.5\, u_{t-1} + \epsilon_t\] where \(u_t\) is the stationary AR(1) deviation from the trend (the same construction as Section 2.3) and \(\epsilon_t\) is white noise. Then:

      1. Detrend correctly by regressing \(y_t\) on a constant and \(t\). Plot the residuals and compute their ACF.
      1. Difference \(y_t\) to get \(\Delta y_t\). Plot it and compute its ACF.
      1. Fit an ARMA(1,1) to the differenced series using arima(diff(y), order = c(1, 0, 1)). What are the estimated \(\hat{\phi}\) and \(\hat{\theta}\)? Is \(\hat{\theta}\) close to \(-1\), and does \(\hat{\phi}\) retain the original AR persistence? Compare this with a pure MA(1) fit and explain why the ARMA(1,1) matches the DGP more closely.
    • Discuss what each treatment tells you.
  4. ADF specification matters. Generate two series, both with \(T = 200\):

    • Series A (trend-stationary): 0.05 * (1:200) + arima.sim(n = 200, list(ar = 0.5))
    • Series B (random walk with drift): cumsum(rnorm(200) + 0.05)

    Plot both. They will look similar. Now run the ADF test on each, with both the drift specification and the trend specification. What do you conclude in each case? Which specification gives you the right answer for which series?

  5. Real data workflow. Pull a different FRED series (suggestions: real GDP GDPC1, industrial production INDPRO, the federal funds rate FEDFUNDS, the CPI CPIAUCSL). Run the full stationarity workflow: visual inspection, ACF, ADF (with the appropriate specification), KPSS, and the rule of thumb. Synthesize the evidence and write up a verdict in 2-3 paragraphs. Include the cost asymmetry in your reasoning.


Key Takeaways

  1. The difference operator \(\Delta = (1 - L)\) is the discrete-time analog of a derivative. It is both the dependent variable in the DF test and the tool we use to make non-stationary series stationary.

  2. Visual inspection of the series and the ACF are the cheapest evidence and are sufficient for the easy cases. They cannot distinguish trend-stationary from difference-stationary, and they fail in small samples.

  3. The Dickey-Fuller test rearranges the AR(1) into a regression of \(\Delta y_t\) on \(y_{t-1}\). The test of \(H_0: \phi = 1\) becomes a test of \(H_0: \gamma = 0\) where \(\gamma = \phi - 1\). The \(t\)-statistic does not follow a standard \(t\)-distribution under the null because non-stationary regressors break OLS asymptotics — you must use MacKinnon critical values.

  4. There are three ADF specifications (none, drift, trend). Choose based on what you see in the plot. Critical values differ across specifications.

  5. KPSS is the reverse-null test: \(H_0\) = stationary, \(H_1\) = unit root. Use ADF and KPSS together for confirmatory analysis.

  6. All unit root tests have low power against highly persistent stationary alternatives. With \(\phi = 0.95\) and \(T = 100\), ADF rejects less than 10% of the time. This is why we use multiple sources of evidence.

  7. The rule of thumb \(\sigma_{\Delta y} / \sigma_y < 0.5\) flags series with persistence above roughly \(\phi > 0.875\) as candidates for differencing.

  8. Order of integration \(I(d)\): the number of differences needed to achieve stationarity. Most economic series are \(I(0)\) or \(I(1)\); \(I(2)\) is rare; higher orders are essentially never seen.

  9. The over-differencing trap: differencing a trend-stationary series adds a non-invertible \((1-L)\) MA factor. It yields a pure MA(1) with \(\theta=-1\) only for white-noise deviations; otherwise the original ARMA dynamics remain. Use the trend-augmented ADF as evidence when choosing between detrending and differencing, then diagnose the resulting model.

  10. The cost asymmetry: when in doubt, difference. Spurious regression is a more dangerous inference error than over-differencing. But be aware of the over-differencing tax — it shows up in your ARMA model fits later.


Looking Ahead — Module 3

Two loose ends from this module become Module 3’s opening material:

  • The over-differencing trap named a boundary. We showed that differencing a trend-stationary series adds an MA unit-root factor \((1-L)\) — a violation of invertibility. In the white-noise special case this is a pure MA(1); with ARMA deviations, the original dynamics remain. Module 3 defines invertibility properly and probes that boundary experimentally with a nearly non-invertible MA(1).
  • The power problem needs a complementary tool. Formal tests cannot reliably separate \(\phi = 0.95\) from \(\phi = 1.0\) in realistic samples. Module 3 adds the correlogram fingerprint — reading ACF/PACF pairs to identify AR(\(p\)) and MA(\(q\)) structure — which works alongside testing rather than replacing it. The mixing board gets its first new sliders: higher-order AR memory (\(\phi_1, \ldots, \phi_p\)) and memory of past shocks (\(\theta_1, \ldots, \theta_q\)).