library(ggplot2)
library(urca)
# Stamped copies of the canonical course helpers.
source("R/ar1_simulator.R") # the Module 1 AR(1) DGP, written as a loop
source("R/mackinnon_cv.R") # MacKinnon critical values, drift specification
source("R/mystery_series.R") # four sealed series -- source it, don't read it
# Cached FRED pulls, so this lab renders with no key and no internet.
# See data/README.md for provenance and the one known gap.
unrate_raw <- read.csv("data/UNRATE.csv", stringsAsFactors = FALSE)
theme_set(theme_bw(base_size = 13))Module 2 Lab: Testing for Stationarity
Econ 6376 — Applied Time Series Econometrics · AI Learning Pack
How to use this lab
Everything here runs as shipped. No API key, no internet, no blanks to fill in. That is deliberate: the point of this lab is not to get the code working, it is to find out whether your intuition about stationarity testing is any good.
So there is one rule that makes the whole thing worth doing:
Write your prediction down before you run the chunk.
Not in your head — in your head you will “sort of” have predicted whatever happens. Write it in LEARNING_LOG.md, or in a comment, or on paper. Being wrong on paper is the single most useful thing that can happen to you today. This module ends with four sealed mystery series where that habit stops being practice and becomes the whole exercise.
If you are working with an AI tutor, it will ask you for that prediction before it runs anything meaningful, and it will run the code either way. It is not gatekeeping you — it is trying to make sure you get the “huh, I didn’t expect that” moment, because that is the one you remember in November.
Where you are on the mixing board: the master equation is
\[ y_t = \alpha + \delta t + \sum_{j=1}^{p}\phi_j y_{t-j} + \sum_{l=1}^{q}\theta_l \epsilon_{t-l} + \epsilon_t . \]
Module 2 turns on no new sliders. The DGP is still Module 1’s AR(1) — \(\alpha\) and \(\phi_1\) — and what changes is the question we put to it:
\[ \textbf{is } \phi_1 = 1 \textbf{?} \]
Along the way the trend slider \(\delta t\) makes its first cameo, because “trend-stationary or difference-stationary?” turns out to be part of the same testing problem. Everything below — a test with a broken null distribution, a mirror-image test, a power study, an over-differencing autopsy, and a string of verdicts — is that one question, asked carefully.
1. Setup
Three functions and two cached data sets.
Two warnings about R/mystery_series.R, both worth reading now:
- Source it, don’t open it. The bottom half of the file, below a marked
SPOILERSbanner, contains the true DGP of every mystery series. Section 8 is only worth doing if you haven’t peeked. (If you must look at the code, read only above the banner.) - It re-seeds your session. Unlike every other course helper,
mystery_series()sets a seed internally — that is what makes the series identical for every student and every tutor, so committed verdicts are comparable. The side effect: calling it overwrites your session’s RNG state. If your own code depends on a seed, callset.seed()aftermystery_series(), not before.
So what: the toolkit is loaded — a DGP whose \(\phi\) we control, the correct critical values for the test we’re about to build, and four series that will not tell you what they are.
2. The regression you can’t test naively
We want to test \(H_0\!: \phi = 1\) against \(H_1\!: \phi < 1\) in
\[ y_t = \alpha + \phi\, y_{t-1} + \epsilon_t . \]
The obvious move — estimate \(\hat\phi\) by OLS, form a \(t\)-statistic for \(\phi = 1\), look up a \(t\)-table — fails, and it fails for a reason you already met: under \(H_0\) the regressor \(y_{t-1}\) is a random walk, and Module 1’s spurious-regression disaster was exactly what happens when OLS meets non-stationary regressors. The sample moments that standard asymptotics lean on do not converge the way the \(t\)-table assumes.
The Dickey-Fuller fix starts with an honest rearrangement. Subtract \(y_{t-1}\) from both sides:
\[ \Delta y_t = \alpha + \gamma\, y_{t-1} + \epsilon_t, \qquad \gamma = \phi - 1 . \]
Now the test is \(H_0\!: \gamma = 0\) (unit root) against \(H_1\!: \gamma < 0\) — one-sided, left tail. The dependent variable \(\Delta y_t\) is stationary whether the null or the alternative is true, and the hypothesis maps onto a single coefficient. The \(t\)-ratio on \(\hat\gamma\) gets its own name: \(\tau\), the Dickey-Fuller statistic.
One notation flag before anything runs: this unsubscripted \(\gamma\) is a regression coefficient. It is not the autocovariance — those are \(\gamma_k\), always with a lag subscript. Same letter, unrelated objects, both standard.
We’re about to simulate one random walk (\(T = 200\), so \(H_0\) is true by construction) and run the DF regression on it. Two predictions: (a) roughly what value will \(\hat\phi = \hat\gamma + 1\) take? (b) will the \(t\)-ratio \(\tau\) be near \(0\), near \(-1\) to \(-2\), or below \(-3\)? Commit to both.
set.seed(6376)
T_n <- 200
y <- cumsum(rnorm(T_n)) # random walk: phi = 1, alpha = 0
dy <- diff(y) # Delta y_t, length T - 1
y_lag <- y[-T_n] # y_{t-1}, aligned
df_reg <- lm(dy ~ y_lag)
round(coef(summary(df_reg)), 4) Estimate Std. Error t value Pr(>|t|)
(Intercept) 0.0901 0.0840 1.0718 0.2851
y_lag -0.0135 0.0114 -1.1884 0.2361
tau_hat <- summary(df_reg)$coefficients["y_lag", "t value"]
c(phi_hat = round(1 + coef(df_reg)["y_lag"], 4), tau = round(tau_hat, 3))phi_hat.y_lag tau
0.9865 -1.1880
\(\hat\phi \approx 0.986\) — close to 1, as it should be, since OLS is still consistent here. And \(\tau \approx -1.19\).
Now the trap. If you compared \(-1.19\) to the one-sided 5% critical value from the standard normal, \(-1.645\), you would say “fail to reject” and feel like you had run a proper test. You would be right by accident: the number \(-1.645\) has nothing to do with this regression. The question “what distribution does \(\tau\) actually follow when \(H_0\) is true?” cannot be answered by habit — Section 3 answers it by brute force.
So what: the DF rearrangement gives us a clean statistic to compute; it does not give us permission to use the tables we already own.
3. The assumption-break: the null distribution isn’t the one you were promised
This is the module’s centerpiece. Every \(t\)-statistic you have ever compared to \(\pm 1.96\) leaned on an assumption — well-behaved regressors — that the unit-root null destroys by construction. So we do the honest thing: simulate the null thousands of times and look at the distribution of \(\tau\).
2000 independent random walks, one DF regression each, 2000 values of \(\tau\). Where is the histogram centered — (a) at \(0\), like a \(t\) distribution, (b) around \(-1.5\), or (c) around \(-3\)? And is it symmetric? Pick before you run.
set.seed(42)
n_sims <- 2000
T_n <- 200
df_stats <- replicate(n_sims, {
y <- cumsum(rnorm(T_n)) # H_0 true by construction
dy <- diff(y)
y_lag <- y[-T_n]
summary(lm(dy ~ y_lag))$coefficients["y_lag", "t value"]
})
c(mean = round(mean(df_stats), 3),
sd = round(sd(df_stats), 3),
q05 = round(unname(quantile(df_stats, 0.05)), 3)) mean sd q05
-1.542 0.853 -2.883
ggplot(data.frame(tau = df_stats), aes(tau)) +
geom_histogram(aes(y = after_stat(density)), bins = 60,
fill = "steelblue4", alpha = 0.65) +
stat_function(fun = dnorm, colour = "firebrick", linewidth = 0.9) +
geom_vline(xintercept = -1.645, linetype = "dashed",
colour = "firebrick") +
geom_vline(xintercept = quantile(df_stats, 0.05), linetype = "dashed",
colour = "steelblue4") +
labs(title = "2000 DF t-statistics under a true unit-root null",
subtitle = paste("Red curve: the N(0,1) your t-table assumes.",
"Dashed lines: the two competing 5% cutoffs"),
x = expression(tau ~ "under" ~ H[0]: phi == 1), y = "density")The histogram is centered near \(-1.54\), not zero, with a standard deviation around \(0.85\), not one. It is a different distribution — the Dickey-Fuller distribution — and its 5% quantile is nowhere near \(-1.645\).
Now the two payoffs. First: MacKinnon’s critical values are nothing more than a very careful, very large version of the simulation you just ran.
c(empirical_q05 = round(unname(quantile(df_stats, 0.05)), 3),
mackinnon_cv = round(mackinnon_cv(T = 200, level = 0.05), 3))empirical_q05 mackinnon_cv
-2.883 -2.876
Our 2000-replication quantile lands at \(-2.883\); MacKinnon’s response surface says \(-2.876\). Off by less than a hundredth — and his number comes from vastly more replications across many sample sizes, smoothed into the formula sitting in R/mackinnon_cv.R. Nothing mystical: empirical quantiles of exactly this experiment.
Second: the price of using the wrong table.
c(size_wrong_cv = mean(df_stats < -1.645),
size_correct_cv = mean(df_stats < mackinnon_cv(200))) size_wrong_cv size_correct_cv
0.4680 0.0515
A “5% test” run against the normal critical value rejects a true unit-root null about 47% of the time. Not five percent — forty-seven. Every one of those rejections would send an analyst off to run a levels regression on data that needed differencing, which is how Module 1’s spurious-regression trap gets baited in practice. Against the correct critical value the rejection rate sits where it belongs, near 5%.
Change T_n to 50 and re-run the three chunks above. Predict first: does the empirical 5% quantile move left (more negative) or right? Then check it against mackinnon_cv(T = 50) — the formula’s whole job is tracking exactly this sample-size dependence.
So what: the DF test is not a new test statistic — it is the old \(t\)-ratio equipped with the critical values it actually has under the unit-root null. Use the right table and the machinery works; use the familiar one and your error rate is off by a factor of nine.
4. Building the ADF, and choosing its specification
Real series are richer than an AR(1), and leftover serial correlation in \(\epsilon_t\) distorts the DF test. The fix — the augmented DF — adds lagged differences to soak up the short-run dynamics:
\[ \Delta y_t = \alpha + \gamma\, y_{t-1} + \sum_{i=1}^{k}\beta_i\, \Delta y_{t-i} + \epsilon_t . \]
The test is still \(H_0\!: \gamma = 0\). Let’s build it from scratch once — diff(), embed(), lm() — so urca::ur.df() never gets to be a black box.
The input is a stationary AR(1) with \(\phi = 0.7\) and \(T = 250\) — far from the boundary. Does the ADF reject the unit root: (a) comfortably, \(\tau\) well below the critical value, (b) barely, or (c) not at all? Pick one.
set.seed(6376)
y <- ar1_simulator(n = 250, alpha = 0, phi = 0.7, sigma = 1, burn_in = 200)
T_n <- length(y)
k_lags <- 4
dy <- diff(y)
X <- embed(dy, k_lags + 1) # col 1: dy_t; cols 2..: dy_{t-1}..dy_{t-k}
dy_t <- X[, 1]
dy_lags <- X[, -1, drop = FALSE]
y_lag <- as.numeric(y[(k_lags + 1):(T_n - 1)]) # y_{t-1}, aligned
adf_reg <- lm(dy_t ~ y_lag + dy_lags)
tau_hat <- summary(adf_reg)$coefficients["y_lag", "t value"]
critical <- mackinnon_cv(T = T_n, level = 0.05)
c(tau = round(tau_hat, 3), cv_05 = round(critical, 3),
reject = tau_hat < critical) tau cv_05 reject
-5.293 -2.873 1.000
\(\tau \approx -5.29\) against a critical value of \(-2.87\): a comfortable, correct rejection. Now confirm the package computes the identical number:
adf_pkg <- ur.df(y, type = "drift", lags = 4)
c(by_hand = round(tau_hat, 4), ur_df = round(adf_pkg@teststat[1], 4))by_hand ur_df
-5.293 -5.293
adf_pkg@cval 1pct 5pct 10pct
tau2 -3.46 -2.88 -2.57
phi1 6.52 4.63 3.81
Same statistic to four decimals. (ur.df() prints tabulated critical values — \(-2.88\) at 5% — where our formula gives \(-2.873\); same decision, slightly different rounding conventions.) In real work you’d let BIC pick \(k\): ur.df(y, type = "drift", lags = 12, selectlags = "BIC"). One software detail worth knowing: that search runs over \(k = 1, \ldots, 12\) only — if \(k = 0\) is a plausible candidate, fit it separately.
The three specifications
What you just ran is the drift specification. There are three, and the choice is where most applied mistakes happen:
- “none” — \(\Delta y_t = \gamma y_{t-1} + \sum \beta_i \Delta y_{t-i} + \epsilon_t\). Zero mean, no trend. Rare in economics.
- “drift” — adds \(\alpha\). The default for series with a level but no trend: unemployment, interest rates, inflation.
- “trend” — adds \(\alpha + \delta t\). For visibly trending series — and this is \(\delta t\)’s cameo: the trend specification is the tool that separates trend-stationary (stationary wiggles around a deterministic trend; detrend it) from difference-stationary (a unit root, possibly with drift; difference it).
Each specification has its own critical values — more deterministic terms push them further left. Reading a trend-spec statistic against drift-spec critical values is a genuine, common error.
Does the choice actually matter? Watch it flip a verdict.
The next series is trend-stationary by construction: \(y_t = 0.05t + u_t\) with \(u_t = 0.5 u_{t-1} + \epsilon_t\) — no unit root anywhere. We run the ADF twice, once with drift, once with trend. Which is true: (a) both reject the unit root, (b) neither does, or (c) trend rejects and drift doesn’t? Commit.
set.seed(456)
n <- 250
u <- numeric(n)
u[1] <- 0
for (t in 2:n) u[t] <- 0.5 * u[t - 1] + rnorm(1) # stationary deviations
y_trendstat <- ts(0.05 * (1:n) + u) # riding a deterministic trend
drift_test <- ur.df(y_trendstat, type = "drift", lags = 4)
trend_test <- ur.df(y_trendstat, type = "trend", lags = 4)
data.frame(
spec = c("drift", "trend"),
tau = round(c(drift_test@teststat[1], trend_test@teststat[1]), 3),
cv_05 = c(drift_test@cval[1, 2], trend_test@cval[1, 2]),
reject = c(drift_test@teststat[1] < drift_test@cval[1, 2],
trend_test@teststat[1] < trend_test@cval[1, 2])
) spec tau cv_05 reject
1 drift -1.461 -2.88 FALSE
2 trend -6.400 -3.43 TRUE
The drift specification sees \(\tau \approx -1.46\) and cannot reject — it has no trend term, so the deterministic climb masquerades as a stochastic trend. The trend specification sees \(\tau \approx -6.40\) and rejects emphatically. Same data, opposite verdicts, and only one specification matches the DGP.
The practical rule: choose the specification from the plot — level but no trend, use drift; visible trend, use trend — knowing that unnecessary deterministic terms drain power while missing ones wreck the test entirely.
So what: the ADF statistic is only half the test; the deterministic specification is the other half, and it can single-handedly flip your conclusion.
5. KPSS: the mirror image
Everything so far shares one structural feature: the null is the unit root. Failing to reject therefore proves nothing — maybe there’s a unit root, maybe the test just couldn’t see past it (Section 6 shows how often that happens). The KPSS test flips the burden of proof: its null is stationarity, its alternative is the unit root.
Two mechanical consequences, both frequent sources of error:
- A large KPSS statistic is evidence against stationarity — the opposite comparison direction from ADF, where rejection lives in the far-left tail.
- KPSS has its own specification choice:
type = "mu"(stationary around a level — pairs with ADF drift) andtype = "tau"(stationary around a trend — pairs with ADF trend).
Two clean series: a long stationary AR(1) (\(\phi = 0.7\), \(T = 1000\)) and a random walk (\(T = 300\)). For each, ADF and KPSS both run. Sketch the 2×2 outcome you expect — which tests reject on which series?
set.seed(101)
y_stat <- ar1_simulator(n = 1000, alpha = 0, phi = 0.7, sigma = 1, burn_in = 200)
set.seed(202)
y_rw <- ts(cumsum(rnorm(300)))
grid <- function(y, label) {
adf <- ur.df(y, type = "drift", lags = 4)
kpss <- ur.kpss(y, type = "mu")
data.frame(
series = label,
adf_tau = round(unname(adf@teststat[1]), 3),
adf_reject = unname(adf@teststat[1]) < adf@cval[1, 2],
kpss_stat = round(unname(kpss@teststat), 3),
kpss_reject = unname(kpss@teststat) > kpss@cval[1, "5pct"]
)
}
rbind(grid(y_stat, "stationary AR(1), T = 1000"),
grid(y_rw, "random walk, T = 300")) series adf_tau adf_reject kpss_stat kpss_reject
1 stationary AR(1), T = 1000 -10.711 TRUE 0.171 FALSE
2 random walk, T = 300 -1.995 FALSE 3.578 TRUE
Both series land on the diagonal of the confirmatory table: the stationary series has ADF rejecting (\(\tau \approx -10.7\)) while KPSS stays quiet (\(0.17\), under its 5% critical value of \(0.463\)); the random walk has ADF silent (\(\tau \approx -2.0\)) while KPSS shouts (\(3.58\)). When the two tests agree like this, you have a strong verdict. The full grid:
| ADF says | KPSS says | Conclusion |
|---|---|---|
| reject unit root | fail to reject | strong evidence of stationarity |
| fail to reject | reject stationarity | strong evidence of a unit root |
| reject | reject | contradiction — break or misspecification suspected |
| fail to reject | fail to reject | inconclusive — defer to the other evidence |
The off-diagonal rows are not failures of the method — they are the method telling you something you need to know. You will meet the contradiction row on real data before this lab ends.
One honest wrinkle before moving on. KPSS is not magic either — watch what it does to a series we know is stationary, at a realistic sample size:
set.seed(6376)
y_short <- ar1_simulator(n = 250, alpha = 0, phi = 0.7, sigma = 1, burn_in = 200)
kpss_short <- ur.kpss(y_short, type = "mu")
c(kpss = round(unname(kpss_short@teststat), 3),
cv_05 = kpss_short@cval[1, "5pct"],
rejects_a_true_null = unname(kpss_short@teststat) > kpss_short@cval[1, "5pct"]) kpss cv_05 rejects_a_true_null
0.618 0.463 1.000
This is the same \(\phi = 0.7\) DGP that Section 4’s ADF called correctly — and KPSS rejects stationarity on it (\(0.618 > 0.463\)). At moderate \(T\), KPSS over-rejects persistent-but-stationary series. Neither test is the referee; they are two witnesses with opposite biases, which is exactly why the course makes them testify together.
So what: ADF and KPSS asking opposite questions is what makes their agreement informative — and their individual verdicts fallible, which the next section quantifies.
6. Power: the most important practical fact in this module
Here is the question that decides how much weight a “fail to reject” can carry: if the series really is stationary but barely — \(\phi = 0.95\) — how often does the ADF test figure that out?
At the 5% level, with \(\phi = 0.95\) truly stationary: what fraction of the time does ADF correctly reject the unit root at \(T = 100\)? (a) around 60–80%, (b) around 30–50%, (c) around 10%. And what about \(T = 400\)? Commit to both numbers.
set.seed(2024)
n_sims <- 400
sample_sizes <- c(50, 100, 200, 400)
power_at <- sapply(sample_sizes, function(T_p) {
mean(replicate(n_sims, {
y <- ar1_simulator(n = T_p, alpha = 0, phi = 0.95, sigma = 1, burn_in = 200)
ur.df(y, type = "drift", lags = 4)@teststat[1] < mackinnon_cv(T_p, 0.05)
}))
})
power_tab <- data.frame(T = sample_sizes, power = power_at)
power_tab T power
1 50 0.0700
2 100 0.1000
3 200 0.2725
4 400 0.7725
ggplot(power_tab, aes(T, power)) +
geom_hline(yintercept = 0.05, linetype = "dashed", colour = "grey55") +
geom_hline(yintercept = 0.80, linetype = "dashed", colour = "firebrick") +
geom_line(colour = "steelblue4", linewidth = 0.8) +
geom_point(colour = "steelblue4", size = 2.5) +
labs(title = "Power of the 5% ADF test against a stationary phi = 0.95",
subtitle = paste("Dashed grey: the nominal size (what a useless test",
"achieves). Dashed red: the conventional 80% target"),
x = "sample size T", y = "P(reject unit root)")Read the numbers without flinching: at \(T = 50\) power is about \(0.07\) — barely above the 5% a coin-flip test would deliver. At \(T = 100\) it is about \(0.10\): the truly stationary series is misclassified roughly 90% of the time. Even at \(T = 200\) the test finds the truth only about a quarter of the time; not until \(T = 400\) does it become respectable (about \(0.77\) here). And \(\phi = 0.95\) with \(T \le 200\) is not an exotic corner case — it is what typical macroeconomic samples look like.
This is not an ADF defect. KPSS, PP, and ERS all face versions of the same wall: near the boundary, finite samples simply do not contain enough information. Which reframes everything: a failure to reject is weak evidence, borderline verdicts need multiple witnesses, and an honest “the data cannot tell” is sometimes the correct conclusion.
The rule of thumb
The fourth and cheapest source of evidence. Compute \(R = \sigma_{\Delta y} / \sigma_y\) — the ratio of the standard deviation of the differences to that of the levels. The logic is one line of algebra: for a stationary AR(1), \(\operatorname{Var}(\Delta y_t) / \operatorname{Var}(y_t) = 2(1 - \phi)\) exactly, so
\[ R = \sqrt{2(1-\phi)} . \]
Before running: for white noise (\(\phi = 0\)), is \(R\) above or below 1? Most people’s instinct says differencing always shrinks a series. Does it?
phis <- c(0, 0.5, 0.9, 0.95)
thumb <- data.frame(
phi = phis,
sample_R = sapply(phis, function(p) {
set.seed(6376)
y <- ar1_simulator(n = 20000, alpha = 0, phi = p, sigma = 1, burn_in = 500)
sd(diff(y)) / sd(y)
}),
theory_R = sqrt(2 * (1 - phis))
)
round(thumb, 3) phi sample_R theory_R
1 0.00 1.415 1.414
2 0.50 0.999 1.000
3 0.90 0.452 0.447
4 0.95 0.322 0.316
set.seed(99)
y_rw_long <- cumsum(rnorm(20000))
c(random_walk_R = round(sd(diff(y_rw_long)) / sd(y_rw_long), 3))random_walk_R
0.019
The samples track the theory: \(R \approx 1.41\) at \(\phi = 0\) (differencing white noise makes it noisier), \(1.0\) at \(\phi = 0.5\), \(0.45\) at \(\phi = 0.9\), \(0.32\) at \(\phi = 0.95\) — and effectively \(0\) for the random walk, whose levels variance dwarfs its differences.
The course’s working rule: \(R < 0.5\) — which the algebra maps to roughly \(\phi > 0.875\) — flags the series for differencing. But keep its epistemics straight: a small \(R\) is a high-persistence flag, not a unit-root test. It cannot tell \(0.95\) from \(1.0\) (both give small \(R\)), and Section 8 will show a trend-stationary series tripping it too. It earns its keep by being instant, assumption-light, and a good tiebreaker when the formal tests disagree.
So what: near the unit-root boundary the formal tests are nearly blind — so the verdict has to come from converging evidence: plot, ACF, ADF, KPSS, and the ratio, weighed together.
7. The over-differencing trap
The cost asymmetry says: when in doubt, difference. Before you internalize that rule, you should see — precisely — what it costs when the doubt resolves the other way and the series was trend-stationary all along.
The algebra first. Take the simplest trend-stationary DGP, white noise around a line, and difference it:
\[ y_t = \alpha + \delta t + \epsilon_t \;\;\Longrightarrow\;\; \Delta y_t = \delta + \epsilon_t - \epsilon_{t-1} . \]
The trend collapses to the constant \(\delta\) — fine. But the noise is now \(\epsilon_t - \epsilon_{t-1}\): an MA(1) with \(\theta = -1\). Its MA polynomial \(\Theta(L) = 1 - L\) has its root exactly on the unit circle. Differencing did not remove a unit root — there was none — it manufactured one, in the MA polynomial. That is the definition of a non-invertible process, a boundary Module 3 explores properly.
Does estimation actually feel this, or is it a theoretical curiosity?
We difference a trend-stationary series (\(T = 300\), white-noise deviations) and fit an MA(1) to \(\Delta y_t\). Where does \(\hat\theta\) land: (a) scattered somewhere in \((-1, 0)\) depending on the draw, (b) pinned hard against \(-1\), or (c) near \(0\)? Pick one.
set.seed(456)
n <- 300
y_twn <- 5 + 0.05 * (1:n) + rnorm(n) # trend + pure white noise
fit_ma1 <- arima(diff(y_twn), order = c(0, 0, 1))
round(coef(fit_ma1), 4) ma1 intercept
-1.0000 0.0499
\(\hat\theta = -1.0000\). Not “near” \(-1\) — the optimizer slammed into the invertibility boundary and stopped, exactly where the algebra said the truth lives. In the wild this fingerprint arrives with convergence warnings, fragile standard errors, and an estimate that refuses to move. When a fitted MA coefficient sits at \(-1\), your first suspect is always: this series was differenced once too many times.
Now the general case, which is subtler and more important. If the deviations have their own dynamics — say AR(1), \(u_t = 0.6 u_{t-1} + \epsilon_t\) — differencing does not produce a pure MA(1). It keeps the original dynamics and multiplies them by the extra \((1-L)\) factor:
\[ (1 - 0.6L)\,\Delta u_t = (1 - L)\,\epsilon_t \;\;\Longrightarrow\;\; \Delta y_t \sim \text{ARMA}(1,1) \text{ with } \phi = 0.6,\; \theta = -1 . \]
set.seed(456)
u <- numeric(n)
u[1] <- 0
for (t in 2:n) u[t] <- 0.6 * u[t - 1] + rnorm(1)
y_tar <- 5 + 0.05 * (1:n) + u # trend + AR(1) deviations
fit_arma <- arima(diff(y_tar), order = c(1, 0, 1))
fit_ma_only <- arima(diff(y_tar), order = c(0, 0, 1))
data.frame(
model = c("ARMA(1,1) [matches the algebra]", "MA(1) only [misspecified]"),
ar1 = c(round(coef(fit_arma)["ar1"], 3), NA),
ma1 = round(c(coef(fit_arma)["ma1"], coef(fit_ma_only)["ma1"]), 3),
truth = c("phi = 0.6, theta = -1", "no pure MA(1) exists here")
) model ar1 ma1 truth
ar1 ARMA(1,1) [matches the algebra] 0.594 -1.000 phi = 0.6, theta = -1
MA(1) only [misspecified] NA -0.307 no pure MA(1) exists here
The correctly specified ARMA(1,1) recovers both halves of the story: \(\hat\phi \approx 0.59\) — the original persistence, intact — and \(\hat\theta = -1\), the manufactured boundary root. The misspecified pure MA(1) finds \(\hat\theta \approx -0.31\): the AR dynamics contaminate the fit so badly it cannot even locate the boundary. Over-differencing doesn’t erase your series’ dynamics; it smuggles a non-invertible factor in alongside them.
So what: the cost asymmetry stands — spurious regression is still the worse error — but differencing is not free. The \(\hat\theta \approx -1\) fingerprint is how the mistake announces itself later, and the trend-spec ADF from Section 4 is how you avoid making it now.
8. The verdict: from labeled practice to sealed series to real data
Everything so far taught you the instruments. This section is the flight hours. The workflow, every time:
- Plot the series — level? trend? wandering?
- ACF — fast decay or near-flat persistence?
- ADF with the specification the plot suggests; KPSS with its match (
muwith drift,tauwith trend). - Ratio \(\sigma_{\Delta y}/\sigma_y\) as the sanity check.
- Verdict — with the cost asymmetry breaking genuine ties toward differencing.
One helper packages the evidence so we can run it repeatedly. Read it — it is nothing but Sections 4–6 stapled together. Note what it deliberately does not do: it never decides anything. It prints both ADF specifications and both KPSS types; choosing which rows to believe is your job, and the plot is what tells you.
four_sources <- function(y, label) {
y <- ts(as.numeric(y))
T_n <- length(y)
print(
ggplot(data.frame(t = seq_len(T_n), y = as.numeric(y)), aes(t, y)) +
geom_line(colour = "steelblue4", linewidth = 0.4) +
labs(title = paste0(label, ": the series"), x = "t",
y = expression(y[t]))
)
rho <- stats::acf(y, lag.max = 25, plot = FALSE)$acf[-1]
print(
ggplot(data.frame(lag = seq_along(rho), rho = as.numeric(rho)), aes(lag)) +
geom_hline(yintercept = 0, colour = "grey40", linewidth = 0.3) +
geom_hline(yintercept = c(-1.96, 1.96) / sqrt(T_n),
linetype = "dashed", colour = "grey55") +
geom_segment(aes(xend = lag, y = 0, yend = rho),
colour = "steelblue4", linewidth = 1) +
labs(title = paste0(label, ": sample ACF"), x = "lag k",
y = expression(hat(rho)[k]))
)
adf_d <- ur.df(y, type = "drift", lags = 4)
adf_t <- ur.df(y, type = "trend", lags = 4)
kp_mu <- ur.kpss(y, type = "mu")
kp_ta <- ur.kpss(y, type = "tau")
data.frame(
measure = c("ADF tau (drift)", "ADF tau (trend)",
"KPSS (mu)", "KPSS (tau)", "sd(dy)/sd(y)"),
value = round(c(adf_d@teststat[1], adf_t@teststat[1],
kp_mu@teststat, kp_ta@teststat,
sd(diff(y)) / sd(y)), 3),
cv_05 = c(adf_d@cval[1, 2], adf_t@cval[1, 2],
kp_mu@cval[1, "5pct"], kp_ta@cval[1, "5pct"], NA),
reject = c(adf_d@teststat[1] < adf_d@cval[1, 2],
adf_t@teststat[1] < adf_t@cval[1, 2],
kp_mu@teststat > kp_mu@cval[1, "5pct"],
kp_ta@teststat > kp_ta@cval[1, "5pct"], NA)
)
}(The reject column already applies each test’s own direction — left tail for ADF, right tail for KPSS — so you don’t have to re-derive it per row. The lag order is fixed at 4 throughout for comparability. This function also ships as a stamped copy in R/four_sources.R — the canonical course copy lives in the repo’s helpers/ — so later sessions can source() it instead of re-building; the inline build above is the pedagogy.)
8.1 Two labeled warm-ups
These two come with their DGPs printed on the label, because each one teaches the specification decision in a different direction.
Warm-up A: trend-stationary by construction. \(y_t = 10 + 0.08t + u_t\), \(u_t = 0.4u_{t-1} + \epsilon_t\), \(T = 250\).
You know the answer key here — the test is whether the instruments agree with it. Which rows of the evidence table will point at “trend-stationary”? And here’s the sharper one: will the ratio \(\sigma_{\Delta y}/\sigma_y\) land above or below the 0.5 flag, given that this series has no unit root?
set.seed(1234)
n <- 250
uA <- numeric(n)
uA[1] <- 0
for (t in 2:n) uA[t] <- 0.4 * uA[t - 1] + rnorm(1)
warmup_a <- 10 + 0.08 * (1:n) + uA
four_sources(warmup_a, "Warm-up A") measure value cv_05 reject
1 ADF tau (drift) -0.688 -2.880 FALSE
2 ADF tau (trend) -5.962 -3.430 TRUE
3 KPSS (mu) 4.218 0.463 TRUE
4 KPSS (tau) 0.076 0.146 FALSE
5 sd(dy)/sd(y) 0.188 NA NA
The plot shows an unmistakable trend, so the rows that match the plot are the trend-spec ADF and the tau-type KPSS — and they agree beautifully: ADF (trend) rejects the unit root (\(-5.96\) against \(-3.43\)) while KPSS (tau) sees no evidence against trend-stationarity (\(0.076\), under \(0.146\)). The verdict is trend-stationary; the treatment is detrending, not differencing.
Now the two traps the table also documents. The drift-spec rows get it wrong-looking: ADF (drift) can’t reject (\(-0.69\)) and KPSS (mu) rejects loudly (\(4.22\)) — of course they do, because “stationary around a constant” is genuinely false for this series; the specification, not the series, is what failed. And the ratio comes out at \(0.188\) — far below 0.5, on a series with no unit root — because a deterministic trend inflates the levels variance just like high persistence does. The ratio flags; it does not diagnose.
Warm-up B: random walk with drift. \(y_t = 0.05 + y_{t-1} + \epsilon_t\), \(T = 250\).
This one also trends upward in the plot — drift does that. So: does the trend-spec ADF reject here, the way it did for Warm-up A? What’s the one pattern in the table that separates “trends because of drift” from “trends because of \(\delta t\)”?
set.seed(5678)
warmup_b <- cumsum(rnorm(250, mean = 0.05))
four_sources(warmup_b, "Warm-up B") measure value cv_05 reject
1 ADF tau (drift) -1.773 -2.880 FALSE
2 ADF tau (trend) -1.756 -3.430 FALSE
3 KPSS (mu) 0.694 0.463 TRUE
4 KPSS (tau) 0.616 0.146 TRUE
5 sd(dy)/sd(y) 0.205 NA NA
Visually, Warm-up B could be Warm-up A’s sibling — both climb. But the table tells a completely different story: neither ADF specification rejects (\(-1.77\) drift, \(-1.76\) trend), and both KPSS types reject (\(0.694\) and \(0.616\)). Allowing for a deterministic trend didn’t rescue stationarity, because the trend here isn’t deterministic — it’s accumulated drift on a unit root. The verdict is difference-stationary; the treatment is \(\Delta y_t\), which strips the unit root and leaves the drift as a constant.
Side by side, these two warm-ups are the specification lesson in full: a trending plot poses the question, and it is the trend-spec ADF plus its matching KPSS — not the eye — that answers it.
8.2 The four mystery series
Now the label comes off. mystery_series(1) through mystery_series(4) are sealed: deterministic, identical for every student, DGPs hidden below the SPOILERS banner in R/mystery_series.R. Each is one of the module’s suspects — stationary, trend-stationary, or unit root (with or without drift).
The protocol, per series — and the commit step is the entire point:
- Run the evidence chunk.
- Commit a verdict in
LEARNING_LOG.mdbefore doing anything else: (a) stationary around a level, (b) trend-stationary, or (c) unit root — difference it. One sentence on which evidence carried the decision. - Only then, in your console (not by editing this file):
reveal_mystery(id). - Reconcile in the log: verdict vs. truth, and which instrument would have corrected you.
If you are working with an AI tutor, it has been asked to collect your lettered verdict once and then reveal — it will not withhold the answer, and it will not volunteer it before you’ve attempted the series either.
four_sources(mystery_series(1), "Mystery 1") measure value cv_05 reject
1 ADF tau (drift) -8.315 -2.870 TRUE
2 ADF tau (trend) -8.311 -3.420 TRUE
3 KPSS (mu) 0.097 0.463 FALSE
4 KPSS (tau) 0.069 0.146 FALSE
5 sd(dy)/sd(y) 1.122 NA NA
reveal_mystery(1) # console, after your committed verdictfour_sources(mystery_series(2), "Mystery 2") measure value cv_05 reject
1 ADF tau (drift) -0.563 -2.870 FALSE
2 ADF tau (trend) -1.945 -3.420 FALSE
3 KPSS (mu) 4.562 0.463 TRUE
4 KPSS (tau) 0.900 0.146 TRUE
5 sd(dy)/sd(y) 0.108 NA NA
reveal_mystery(2) # console, after your committed verdictfour_sources(mystery_series(3), "Mystery 3") measure value cv_05 reject
1 ADF tau (drift) -1.612 -2.870 FALSE
2 ADF tau (trend) -6.438 -3.420 TRUE
3 KPSS (mu) 4.720 0.463 TRUE
4 KPSS (tau) 0.072 0.146 FALSE
5 sd(dy)/sd(y) 0.373 NA NA
reveal_mystery(3) # console, after your committed verdictfour_sources(mystery_series(4), "Mystery 4") measure value cv_05 reject
1 ADF tau (drift) -2.274 -2.880 FALSE
2 ADF tau (trend) -1.870 -3.430 FALSE
3 KPSS (mu) 1.441 0.463 TRUE
4 KPSS (tau) 0.546 0.146 TRUE
5 sd(dy)/sd(y) 0.252 NA NA
reveal_mystery(4) # console, after your committed verdictNo commentary on these four here — commentary is spoilers. Two process notes only. First, remember from Section 1 that mystery_series() re-seeds the session: if you add your own simulations below a mystery chunk, call set.seed() again. Second, expect at least one of the four to be genuinely hard — if every verdict felt easy, compare your evidence tables against Section 6’s power numbers and ask whether the confidence was earned.
8.3 Real data: the unemployment rate
Module 1 ended on the cliffhanger: UNRATE has \(\hat\phi \approx 0.97\), and no plot or ACF could say whether that is “very persistent” or “unit root.” You now own four instruments the plot doesn’t have. Time to use them.
unrate <- unrate_raw
names(unrate)[names(unrate) == "observation_date"] <- "date"
unrate$date <- as.Date(unrate$date)
# The cache has exactly one missing month (2025-10, a federal data-release gap).
# A monthly ts index cannot carry a hole, so we keep the complete run up to it.
# This is a documented feature of the file, not damage — see data/README.md.
first_na <- which(is.na(unrate$UNRATE))
if (length(first_na) > 0) unrate <- unrate[seq_len(min(first_na) - 1), ]
unrate <- unrate[order(unrate$date), ]
unrate_ts <- ts(unrate$UNRATE,
start = c(as.integer(format(min(unrate$date), "%Y")),
as.integer(format(min(unrate$date), "%m"))),
frequency = 12)
data.frame(n_obs = length(unrate_ts),
first = format(min(unrate$date)),
last = format(max(unrate$date))) n_obs first last
1 933 1948-01-01 2025-09-01
933 monthly observations, \(\hat\phi \approx 0.97\), no visible trend. Before the tests run: which cell of the ADF/KPSS grid does unemployment land in — (a) strong stationarity, (b) strong unit root, (c) the contradiction row, or (d) inconclusive? This one is worth committing to writing.
The series has a level but no trend, so the specification decision is easy for once: ADF with drift, KPSS with mu, and BIC choosing the augmentation lags on a real series’s messy dynamics.
print(
ggplot(unrate, aes(date, UNRATE)) +
geom_line(colour = "steelblue4", linewidth = 0.4) +
geom_hline(yintercept = mean(unrate$UNRATE), linetype = "dashed",
colour = "grey45") +
labs(title = "U.S. civilian unemployment rate, 1948-2025",
subtitle = "Monthly, seasonally adjusted. Dashed: full-sample mean",
x = NULL, y = "percent")
)adf_un <- ur.df(unrate_ts, type = "drift", lags = 12, selectlags = "BIC")
kpss_un <- ur.kpss(unrate_ts, type = "mu")
data.frame(
measure = c("ADF tau (drift, BIC lags)", "KPSS (mu)", "sd(dy)/sd(y)"),
value = round(c(adf_un@teststat[1], kpss_un@teststat,
sd(diff(unrate_ts)) / sd(unrate_ts)), 3),
cv_05 = c(adf_un@cval[1, 2], kpss_un@cval[1, "5pct"], NA),
reject = c(adf_un@teststat[1] < adf_un@cval[1, 2],
kpss_un@teststat > kpss_un@cval[1, "5pct"], NA)
) measure value cv_05 reject
1 ADF tau (drift, BIC lags) -3.903 -2.860 TRUE
2 KPSS (mu) 0.970 0.463 TRUE
3 sd(dy)/sd(y) 0.243 NA NA
Line up all four sources:
- Plot / ACF (Module 1’s evidence): high persistence, no trend — inconclusive between \(\phi = 0.97\) and \(\phi = 1\).
- ADF: \(\tau \approx -3.90\), well past \(-2.86\) — rejects the unit root.
- KPSS: \(0.970\), past \(0.463\) — rejects stationarity.
- Ratio: \(0.243\), far below \(0.5\) — the high-persistence flag is up.
Both formal tests reject: this is the contradiction row of the grid. The tests are not malfunctioning — they are jointly reporting that neither “clean unit root” nor “clean stationary AR” fits this series comfortably. The usual suspects for that row: a structural break, an outlier episode (2020 is sitting right there in the plot), or some other misspecification the workflow hasn’t modeled. The workflow establishes the conflict; it does not identify the cause — Exploration 5 chases the leading suspect.
So the verdict is a judgment call, made under the cost asymmetry: spurious regression (treating a unit root as stationary) is confidently, invisibly wrong; over-differencing is visible in later diagnostics via Section 7’s \(\hat\theta \approx -1\) fingerprint. The course’s working decision: treat UNRATE as \(I(1)\) and model \(\Delta y_t\) — stated plainly as a defensible choice under conflicting evidence, not as a fact the tests proved. Economists genuinely disagree about this series (a bounded rate arguably cannot have a literal unit root — see the notes’ Deeper Dive); what they agree on is that 78 years of monthly data cannot settle \(\phi = 0.998\) versus \(1.000\).
One step remains, and it is the second half of the order-of-integration workflow. “Treat as \(I(1)\)” is a claim that one difference is enough — so test it, by hand: difference the series and run the same instruments on \(\Delta y_t\). (This manual two-step — test the levels, difference, re-test — is the course’s \(I(0)\)-vs-\(I(1)\) procedure. If \(I(2)\) were plausible you would instead test downward from the highest order, Dickey-Pantula style, but for an unemployment rate it is not.)
d_unrate <- diff(unrate_ts)
adf_d <- ur.df(d_unrate, type = "drift", lags = 12, selectlags = "BIC")
kpss_d <- ur.kpss(d_unrate, type = "mu")
data.frame(
measure = c("ADF tau (drift, BIC lags)", "KPSS (mu)", "sd(d2y)/sd(dy)"),
value = round(c(adf_d@teststat[1], kpss_d@teststat,
sd(diff(d_unrate)) / sd(d_unrate)), 3),
cv_05 = c(adf_d@cval[1, 2], kpss_d@cval[1, "5pct"], NA),
reject = c(adf_d@teststat[1] < adf_d@cval[1, 2],
kpss_d@teststat > kpss_d@cval[1, "5pct"], NA)
) measure value cv_05 reject
1 ADF tau (drift, BIC lags) -21.953 -2.860 TRUE
2 KPSS (mu) 0.038 0.463 FALSE
3 sd(d2y)/sd(dy) 1.389 NA NA
The differenced series lands on the grid’s good diagonal with room to spare: ADF rejects emphatically (\(\tau \approx -21.95\)), KPSS goes quiet (\(0.038\), nowhere near \(0.463\)), and the ratio jumps to \(1.39\) — right at the \(\sqrt{2} \approx 1.41\) the rule of thumb predicts for a series with little persistence left. One difference was enough: whatever ambiguity the levels carry, \(\Delta y_t\) is comfortably stationary, and \(I(1)\) is a coherent working label.
So what: you just did the whole module on real data — chose specifications from the plot, ran opposite-null tests, hit the contradiction row, made a defensible call anyway, and then tested the call by re-running the workflow on the differenced series. That closed loop, argued honestly, is the skill the problem sets grade.
9. Explorations
Nothing in this section is required, and none of it appears on problem sets or exams — assessments draw only on the Core tier of the module notes. These are here because the interesting questions usually start where the syllabus stops.
If one of these grabs you, chase it. If a different question grabs you, chase that one instead — that is a better outcome than any of the six below. Your AI tutor’s job in this section is to help you turn a hunch into an experiment you can actually run, then connect what you find back to a Core concept.
1. The power frontier. (guided)
Section 6 fixed \(\phi = 0.95\) and varied \(T\). Complete the map: for each \(\phi \in \{0.80, 0.90, 0.95\}\), find the approximate \(T\) at which the ADF test reaches 80% power. Here is a starter for one value of \(\phi\) — extend it, then plot the three curves together:
set.seed(6376)
sizes <- c(50, 100, 200, 400)
frontier_row <- sapply(sizes, function(T_p) {
mean(replicate(200, {
y <- ar1_simulator(n = T_p, alpha = 0, phi = 0.90, sigma = 1, burn_in = 200)
ur.df(y, type = "drift", lags = 4)@teststat[1] < mackinnon_cv(T_p, 0.05)
}))
})
names(frontier_row) <- paste0("T=", sizes)
round(frontier_row, 3) T=50 T=100 T=200 T=400
0.090 0.220 0.675 1.000
Before extending it, predict the shape of the answer: does the 80%-power sample size grow linearly as \(\phi \to 1\), or much faster? Your three numbers settle it.
2. GDP: trend-stationary or difference-stationary? (guided)
This is one of the longest-running empirical fights in macroeconomics — whether U.S. output returns to a deterministic trend after recessions or wanders permanently (Nelson and Plosser started it in 1982). The pack ships the data:
gdp_raw <- read.csv("data/GDPC1.csv", stringsAsFactors = FALSE)
log_gdp <- ts(log(gdp_raw$GDPC1), start = c(1947, 1), frequency = 4)
ggplot(data.frame(t = as.numeric(time(log_gdp)), y = as.numeric(log_gdp)),
aes(t, y)) +
geom_line(colour = "steelblue4", linewidth = 0.4) +
labs(title = "Log U.S. real GDP, quarterly, 1947-2026",
subtitle = "Trend-stationary or difference-stationary? 318 quarters of ambiguity",
x = NULL, y = "log(GDPC1)")Run the full workflow on log_gdp: both ADF specifications, KPSS with type = "tau", the ratio. Take a verdict and defend it in a paragraph — then ask your tutor to argue the other side, because for this series there genuinely is one.
3. The wrong-specification tax.
Section 4 showed one specification flip. Measure the phenomenon properly: simulate many trend-stationary series and record how often the drift-spec ADF wrongly fails to reject; then simulate driftless stationary AR(1)s and measure the power cost of running the trend spec when drift would have sufficed. Which direction of specification error is more expensive?
4. Mapping the inconclusive zone.
As \(\phi \to 1\) from below, the ADF/KPSS grid’s cells shift population. For \(\phi \in \{0.90, 0.95, 0.98, 0.995\}\) at \(T = 200\), simulate a few hundred series each and tabulate which of the four cells the pair lands in. Where does the “strong stationarity” diagonal give way to the inconclusive and contradiction cells? You are drawing the empirical boundary of what testing can know.
5. The break masquerade.
UNRATE landed in the contradiction row, and Section 8 named a suspect this module doesn’t formally model: a structural break. Build the lineup: simulate a stationary AR(1) (\(\phi = 0.5\), \(T = 300\)) whose mean jumps from 0 to 3 at \(t = 150\), and run the four sources on it. Does a mere level shift push the pair of tests into the same contradiction row as UNRATE? (Structural breaks are off this course’s syllabus by design — this exploration is the one handshake with them, and your tutor can take you further if the question bites.)
6. Your own FRED series. (open)
Pick a series you personally care about — inflation, house prices, your industry’s employment. Download its CSV by hand from the FRED website (using data/UNRATE.csv as the format template), or pull it live with R/fetch_fred.R if you have a free API key (the key lives in your environment, never in a file). Run the complete workflow: plot, ACF, specification choice, ADF, matching KPSS, ratio, verdict — written up in 2–3 paragraphs with the cost asymmetry doing the tie-breaking. This is exactly the shape of the problem-set task, practiced on stakes you chose.
# Optional. Requires the fredr package and a FRED_KEY in your environment.
# Not required for anything in this lab.
source("R/fetch_fred.R")
cpi <- fetch_fred("CPIAUCSL", start_date = "1960-01-01", frequency = 12)10. Checks
R/checks.R is the referee. Seven assertions about things this module claims are true — the DF null distribution, the ADF and KPSS verdicts on clean cases, the power problem, the rule-of-thumb algebra, the over-differencing boundary, and the integrity of the data and the sealed series. It does not know what you concluded and it cannot grade you — it just refuses to let an incorrect claim about the code stand, whether that claim came from you, from your AI tutor, or from a typo.
If a check fails, CHECKS.md explains what that check asserts and what a failure usually means.
source("R/checks.R")Module 2 checks
---------------
PASS DF null MC 5% quantile matches mackinnon_cv(200) [empirical q05 = -2.883, MacKinnon CV = -2.876, tol 0.15]
PASS ADF rejects on a long stationary AR(1), fails to reject on a random walk [stationary tau = -10.71 vs -2.86; random walk tau = -2.00 vs -2.87]
PASS KPSS fails to reject the stationary series and rejects the random walk [stationary stat = 0.171, random walk stat = 3.578, 5% CV = 0.463]
PASS ADF power at phi = 0.95, T = 100 is low (rejection rate < 0.25) [rejection rate = 0.103 over 300 replications]
PASS sd(dy)/sd(y) for phi = 0.9 matches sqrt(2(1 - phi)) [sample ratio = 0.4662, theory 0.4472, tol 0.03]
PASS MA(1) fit to a differenced trend-stationary series pins theta near -1 [theta_hat = -1.0000, boundary -1, threshold -0.9]
PASS UNRATE.csv, GDPC1.csv, and the four mystery series are intact [UNRATE ok, GDPC1 ok, mystery series ok]
---------------
7/7 checks passed.
Where you are now
You asked one question — is \(\phi_1 = 1\)? — and built the machinery to answer it honestly:
- The DF rearrangement turns the boundary question into \(H_0\!: \gamma = 0\) in a regression — but under that null the \(t\)-ratio follows the left-shifted DF distribution, and using the normal’s \(-1.645\) rejects a true null about 47% of the time.
- MacKinnon critical values are the empirical quantiles of that distribution; the specification (none / drift / trend) is half the test, and choosing it from the plot can flip a verdict.
- KPSS reverses the null; ADF and KPSS agreeing is strong evidence, both rejecting is a contradiction flag, both silent is an honest “don’t know.”
- Power near the boundary is brutal: \(\phi = 0.95\) at \(T = 100\) is caught about 10% of the time. Hence: four sources, never one.
- \(\sigma_{\Delta y}/\sigma_y = \sqrt{2(1-\phi)}\) for a stationary AR(1); below 0.5 is a persistence flag, not a unit-root proof.
- Over-differencing manufactures a non-invertible \((1-L)\) MA factor — \(\hat\theta\) pinned at \(-1\) is its fingerprint — and that word, invertibility, is the boundary Module 3 opens with.
- UNRATE, tested at last, lands in the contradiction row — and the verdict (treat as \(I(1)\)) is a defended judgment under the cost asymmetry, not a measurement.
Module 3 turns on the next sliders: higher-order AR memory and MA terms, and the correlogram fingerprints that identify them — the complementary tool the power problem demands.
Before you close this file: open LEARNING_LOG.md and check that every mystery series you attempted has both halves written down — what you called it, and what it was. The gap between those two columns, in your own words, is the most valuable thing this lab produces.