Module 1 Lab: Memory, Noise, and Why Time Order Matters
Econ 6376 — Applied Time Series Econometrics · AI Learning Pack
How to use this lab
Everything here runs as shipped. No API key, no internet, no blanks to fill in. That is deliberate: the point of this lab is not to get the code working, it is to find out whether your intuition about time series is any good.
So there is one rule that makes the whole thing worth doing:
Write your prediction down before you run the chunk.
Not in your head — in your head you will “sort of” have predicted whatever happens. Write it in LEARNING_LOG.md, or in a comment, or on paper. Being wrong on paper is the single most useful thing that can happen to you today. It is how you find out which parts of your mental model are load-bearing and which parts are decoration.
If you are working with an AI tutor, it will ask you for that prediction before it runs anything meaningful, and it will run the code either way. It is not gatekeeping you — it is trying to make sure you get the “huh, I didn’t expect that” moment, because that is the one you remember in November.
Where you are on the mixing board: the master equation is
Module 1 turns on exactly two sliders, \(\alpha\) and \(\phi_1\), and leaves every other one at zero:
\[
y_t = \alpha + \phi y_{t-1} + \epsilon_t .
\]
That is the whole model for this lab. Everything you are about to see — persistence, the ACF, nonstationarity, spurious regression — comes out of those two knobs.
1. Setup
Two ingredients: the simulator, and the cached unemployment data.
library(ggplot2)# ar1_simulator() is a stamped copy of the canonical course helper.source("R/ar1_simulator.R")# Cached FRED pull, so this lab renders with no key and no internet.# See data/README.md for provenance and the one known gap.unrate_raw <-read.csv("data/UNRATE.csv", stringsAsFactors =FALSE)theme_set(theme_bw(base_size =13))str(unrate_raw)
Same thing. rnorm(1, mu, sigma) is \(\epsilon_t\), the innovation — the genuinely new information arriving at time \(t\). Note that it is drawn with mean zero. If you gave the innovations a nonzero mean it would silently fold into the intercept and you would no longer be able to interpret \(\alpha\).
One habit worth building now: call this function with named arguments. The course has more than one version of ar1_simulator() floating around (the notes build one live, helpers/ keeps the canonical one), and their argument orders differ. Named arguments make you immune to that.
A note on burn_in
burn_in simulates extra observations at the front and throws them away, so the series you keep has forgotten the arbitrary starting value init. For a stationary AR(1) the influence of init after \(b\) periods is proportional to \(\phi^b\), so it dies fast when \(\phi\) is small and slowly when \(\phi\) is near 1.
For \(\phi = 1\) no burn-in helps, because there is no stationary distribution to settle into. Discarding the first stretch of a random walk just changes which part of the same wandering path you are looking at.
So what: the model is now code, and the code is now running. Everything below is turning the two knobs.
2. Memory: simulating AR(1) across the \(\phi\) regimes
We simulate five processes that differ only in \(\phi\). Same seed each time, so the innovation sequence is the same in every panel — any difference you see is caused by \(\phi\) and nothing else.
Predict first
Before you run this: sketch (or describe in one sentence each) what the five panels will look like for \(\phi = 0, 0.5, 0.9, 0.99, 1\). In particular — which of them will hover around a fixed level, and which will wander off?
phis <-c(0, 0.5, 0.9, 0.99, 1.0)gallery <-do.call(rbind, lapply(phis, function(p) {set.seed(6376) # same shocks in every panel y <-ar1_simulator(n =400, alpha =0, phi = p, sigma =1, burn_in =500)data.frame(t =seq_along(y),y =as.numeric(y),phi =factor(paste0("phi == ", p), levels =paste0("phi == ", phis)) )}))ggplot(gallery, aes(t, y)) +geom_hline(yintercept =0, colour ="grey60", linewidth =0.3) +geom_line(linewidth =0.35, colour ="steelblue4") +facet_wrap(~ phi, ncol =2, scales ="free_y", labeller = label_parsed) +labs(title ="One AR(1) DGP, five settings of the memory knob",subtitle ="Identical innovations in every panel; only phi changes",x ="t", y =expression(y[t]))
Look at the vertical scales — they are free, and that is the first clue.
memory_summary <-do.call(rbind, lapply(phis, function(p) {set.seed(6376) y <-as.numeric(ar1_simulator(n =400, alpha =0, phi = p, sigma =1,burn_in =500)) half <-floor(length(y) /2) # so this still splits in two if you change ndata.frame(phi = p,sd =sd(y),mean_1st_half =mean(y[1:half]),mean_2nd_half =mean(y[(half +1):length(y)]),theory_sd =if (abs(p) <1) sqrt(1/ (1- p^2)) elseNA_real_ )}))round(memory_summary, 3)
Three things to notice, and the third is the interesting one.
The spread grows as \(\phi\) rises. At \(\phi = 1\) the theoretical column returns NA — not because R gave up, but because the quantity does not exist. There is no stationary variance for a random walk.
The two half-sample means tell the same story from a different angle. For \(\phi = 0\) and \(\phi = 0.5\) both halves sit near zero: the process keeps returning to its mean. As \(\phi\) climbs toward 1, the halves drift apart, because “the mean” is losing its grip on the series.
Now the third thing. At \(\phi = 0, 0.5, 0.9\) the realized sd sits close to the theoretical \(\sqrt{\sigma^2/(1-\phi^2)}\) — but at \(\phi = 0.99\) it is about 4.0 against a theoretical 7.1, barely half. The theory is not wrong. The sample is too short to have seen the process’s full range. At \(\phi = 0.99\) a shock takes roughly 69 periods to decay by half, so 500 burn-in periods plus 400 observations is simply not enough wandering to explore the whole stationary distribution. A near-unit-root process looks less variable than it truly is when you only watch it briefly.
Sit with that for a second, because it is a preview of every hard problem in this course: near the unit-root boundary, finite samples systematically understate persistence. Exploration 1 and Exploration 4 both chase this.
Modify and re-run
Change n = 400 to n = 4000 and re-run both chunks. Before you do: does the \(\phi = 0.99\) panel start looking more stationary, or less? Why? (Hint: how many periods does it take for \(0.99^b\) to get small?)
So what:\(\phi\) is not a cosmetic parameter. It controls how long a shock survives, how wide the series swings, and — at the boundary — whether the series has a long-run mean at all.
3. Reading the ACF
The headline result of Module 1 is that for a stationary AR(1),
\[
\rho_k = \phi^k .
\]
The autocorrelation at lag \(k\) is just \(\phi\) raised to the \(k\). That is a strong claim, and it is falsifiable, so let’s try to falsify it.
A correlogram is not magic. It is a bar chart of \(\hat\rho_k\) against \(k\), so we will build it by hand — computing it yourself once is worth more than calling ggAcf() ten times.
Predict first
For \(\phi = 0.9\): roughly how many lags before the ACF drops below 0.5? Write down a number. (You can compute it exactly — \(0.9^k < 0.5\) — but guess first, then check.)
acf_frame <-function(y, phi, lag_max =20) { a <- stats::acf(y, lag.max = lag_max, plot =FALSE)$acf[-1] # drop lag 0data.frame(lag =seq_len(lag_max),sample =as.numeric(a),theory = phi^seq_len(lag_max),phi =factor(paste0("phi == ", phi)) )}T_n <-1000# change this one number and everything below followsacf_data <-do.call(rbind, lapply(c(0.5, 0.9), function(p) {set.seed(6376) y <-ar1_simulator(n = T_n, alpha =0, phi = p, sigma =1, burn_in =500)acf_frame(y, phi = p)}))band <-1.96/sqrt(T_n)ggplot(acf_data, aes(lag)) +geom_hline(yintercept =c(-band, band), linetype ="dashed",colour ="grey55") +geom_hline(yintercept =0, colour ="grey40", linewidth =0.3) +geom_segment(aes(xend = lag, y =0, yend = sample),colour ="steelblue4", linewidth =1.1) +geom_line(aes(y = theory), colour ="firebrick", linewidth =0.7) +geom_point(aes(y = theory), colour ="firebrick", size =1.6) +facet_wrap(~ phi, labeller = label_parsed) +labs(title ="Sample ACF (bars) against the theoretical ACF (red)",subtitle =bquote("Theory:"~ rho[k] == phi^k *"; T ="~ .(T_n)),x ="lag k", y =expression(rho[k]))
The red line is theory; the bars are what one sample of 1000 observations actually produced. They are close but not identical, and that gap is the whole reason we bother with inference. \(\hat\rho_k\) is an estimate.
The dashed bands sit at \(\pm 1.96/\sqrt{T}\). They are approximate, individual-lag reference lines drawn under an iid white-noise null — not a confidence band for a persistent process, and not a multiple-comparison correction. With 20 lags plotted, one bar poking outside is not news.
Modify and re-run
Change T_n <- 1000 to T_n <- 100 and re-run (the significance bands rescale with it automatically). Predict first: does the theoretical red line move? Does the sample ACF get noticeably worse? Which lags degrade first — short ones or long ones?
So what: the ACF is the fingerprint. Geometric decay says “stationary AR.” Decay so slow it looks flat says “highly persistent, and possibly not stationary at all.” Which brings us to the interesting part.
4. The assumption-break: push \(\phi \to 1\) and watch stationarity die
Every simulation so far has confirmed the theory. That is only half an exercise. Now we break it.
Weak stationarity asks for three things: a constant mean, a constant finite variance, and an autocovariance that depends only on the lag \(k\) and not on calendar time \(t\). At \(\phi = 1\) the model becomes
so every shock that has ever happened is still in the level, at full weight, forever. The claim is that \(\operatorname{Var}(y_t) = t\sigma^2\).
Here is the crucial move: that is a statement about the process, not about one path. You cannot see it in a single series. So we simulate the same random walk 500 times and measure the spread across paths at each date.
Predict first
Sketch what “500 random walks on one plot” looks like. Does the cloud of paths get wider, narrower, or stay the same width as \(t\) increases? And if it widens, does it widen linearly, or like \(\sqrt{t}\)?
set.seed(6376)n_paths <-500# burn_in = 0 keeps y_0 = 0, so path element j is the sum of j innovations.rw_paths <-replicate( n_paths,as.numeric(ar1_simulator(n =201, alpha =0, phi =1, init =0,sigma =1, burn_in =0)))dim(rw_paths) # 200 dates x 500 paths
[1] 200 500
show_n <-60fan <-data.frame(t =rep(seq_len(nrow(rw_paths)), times = show_n),y =as.numeric(rw_paths[, seq_len(show_n)]),path =rep(seq_len(show_n), each =nrow(rw_paths)))envelope <-data.frame(t =seq_len(nrow(rw_paths)),sd =apply(rw_paths, 1, sd))ggplot() +geom_line(data = fan, aes(t, y, group = path),alpha =0.18, linewidth =0.3, colour ="steelblue4") +geom_line(data = envelope, aes(t, 2* sd), colour ="firebrick",linewidth =0.9) +geom_line(data = envelope, aes(t, -2* sd), colour ="firebrick",linewidth =0.9) +geom_line(data = envelope, aes(t, 2*sqrt(t)), colour ="black",linetype ="dashed", linewidth =0.7) +geom_line(data = envelope, aes(t, -2*sqrt(t)), colour ="black",linetype ="dashed", linewidth =0.7) +labs(title ="500 random walks, and the fan they open into",subtitle =paste("Red: +/- 2 simulated SD at each date.","Dashed: +/- 2 sqrt(t), the theoretical spread"),x ="t", y =expression(y[t]))
The simulated variance tracks \(t\). Not “approximately constant with some noise” — it grows without bound, in a straight line, exactly as \(\operatorname{Var}(y_t)
= t\sigma^2\) predicts. Stationarity condition 2 is dead, and condition 1 goes with it: there is no fixed level for the process to return to, so asking “what is the mean of a random walk?” is asking about something that isn’t there.
Now compare the two cases that look alike and are not:
compare <-do.call(rbind, lapply(c(0.99, 1.0), function(p) {set.seed(6376) y <-as.numeric(ar1_simulator(n =400, alpha =0, phi = p, sigma =1,burn_in =500))data.frame(t =seq_along(y), y = y,phi =factor(paste0("phi == ", p)))}))ggplot(compare, aes(t, y, colour = phi)) +geom_hline(yintercept =0, colour ="grey60") +geom_line(linewidth =0.45) +scale_colour_manual(values =c("steelblue4", "firebrick"),labels =c(expression(phi ==0.99), expression(phi ==1)),name =NULL) +labs(title ="Stationary or not? You cannot tell by looking",subtitle ="T = 400. One of these has a long-run mean. One does not.",x ="t", y =expression(y[t]))
They look like siblings. But at \(\phi = 0.99\) a shock has decayed to \(0.99^{100} \approx 0.37\) after a century of months and the process does have a stationary variance, \(\sigma^2/(1-\phi^2) \approx 50\). At \(\phi = 1\) the shock is still at full strength and there is no such variance.
This is the honest, uncomfortable core of Module 1: a plot cannot settle this question, and neither can an ACF. Module 2 exists because we need a test.
So what: “highly persistent” and “nonstationary” are different claims that produce nearly identical pictures in a finite sample. Anyone who tells you they can eyeball a unit root is telling you about their confidence, not about the data.
5. Spurious regression: the high-\(R^2\) trap
Here is why any of this matters for applied work.
Take two random walks that have nothing to do with each other — separate, independent innovation sequences, no relationship of any kind in the population. Regress one on the other with lm().
Predict first
The true slope is zero. If you run this experiment 1000 times at the 5% level, how often should a well-behaved test reject? How often do you think it will? Commit to a number before scrolling.
Call:
lm(formula = y ~ x)
Residuals:
Min 1Q Median 3Q Max
-4.4415 -1.4122 0.1288 1.2440 3.8604
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 3.97691 0.30393 13.09 <2e-16 ***
x 0.59072 0.04288 13.78 <2e-16 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 1.809 on 98 degrees of freedom
Multiple R-squared: 0.6594, Adjusted R-squared: 0.656
F-statistic: 189.8 on 1 and 98 DF, p-value: < 2.2e-16
one_pair <-data.frame(t =seq_len(n_obs), x = x, y = y)ggplot(one_pair, aes(t)) +geom_line(aes(y = x, colour ="x"), linewidth =0.5) +geom_line(aes(y = y, colour ="y"), linewidth =0.5) +scale_colour_manual(values =c(x ="steelblue4", y ="firebrick"),name =NULL) +labs(title ="Two completely unrelated random walks",subtitle =sprintf("R-squared = %.3f, p-value on x = %.4g",summary(fit)$r.squared,coef(summary(fit))["x", "Pr(>|t|)"]),x ="t", y =NULL)
One draw proves nothing — maybe we got unlucky. So do it a thousand times.
set.seed(42)n_sims <-1000reject_rw <-mean(replicate(n_sims, { x <-cumsum(rnorm(n_obs)) y <-cumsum(rnorm(n_obs))coef(summary(lm(y ~ x)))["x", "Pr(>|t|)"] <0.05}))# Contrast: two independent STATIONARY AR(1)s, same everything else.set.seed(42)reject_ar <-mean(replicate(n_sims, { x <-as.numeric(ar1_simulator(n = n_obs, phi =0.5, burn_in =200)) y <-as.numeric(ar1_simulator(n = n_obs, phi =0.5, burn_in =200))coef(summary(lm(y ~ x)))["x", "Pr(>|t|)"] <0.05}))data.frame(design =c("two independent random walks", "two independent AR(1), phi = 0.5"),nominal_level =0.05,rejection_rate =c(reject_rw, reject_ar))
design nominal_level rejection_rate
1 two independent random walks 0.05 0.769
2 two independent AR(1), phi = 0.5 0.05 0.117
Roughly three quarters of the time, a 5% test announces a significant relationship between two series that have no relationship whatsoever.
Be precise about why, because the usual one-liner (“they both trend”) is too weak. The regressor and the disturbance are both integrated, their sample cross-products accumulate shocks instead of averaging them away, and the normalization that makes the conventional \(t\) statistic converge to a standard normal simply does not apply here. The \(t\) statistic diverges. That is why the problem gets worse with more data, not better — the one situation where your instinct to collect a longer sample actively hurts you.
The stationary contrast is the useful control, and it separates two lessons that usually get mashed into one. Dropping from a unit root to \(\phi = 0.5\) takes the rejection rate from about 0.77 down to about 0.12 — most of the damage really was the unit root, not “dependence” in general. But 0.12 is still more than twice the 5% you asked for, because lm() assumes the disturbance is serially uncorrelated and it isn’t. So: a unit root wrecks the asymptotics outright, and ordinary serial correlation quietly miscalibrates your standard errors. Different problems, different fixes, both waiting for you later in the course.
Modify and re-run
Change n_obs to 50, then 200, then 500, and re-run the Monte Carlo. Predict the direction first. A correctly calibrated test’s rejection rate should settle toward 0.05 as \(T\) grows — write down what you expect to see instead, and why.
Do not over-learn this
The lesson is not “difference everything nonstationary before regressing.” If a linear combination of two nonstationary series is stationary, they are cointegrated and the levels regression describes a genuine long-run equilibrium — differencing it away would destroy the very thing you wanted. Diagnose first, transform second. Cointegration arrives later in the course.
So what: a high \(R^2\) and a tiny \(p\)-value are not evidence of a relationship between persistent series. They are what you get by default from persistent series. This is the single best reason to care about stationarity testing.
6. Real data: where does unemployment sit on the memory spectrum?
Simulation is where you learn the rules. Real data is where you find out how much the rules help.
UNRATE is the monthly U.S. civilian unemployment rate, seasonally adjusted, from January 1948. We load the cached copy so this runs for everyone.
unrate <- unrate_rawnames(unrate)[names(unrate) =="observation_date"] <-"date"unrate$date <-as.Date(unrate$date)# The cache has exactly one missing month (2025-10, a federal data-release gap).# A monthly ts index cannot carry a hole, so we keep the complete run up to it.# This is a documented feature of the file, not damage — see data/README.md.first_na <-which(is.na(unrate$UNRATE))if (length(first_na) >0) unrate <- unrate[seq_len(min(first_na) -1), ]unrate <- unrate[order(unrate$date), ]unrate_ts <-ts(unrate$UNRATE,start =c(as.integer(format(min(unrate$date), "%Y")),as.integer(format(min(unrate$date), "%m"))),frequency =12)data.frame(n_obs =length(unrate_ts),first =format(min(unrate$date)),last =format(max(unrate$date)))
n_obs first last
1 933 1948-01-01 2025-09-01
Predict first
You know what unemployment looks like. Before plotting: is its ACF going to look more like the \(\phi = 0.5\) panel or the \(\phi = 0.99\) panel from section 2? And if you had to name a single number \(\hat\phi\) for it, what would you guess?
\(\hat\phi \approx 0.97\). Read that against section 2: unemployment sits far to the right on the memory spectrum, in the neighborhood where \(\phi = 0.99\) and \(\phi = 1\) were visually indistinguishable. The implied long-run mean \(\hat\alpha/(1-\hat\phi)\) lands almost exactly on the sample mean, which is a nice consistency check on the arithmetic — though not evidence the model is right.
Now the honest part. Does \(\hat\rho_k = \hat\phi^{\,k}\) actually describe this series?
k <-1:48compare_acf <-data.frame(lag =rep(k, 2),rho =c(as.numeric(unrate_acf), phi_hat^k),src =rep(c("UNRATE sample ACF", "AR(1) theory: phi-hat^k"), each =length(k)))ggplot(compare_acf, aes(lag, rho, colour = src)) +geom_hline(yintercept =0, colour ="grey60") +geom_line(linewidth =0.8) +scale_colour_manual(values =c("firebrick", "steelblue4"), name =NULL) +labs(title ="Does an AR(1) actually describe unemployment?",subtitle =sprintf("Fitted phi-hat = %.3f", phi_hat),x ="lag k (months)", y =expression(rho[k]))
The two curves agree beautifully at short lags and then separate: the real ACF falls off faster than the AR(1) says it should. An AR(1) captures the month-to-month persistence and then overstates how long that persistence lasts.
That is not a failure of the lab — it is the correct finding, and it is the handoff to the rest of the course. Real series are richer than one lag.
Questions worth sitting with:
Does unemployment look like it has a constant mean across all 77 years?
Recessions appear as sharp spikes followed by long declines. Is that symmetry consistent with a linear AR(1), which treats up-moves and down-moves identically?
\(\hat\phi = 0.97\) is close to 1. Is it close enough to 1 that you would call this series nonstationary? What evidence would let you decide?
You cannot answer question 3 with the tools in this module. That is Module 2.
So what: you can now take an unfamiliar series, plot it, read its ACF, fit an AR(1), and say where it sits on the memory spectrum — and, just as importantly, say what that evidence does not settle.
7. Explorations
Optional, and outside the assessable surface
Nothing in this section is required, and none of it appears on problem sets or exams — assessments draw only on the Core tier of the module notes. These are here because the interesting questions usually start where the syllabus stops.
If one of these grabs you, chase it. If a different question grabs you, chase that one instead — that is a better outcome than any of the six below. Your AI tutor’s job in this section is to help you turn a hunch into an experiment you can actually run, then connect what you find back to a Core concept.
1. How much data do you need to pin down \(\phi\)? (guided)
We estimated \(\hat\phi\) from 933 months of unemployment. How much would you trust it from 60? Here is a starter you can run and extend:
Notice that \(\hat\phi\) is biased downward in small samples — the mean is below 0.9 and creeps up with \(T\). Why would OLS systematically understate persistence? (This is a real and well-known result, not a bug in the simulation.)
2. Does the spurious-regression problem get better or worse with more data? (guided)
Run the section 5 Monte Carlo at \(T \in \{50, 100, 200, 400, 800\}\) and plot the rejection rate against \(T\). Then articulate why this curve is the opposite shape from every consistency result you learned in your OLS course.
3. Is unemployment more persistent in some eras than others?
Split unrate_ts — pre-1985 versus post-1985, or expansions versus recessions if you can define them — and estimate \(\hat\phi\) separately on each piece. If the two numbers differ a lot, what does that do to the claim that this series is covariance stationary? (Careful: short subsamples make \(\hat\phi\) noisy. See exploration 1 before you get excited about a difference.)
4. How long does burn-in actually need to be?
Set init = 50 with \(\phi = 0.99\) and vary burn_in from 0 to 5000. How many periods until the starting value is genuinely forgotten? Now try the same thing with \(\phi = 1\) and explain why the question has no answer there.
5. Pick a FRED series you personally care about.
Inflation, housing starts, your industry’s employment, the price of eggs. Plot it, read its ACF, fit an AR(1), and defend where it sits on the memory spectrum — and say what you cannot conclude from that evidence.
You can use the cached UNRATE.csv as a template for any CSV you download by hand from the FRED website; no key needed. If you do have a free FRED API key, R/fetch_fred.R will pull one live. It reads the key from the environment — never paste a key into a file you might share.
# Optional. Requires the fredr package and a FRED_KEY in your environment.# Not required for anything in this lab.source("R/fetch_fred.R")Sys.setenv(FRED_KEY ="put_your_key_in_.Renviron_instead")cpi <-fetch_fred("CPIAUCSL", start_date ="1960-01-01", frequency =12)
6. When can you tell \(\phi = 0.99\) from \(\phi = 1\)?
Design your own experiment. Simulate both, look at whatever statistic you think should separate them, and find the sample size at which you can reliably call it. Be honest about your success rate. Then go read what Module 2 does instead — you will appreciate why the test is built the way it is.
8. Checks
R/checks.R is the referee. It is six assertions about things this module claims are true, written in plain base R so you can read every line. It does not know what you concluded and it cannot grade you — it just refuses to let an incorrect claim about the code stand, whether that claim came from you, from your AI tutor, or from a typo.
If a check fails, CHECKS.md explains what that check asserts and what a failure usually means.
source("R/checks.R")
Module 1 checks
---------------
PASS rho_1 of a phi = 0.5 AR(1) is close to 0.5 [sample rho_1 = 0.5013, target 0.500, tol 0.05]
PASS ar1_simulator(n = 500, burn_in = 200) returns 500 finite values as a ts [class = ts, length = 500, NAs = 0]
PASS a phi = 0.5, alpha = 2 AR(1) has sample mean near alpha/(1 - phi) = 4 [sample mean = 3.9974, target 4.000, tol 0.15]
PASS random walk spread at t = 200 is far wider than at t = 10 [var(y_10) = 9.2, var(y_200) = 198.1, ratio = 21.6 (theory 20)]
PASS independent random walks reject at 5% far more than 5% of the time [rejection rate = 0.780, nominal 0.05, floor for this check 0.50]
PASS data/UNRATE.csv has the expected columns and its one 2025-10 gap [columns = observation_date+UNRATE, rows = 940, NAs = 1 at 2025-10-01]
---------------
6/6 checks passed.
Where you are now
You turned on two sliders of the master equation and found:
\(\phi\) controls memory. A shock at time \(t\) is worth \(\phi^k\) at time \(t+k\).
\(\alpha\) controls level, not memory. \(\mu = \alpha/(1-\phi)\) when the process is stationary.
For a stationary AR(1), \(\rho_k = \phi^k\) — and the sample ACF finds it.
At \(\phi = 1\) stationarity fails: \(\operatorname{Var}(y_t) = t\sigma^2\) grows without bound and there is no mean to return to.
Regressing independent random walks on each other rejects a true null about three quarters of the time, and gets worse with more data.
Unemployment has \(\hat\phi \approx 0.97\), which is exactly the ambiguous region where plots and ACFs cannot tell you what you need to know.
That last point is the open question you carry into Module 2: we need a test, not a picture.
Before you close this file: open LEARNING_LOG.md and write down, in your own words, one prediction you got wrong today and what you now think instead. Your words, not your tutor’s — that is the entire value of the exercise.