library(ggplot2)
library(forecast)
library(patchwork)
# Stamped copies of the canonical course helpers.
source("R/arma_simulator.R")
source("R/diagnostics.R")
source("R/fingerprint_mysteries.R") # source it; do not read below its spoiler banner
theme_set(theme_bw(base_size = 12))Module 3 Lab: ARMA Processes and Their Fingerprints
Econ 6376 — Applied Time Series Econometrics · AI Learning Pack
How to use this lab
Everything in the Core lab runs as shipped: no API key, no internet, and no fill-in-the-blank code. Your work is to predict, run, explain, and modify. Before a plot appears, write down what you expect its ACF and PACF to do. Use LEARNING_LOG.md, a code comment, or paper. A prediction made only after seeing the plot is not a prediction.
If you are working with an AI tutor, it will ask you for that prediction before it runs anything meaningful, and it will run the code either way. It is not gatekeeping you — it is trying to make sure you get the “huh, I didn’t expect that” moment, because that is the one you remember when an unlabeled correlogram is in front of you and there is no tutor to ask.
The master equation is
\[ y_t = \alpha + \sum_{j=1}^{p}\phi_j y_{t-j} + \epsilon_t + \sum_{l=1}^{q}\theta_l\epsilon_{t-l}. \]
Module 3 turns on the full AR and MA banks. The central distinction is simple:
- AR terms are memory of past outcomes.
- MA terms are memory of past shocks.
The practical problem is harder: infer which memory mechanism could have left the ACF/PACF pattern you observe. The four sealed fingerprints near the end are the test of that skill.
1. Setup and the Module 2 handoff
Module 2 comes first in the workflow even when Module 3 is the destination. Before using a correlogram to choose \((p,q)\), settle the transformation order \(d\). A strongly persistent levels series and a stationary differenced series ask different ARMA questions. An ACF/PACF fingerprint is a model-identification aid after that decision, not a substitute for thinking about stationarity.
The working sequence is therefore
\[ \text{plot and test} \longrightarrow \text{settle } d \longrightarrow \text{read ACF/PACF} \longrightarrow \text{carry candidate }(p,q)\text{ models to Module 4}. \]
2. The master equation in code
The simulator below is the same course helper used in the canonical notes. Read it as an implementation of the equation above. In each loop iteration it adds an AR part, an MA part, and the current innovation.
arma_simulatorfunction (n = 500, alpha = 0, phi = c(), theta = c(), sigma = 1,
burn_in = 200)
{
p <- length(phi)
q <- length(theta)
total <- burn_in + n
eps <- rnorm(total, 0, sigma)
y <- numeric(total)
for (t in seq_len(total)) {
ar_part <- 0
if (p > 0 && t > p) {
ar_part <- sum(phi * y[(t - 1):(t - p)])
}
ma_part <- 0
if (q > 0 && t > q) {
ma_part <- sum(theta * eps[(t - 1):(t - q)])
}
y[t] <- alpha + ar_part + ma_part + eps[t]
}
ts(y[(burn_in + 1):total])
}
Empty phi means no AR terms; empty theta means no MA terms. Here are all three families through one interface.
set.seed(301)
sim_ar <- arma_simulator(n = 120, phi = 0.70)
sim_ma <- arma_simulator(n = 120, theta = 0.70)
sim_arma <- arma_simulator(n = 120, phi = 0.60, theta = 0.40)
data.frame(
family = c("AR(1)", "MA(1)", "ARMA(1,1)"),
mean = round(c(mean(sim_ar), mean(sim_ma), mean(sim_arma)), 3),
sd = round(c(sd(sim_ar), sd(sim_ma), sd(sim_arma)), 3),
rho_1 = round(c(acf(sim_ar, plot = FALSE)$acf[2],
acf(sim_ma, plot = FALSE)$acf[2],
acf(sim_arma, plot = FALSE)$acf[2]), 3)
) family mean sd rho_1
1 AR(1) -0.281 1.141 0.578
2 MA(1) -0.059 1.210 0.413
3 ARMA(1,1) -0.066 1.674 0.787
Those four summaries cannot identify the family. Identification needs the shape across lags, which is why we build the correlogram fingerprint.
3. One innovation stream, two kinds of memory
Predict what happens after the unusually large shock at \(t=40\):
- In the AR system, for how long does the shock remain visible?
- In the MA(3), after which observation must its direct effect be gone?
The next chunk passes the exact same innovations through both systems.
set.seed(302)
n_common <- 100
eps_common <- rnorm(n_common, sd = 0.55)
eps_common[40] <- eps_common[40] + 4
# AR(1): the shock changes y_40, which becomes an input to y_41, and so on.
y_ar_common <- numeric(n_common)
y_ar_common[1] <- eps_common[1]
for (t in 2:n_common) {
y_ar_common[t] <- 0.75 * y_ar_common[t - 1] + eps_common[t]
}
# MA(3): the shock enters only at t, t+1, t+2, and t+3.
lag_one <- c(0, eps_common[-n_common])
lag_two <- c(0, 0, eps_common[seq_len(n_common - 2)])
lag_three <- c(0, 0, 0, eps_common[seq_len(n_common - 3)])
y_ma_common <- eps_common + 0.75 * lag_one + 0.50 * lag_two + 0.25 * lag_three
common_paths <- rbind(
data.frame(t = seq_len(n_common), value = y_ar_common,
system = "AR(1): memory of y"),
data.frame(t = seq_len(n_common), value = y_ma_common,
system = "MA(3): memory of shocks")
)
ggplot(common_paths, aes(t, value)) +
geom_hline(yintercept = 0, color = "grey75") +
geom_vline(xintercept = 40, linetype = 2, color = "firebrick") +
geom_line(linewidth = 0.55) +
facet_wrap(~system, ncol = 1) +
labs(title = "The same innovations through two memory systems",
subtitle = "The dashed line marks the deliberately enlarged innovation",
x = "t", y = expression(y[t]))The AR response fades geometrically because each affected outcome feeds the next outcome. The MA response has finite memory: after lag 3, that particular shock is no longer present in the equation. New shocks keep arriving in both panels, so the observed series does not become literally flat.
4. AR roots: persistence and oscillation
For an AR(2), the course uses two equivalent root conventions:
\[ \Phi(L)=1-\phi_1L-\phi_2L^2=0 \quad\Longleftrightarrow\quad r^2-\phi_1r-\phi_2=0. \]
Lag-polynomial roots must lie outside the unit circle. Their reciprocals, the characteristic roots \(r\), must lie inside. The next table uses the characteristic-root convention because its quadratic formula is cleaner.
Before running it, classify each row: stationary or not; real or complex.
characteristic_roots <- function(phi) {
polyroot(c(-phi[2], -phi[1], 1))
}
root_cases <- list(
"phi = (0.5, 0.3)" = c(0.50, 0.30),
"phi = (0.8, 0.3)" = c(0.80, 0.30),
"phi = (0.6, -0.5)" = c(0.60, -0.50)
)
root_table <- do.call(rbind, lapply(names(root_cases), function(label) {
phi <- root_cases[[label]]
roots <- characteristic_roots(phi)
data.frame(
case = label,
discriminant = round(phi[1]^2 + 4 * phi[2], 3),
root_1 = sprintf("%.3f%+.3fi", Re(roots[1]), Im(roots[1])),
root_2 = sprintf("%.3f%+.3fi", Re(roots[2]), Im(roots[2])),
largest_modulus = round(max(Mod(roots)), 3),
stationary = all(Mod(roots) < 1),
root_shape = if (any(abs(Im(roots)) > 1e-8)) "complex" else "real"
)
}))
rownames(root_table) <- NULL
root_table case discriminant root_1 root_2 largest_modulus
1 phi = (0.5, 0.3) 1.45 -0.352-0.000i 0.852+0.000i 0.852
2 phi = (0.8, 0.3) 1.84 -0.278+0.000i 1.078-0.000i 1.078
3 phi = (0.6, -0.5) -1.64 0.300+0.640i 0.300-0.640i 0.707
stationary root_shape
1 TRUE real
2 FALSE real
3 TRUE complex
The sign of \(\phi_1^2+4\phi_2\) tells you real versus complex; it does not tell you stationary versus nonstationary. Modulus answers the latter question. Stationary real roots yield mixtures of geometric decay. Stationary complex roots yield damped oscillation.
set.seed(303)
ar2_real <- arma_simulator(n = 500, phi = c(0.50, 0.30))
set.seed(304)
ar2_complex <- arma_simulator(n = 500, phi = c(0.60, -0.50))
print(show_fingerprint(ar2_real, "Stationary AR(2): real characteristic roots"))print(show_fingerprint(ar2_complex, "Stationary AR(2): complex characteristic roots"))Both PACFs should cut off after lag 2 in the population. The ACF shapes differ because the roots differ. Sample spikes beyond the population cutoff are noise, not a reason to invent a higher order whenever one bar crosses a band.
5. MA(1): cutoff, half-bound, and invertibility
For
\[ y_t=\epsilon_t+\theta\epsilon_{t-1}, \]
the only shared innovation in \(y_t\) and \(y_{t-1}\) is \(\epsilon_{t-1}\). There is no shared innovation once the separation exceeds one period. Therefore
\[ \rho(1)=\frac{\theta}{1+\theta^2}, \qquad \rho(k)=0\quad(k\ge 2), \qquad |\rho(1)|\le \frac12. \]
Population cutoff and the half-bound
theta_grid <- c(-2, -1, -0.7, 0, 0.7, 1, 2)
data.frame(
theta = theta_grid,
rho_1_formula = round(theta_grid / (1 + theta_grid^2), 4)
) theta rho_1_formula
1 -2.0 -0.4000
2 -1.0 -0.5000
3 -0.7 -0.4698
4 0.0 0.0000
5 0.7 0.4698
6 1.0 0.5000
7 2.0 0.4000
ARMAacf(ma = 0.7, lag.max = 8) 0 1 2 3 4 5 6 7
1.0000000 0.4697987 0.0000000 0.0000000 0.0000000 0.0000000 0.0000000 0.0000000
8
0.0000000
The bound is a statement about the population ACF. A finite-sample \(\hat\rho(1)\) can cross \(1/2\) by sampling variation.
Reciprocal representations
MA processes are stationary for every finite coefficient vector. Their separate restriction is invertibility: every root of \(\Theta(L)=1+\theta_1L+\cdots+\theta_qL^q\) must lie outside the unit circle. For an MA(1), the root is \(-1/\theta\).
For nonzero \(\theta\), the values \(\theta\) and \(1/\theta\) give the same ACF. With an appropriate change in innovation variance, they give the same autocovariances too. Away from the boundary \(|\theta|=1\), only one member of the reciprocal pair is invertible, which gives us a unique, recoverable shock representation.
reciprocal_pair <- c(2, 1 / 2)
data.frame(
theta = reciprocal_pair,
ma_root = -1 / reciprocal_pair,
invertible = abs(-1 / reciprocal_pair) > 1,
rho_1 = reciprocal_pair / (1 + reciprocal_pair^2),
innovation_variance_for_same_gamma = c(1, 4)
) theta ma_root invertible rho_1 innovation_variance_for_same_gamma
1 2.0 -0.5 FALSE 0.4 1
2 0.5 -2.0 TRUE 0.4 4
A stationary AR(\(p\)) has an MA(\(\infty\)) representation. More precisely, Wold decomposes a covariance-stationary process into a deterministic component and a one-sided MA(\(\infty\)) innovation component; the stationary ARMA models here are the purely nondeterministic applied case. In the other direction, invertibility makes an MA(\(q\)) an AR(\(\infty\)): past observations can recover past innovations through a convergent inverse polynomial. This two-way isomorphism is the parsimony payoff. A small ARMA model can encode dynamics that a pure AR or pure MA representation would need many coefficients to approximate. The longer high-order AR demonstration remains optional in the Explorations section.
6. PACF as the deepest-lag coefficient
At lag \(k\), regress \(y_t\) on \(y_{t-1},\ldots,y_{t-k}\) and read the coefficient on the deepest lag. That coefficient is the sample analogue of \(\phi_{kk}\):
\[ y_t=b_0+b_1y_{t-1}+\cdots+b_ky_{t-k}+e_t, \qquad \widehat{\operatorname{PACF}}(k)=\hat b_k. \]
The course helper fits those regressions on one common sample window. Base R’s pacf() uses a different finite-sample convention, so close agreement rather than byte equality is the right target.
set.seed(8675309)
pacf_series <- arima.sim(n = 500, model = list(ar = c(0.60, -0.30)))
pacf_ols <- pacf_by_regression(pacf_series, k_max = 6)
pacf_builtin <- as.numeric(pacf(pacf_series, lag.max = 6, plot = FALSE)$acf)
data.frame(
lag = 1:6,
deepest_lag_ols = round(pacf_ols, 4),
stats_pacf = round(pacf_builtin, 4),
absolute_gap = round(abs(pacf_ols - pacf_builtin), 4)
) lag deepest_lag_ols stats_pacf absolute_gap
phi_11 1 0.4514 0.4459 0.0054
phi_22 2 -0.2582 -0.2571 0.0010
phi_33 3 0.0352 0.0404 0.0052
phi_44 4 -0.0493 -0.0520 0.0027
phi_55 5 0.0156 0.0158 0.0002
phi_66 6 -0.0057 -0.0058 0.0001
For a population AR(2), lags 1 and 2 can add direct information; deeper lags do not once the first two have been controlled for. Thus the PACF cuts off at \(p\) for AR(\(p\)), while its ACF tails off. For MA(\(q\)), reverse the diagnostic: the ACF cuts off at \(q\), while its PACF tails off.
7. A labeled fingerprint gallery
Commit one prediction per process before rendering the gallery. The useful questions are: Which correlogram cuts off? Where? Does the other decay smoothly, alternate, oscillate, or remain hard to classify?
set.seed(1985)
y_ar1 <- arma_simulator(n = 500, phi = 0.70)
y_ar2_real <- arma_simulator(n = 500, phi = c(0.50, 0.30))
y_ar2_complex <- arma_simulator(n = 500, phi = c(0.60, -0.50))
y_ma1 <- arma_simulator(n = 500, theta = 0.70)
y_ma2 <- arma_simulator(n = 500, theta = c(0.80, 0.50))
y_arma11 <- arma_simulator(n = 500, phi = 0.60, theta = 0.40)
print(show_fingerprint(y_ar1, "AR(1), phi = 0.70"))print(show_fingerprint(y_ar2_real, "AR(2), real roots: phi = (0.50, 0.30)"))print(show_fingerprint(y_ar2_complex, "AR(2), complex roots: phi = (0.60, -0.50)"))print(show_fingerprint(y_ma1, "MA(1), theta = 0.70"))print(show_fingerprint(y_ma2, "MA(2), theta = (0.80, 0.50)"))print(show_fingerprint(y_arma11, "ARMA(1,1), phi = 0.60, theta = 0.40"))Use the population rules as tendencies, not a bar-counting algorithm:
| Family | Population ACF | Population PACF |
|---|---|---|
| AR(\(p\)) | tails off | cuts off after \(p\) |
| MA(\(q\)) | cuts off after \(q\) | tails off |
| ARMA(\(p,q\)) | tails off | tails off |
The mixed case is the problem Module 4 will solve. When both sides tail off, a correlogram may narrow the candidate set but generally does not uniquely identify \((p,q)\).
Shrink the sample and blur the signature
The next pair uses the same DGP and seed. The short draw is the same path observed for far fewer periods. Predict which apparent cutoff will be easiest to misread.
set.seed(1986)
ma2_long <- arma_simulator(n = 500, theta = c(0.80, 0.50))
set.seed(1986)
ma2_short <- arma_simulator(n = 60, theta = c(0.80, 0.50))
print(show_fingerprint(ma2_long, "Same MA(2), T = 500"))print(show_fingerprint(ma2_short, "Same MA(2), T = 60"))At \(T=60\), the approximate bands are wider and the population zeros are replaced by noisier sample estimates. “Ambiguous at this sample size” is a legitimate, calibrated conclusion. It is better than a confident order chosen from one stray spike.
8. Assumption break: nearly non-invertible MA(1)
Now push the MA root toward the unit circle. With \(\theta=-0.97\), \(\Theta(L)=1-0.97L\) has root \(1/0.97\approx1.031\): technically invertible, but close to the boundary produced by over-differencing a trend-stationary process.
Before running the fit, predict where \(\hat\theta\) will land and whether a clean optimizer return would make an ordinary Wald interval trustworthy.
set.seed(1999)
y_near_boundary <- arma_simulator(n = 500, theta = -0.97)
show_fingerprint(y_near_boundary,
"MA(1), theta = -0.97 — nearly non-invertible")fit_near_boundary <- arima(
y_near_boundary,
order = c(0, 0, 1),
include.mean = FALSE
)
data.frame(
estimate = unname(fit_near_boundary$coef["ma1"]),
reported_se = sqrt(diag(fit_near_boundary$var.coef))["ma1"],
convergence_code = fit_near_boundary$code
) estimate reported_se convergence_code
ma1 -0.9817495 0.008120466 0
The estimate piles up near \(-1\). The printed standard error and convergence code do not repair boundary asymptotics: the usual interior-parameter likelihood approximation is unreliable there. In applied work, \(\hat\theta\approx-1\) after differencing is a warning to revisit the transformation and specification.
9. Four sealed fingerprint mysteries
For each series, inspect the time plot, ACF, and PACF, then commit to one of:
- AR(\(p\)), with a proposed \(p\);
- MA(\(q\)), with a proposed \(q\);
- ARMA, with the order unresolved from the plots;
- ambiguous at this sample size.
Write the evidence, not just the label. If you are using an AI tutor, it may ask once for that prediction immediately before a reveal. It will always run evidence requests, and if you ask directly for a reveal it will comply in that turn.
Mystery 1
show_fingerprint(fingerprint_mystery(1), "Fingerprint mystery 1")# Deliberately not run during render: execute after committing your diagnosis.
reveal_fingerprint(1)Mystery 2
show_fingerprint(fingerprint_mystery(2), "Fingerprint mystery 2")# Deliberately not run during render: execute after committing your diagnosis.
reveal_fingerprint(2)Mystery 3
show_fingerprint(fingerprint_mystery(3), "Fingerprint mystery 3")# Deliberately not run during render: execute after committing your diagnosis.
reveal_fingerprint(3)Mystery 4
show_fingerprint(fingerprint_mystery(4), "Fingerprint mystery 4")# Deliberately not run during render: execute after committing your diagnosis.
reveal_fingerprint(4)The four reveal chunks are the only non-running Core chunks. Running them during render would publish the answers beside the evidence and destroy the exercise. Each is complete, runnable code for the console after a prediction.
10. UNRATE: full evidence and a labeled sensitivity
The cached FRED file makes this section reproducible offline. It contains one missing release at October 2025, so we use the complete monthly run through September 2025 rather than allowing a ts calendar to jump over a month.
unrate <- read.csv("data/UNRATE.csv", stringsAsFactors = FALSE)
unrate$observation_date <- as.Date(unrate$observation_date)
unrate <- unrate[order(unrate$observation_date), ]
first_missing <- which(is.na(unrate$UNRATE))[1]
if (!is.na(first_missing)) {
unrate <- unrate[seq_len(first_missing - 1L), ]
}
unrate_level <- ts(unrate$UNRATE, start = c(1948, 1), frequency = 12)
delta_unrate <- diff(unrate_level)The Module 2 callback is visible immediately: levels are highly persistent, so we read the fingerprint of the differenced series.
p_unrate_levels <- ggAcf(unrate_level, lag.max = 36) +
ggtitle("UNRATE levels — slow ACF decay")
p_unrate_difference <- ggAcf(delta_unrate, lag.max = 36) +
ggtitle("Differenced UNRATE — full sample")
p_unrate_levels / p_unrate_differencePrimary analysis: keep the full sample
show_fingerprint(delta_unrate, "Differenced US unemployment rate — full sample")The full-sample ACF is nearly flat: none of its first 24 spikes clears the usual approximate band and \(\hat\rho(1)\) is about 0.036. The time plot prevents an overconfident “white noise” verdict. The unemployment rate rose 10.4 percentage points in April 2020; that observation inflates the variance used to normalize every sample autocovariance, compressing the correlogram.
Sensitivity: exclude March 2020 through December 2021 changes
This is explicitly a sensitivity analysis, not a replacement data set. The full sample remains the primary evidence. We ask what dependence becomes visible when one extraordinary episode is withheld.
delta_dates <- unrate$observation_date[-1]
keep_sensitivity <- !(
delta_dates >= as.Date("2020-03-01") &
delta_dates <= as.Date("2021-12-01")
)
delta_unrate_sensitivity <- as.numeric(delta_unrate)[keep_sensitivity]
p_full_acf <- ggAcf(delta_unrate, lag.max = 24) +
ggtitle("ACF — full sample")
p_full_pacf <- ggPacf(delta_unrate, lag.max = 24) +
ggtitle("PACF — full sample")
p_sensitivity_acf <- ggAcf(delta_unrate_sensitivity, lag.max = 24) +
ggtitle("ACF — excluding 2020–21 (sensitivity)")
p_sensitivity_pacf <- ggPacf(delta_unrate_sensitivity, lag.max = 24) +
ggtitle("PACF — excluding 2020–21 (sensitivity)")
(p_full_acf + p_full_pacf) / (p_sensitivity_acf + p_sensitivity_pacf)acf_summary <- function(x, lag_max = 24) {
values <- as.numeric(acf(x, lag.max = lag_max, plot = FALSE)$acf)[-1]
data.frame(
observations = length(x),
rho_1 = values[1],
spikes_outside_approximate_band =
sum(abs(values) > 1.96 / sqrt(length(x)))
)
}
rbind(
cbind(sample = "full", acf_summary(delta_unrate)),
cbind(sample = "exclude 2020-03 through 2021-12",
acf_summary(delta_unrate_sensitivity))
) sample observations rho_1
1 full 932 0.03567108
2 exclude 2020-03 through 2021-12 910 0.10738723
spikes_outside_approximate_band
1 0
2 8
With the window excluded, 8 of the first 24 ACF lags clear the approximate band; the ACF is positive through lag 6, and both correlograms tail off with a negative annual echo around lag 12. That makes a low-order ARMA with possible seasonal structure a defensible candidate, not an identified truth. Module 4 compares candidate orders; Module 5 introduces explicit seasonal operators.
The larger lesson is methodological: read the evidence actually observed, inspect why it may be distorted, and label sensitivity analyses transparently.
11. Explorations
Nothing in this section is required, and none of it appears on problem sets or exams — assessments draw only on the Core tier of the module notes. These are here because the interesting questions usually start where the syllabus stops.
If one of these grabs you, chase it. If a different question grabs you, chase that one instead — that is a better outcome than any of the five below. Your AI tutor’s job in this section is to help you turn a hunch into an experiment you can actually run, then connect what you find back to a Core concept.
Exploration 1 — Design an oscillation (guided)
Choose complex characteristic roots with modulus \(0.9\), translate them into an AR(2) coefficient pair, and simulate it. How does the angle of the roots change the period of the damped oscillation while the modulus controls its decay?
modulus <- 0.9
angle <- pi / 4
root <- modulus * exp(1i * angle)
phi_designed <- c(2 * Re(root), -Mod(root)^2)
phi_designed[1] 1.272792 -0.810000
set.seed(311)
show_fingerprint(
arma_simulator(n = 500, phi = phi_designed),
"AR(2) designed from complex roots"
)Exploration 2 — Population rule versus sample bar counting (guided)
Repeat a short MA(1) simulation many times. How often does a lag beyond 1 cross the approximate band even though its population autocorrelation is exactly zero?
set.seed(312)
false_tail_spikes <- replicate(500, {
draw <- arma_simulator(n = 60, theta = 0.7)
sample_acf <- as.numeric(acf(draw, lag.max = 12, plot = FALSE)$acf)[-1]
any(abs(sample_acf[2:12]) > 1.96 / sqrt(length(draw)))
})
mean(false_tail_spikes)[1] 0.5
Exploration 3 — Optional high-order AR approximation (guided)
This numerical approximation is useful confirmation, but it is not needed to understand invertibility or to complete the Core lab. For an invertible MA(1) with \(\theta=0.7\), the implied coefficients in the equation with \(y_t\) on the left are \(0.7,-0.49,0.343,\ldots\).
set.seed(313)
long_ma <- arma_simulator(n = 5000, theta = 0.7)
high_ar <- ar(long_ma, order.max = 8, aic = FALSE, method = "ols")
data.frame(
lag = 1:8,
fitted_ar = as.numeric(high_ar$ar),
implied_ar_infinity = (-1)^(1:8 + 1) * 0.7^(1:8)
) lag fitted_ar implied_ar_infinity
1 1 0.710307532 0.70000000
2 2 -0.507419291 -0.49000000
3 3 0.330470686 0.34300000
4 4 -0.206670104 -0.24010000
5 5 0.122113514 0.16807000
6 6 -0.076742283 -0.11764900
7 7 0.053408381 0.08235430
8 8 -0.004716653 -0.05764801
Exploration 4 — Change the innovation distribution (guided)
arima.sim() accepts a supplied innovation vector. Compare Gaussian and heavy-tailed innovations under the same ARMA coefficients. Which features of the population fingerprint survive, and which sample spikes become less stable?
set.seed(314)
n_heavy <- 500
gaussian <- arima.sim(n = n_heavy, model = list(ar = 0.6, ma = 0.4),
innov = rnorm(n_heavy))
heavy_tail <- arima.sim(n = n_heavy, model = list(ar = 0.6, ma = 0.4),
innov = rt(n_heavy, df = 3) / sqrt(3))
print(show_fingerprint(gaussian, "ARMA(1,1), Gaussian innovations"))print(show_fingerprint(heavy_tail, "ARMA(1,1), standardized t(3) innovations"))Exploration 5 — Refresh UNRATE live (open)
Only do this if you have the optional fredr package, internet access, and a FRED key in FRED_KEY. R/fetch_fred.R contains the live pull. Compare the refreshed coverage and missing-release handling with data/README.md; never put the key in a script or learning log.
12. Checks
R/checks.R is the referee. Seven assertions about things this module claims are true — the simulator contract, the AR(2) root cases, the MA(1) algebra, the population cutoff identities, the PACF construction, the near-boundary fit, and the integrity of the sealed series and the cached UNRATE contrast. It does not know what you concluded and it cannot grade you — it just refuses to let an incorrect claim about the code stand, whether that claim came from you, from your AI tutor, or from a typo. It exposes no mystery answers.
If a check fails, CHECKS.md explains what that check asserts and what a failure usually means.
source("R/checks.R")Module 3 checks
---------------
PASS arma_simulator returns a deterministic length-n time series [class ts, length 80, identical on re-seed: TRUE]
PASS AR(2) root checks distinguish stationary, nonstationary, and complex cases [min |root|: (0.5, 0.3) = 1.174; (0.8, 0.3) = 0.927; (0.6, -0.5) = 1.414 with |Im| = 1.281]
PASS MA(1) rho(1) = theta/(1+theta^2), rho(h)=0 after lag 1, and |rho(1)| <= 1/2 [ARMAacf rho(1) = 0.469799 vs formula 0.469799; max |rho(2..6)| = 0.0e+00; max |rho(1)| on the theta grid = 0.500]
PASS population AR(2) PACF and MA(2) ACF cut off after lag 2 [max |AR(2) PACF| beyond lag 2 = 1.2e-16; max |MA(2) ACF| beyond lag 2 = 0.0e+00]
PASS pacf_by_regression agrees with stats::pacf within 0.02 at lags 1-4 [maximum absolute difference = 0.0045, tol 0.02]
PASS the near-noninvertible MA(1) example estimates theta close to -1 [estimated theta = -0.98175, window (-1.001, -0.95), true theta -0.97]
PASS sealed fingerprints and the cached UNRATE full/sensitivity contrast are reproducible [mystery pins: ok | raw cache: ok | raw end: 2026-04-01 | raw rows: 940 | modeling rows: 933 | full-sample hits: 0 | sensitivity hits: 8 | full-sample rho(1): 0.036]
---------------
7/7 checks passed.
Where you are now
You turned on the full AR and MA banks of the master equation and found:
- AR terms are memory of past outcomes; MA terms are memory of past shocks. The same innovation stream fades geometrically through the AR(1) and drops out of the MA(3) after three periods.
- Stationarity of an AR(\(p\)) is a joint root condition — every root of \(\Phi(L) = 0\) outside the unit circle, equivalently every characteristic root inside — and \((0.8, 0.3)\) fails it with both coefficients below one. Complex roots inside the circle give damped oscillation, not a broken model.
- A finite MA(\(q\)) is always stationary; its separate restriction is invertibility, which selects one member of each reciprocal \(\theta\) pair and makes past shocks recoverable from past observations.
- In the population, the PACF cuts off at \(p\) for AR(\(p\)) and the ACF cuts off at \(q\) for MA(\(q\)); ARMA tails off on both sides. In a sample, bars near a band are evidence to weigh, not a counter, and “ambiguous at this \(T\)” is a legitimate call.
- \(\hat\theta\) piled up near \(-1\) is a warning about the transformation, and the printed standard error does not rescue it.
- Differenced UNRATE shows a nearly flat full-sample correlogram, compressed by the April 2020 jump, and visible low-order dependence once the March 2020–December 2021 changes are withheld — a labeled sensitivity result, not a replacement data set.
If all seven checks pass, you have a reproducible Module 3 environment. The next step is interpretive: turn the correlogram into a small candidate set, then let Module 4 compare those candidates using likelihood, information criteria, and residual diagnostics.
Before you close this file: open LEARNING_LOG.md and check that every mystery series you attempted has both halves written down — what you called it, and what it was. The gap between those two columns, in your own words, is the most valuable thing this lab produces.