Ramey and Zubairy (2018), linear model: US quarterly data, 1889Q1–2015Q4
Ramey and Zubairy (2018, JPE) estimate how much GDP rises when government spending rises, using more than a century of US quarterly data. This page replicates the baseline linear model from their Stata program jordagk.do: the impulse responses, the one-step cumulative multiplier, the two-step multiplier and the weak-instrument diagnostic. The authors build every lead, cumulated sum and lag by hand and loop ivreg2 over horizons. Here each is a single call.
Data and specification
The cleaned dataset ships with the package as docs/src/data/ramey_zubairy.csv, built from the replication package with the variable definitions of jordagk.do. GDP and spending are divided by potential GDP (the Gordon–Krenn normalization), so regression coefficients are already in dollars of GDP per dollar of spending.
newsy: news about future military spending, over lagged potential GDP
4 lags of newsy, y, g
Blanchard–Perotti
bp: current spending, taken as predetermined
4 lags of y, g
newsy starts in 1890Q1: the file begins in 1889Q1, and its first lag does not exist.
Impulse responses
The responses of spending and GDP at horizon \(h\) are the coefficients on the shock in \(x_{t+h} = \alpha_h + \beta_h s_t + \gamma_h' w_{t-1} + u_{t+h}\). leads(x) puts \(x_{t+h}\) on the left and lp runs the regression for every \(h\) up to horizon:
Responses of government spending and GDP. 90% bands, Newey–West automatic bandwidth.
After a news shock both responses are hump-shaped: spending peaks at 0.37 after 11 quarters, GDP at 0.29 after 10. Under Blanchard–Perotti the shock is current spending, so spending moves one for one on impact, keeps rising to 2.25 after 3 quarters, then declines.
The one-step cumulative multiplier
The multiplier at horizon \(h\) is cumulative extra GDP per cumulative extra dollar of spending. Ramey and Zubairy estimate it in one step, as \(m_h\) in
with cumulative spending instrumented by the shock \(s_t\). Both sides move with the horizon, and cumul handles either side, so lpiv returns the whole multiplier path at once:
mult_news =lpiv(@formula(cumul(y) ~ (cumul(g) ~ newsy) +lags(newsy, 4) +lags(y, 4) +lags(g, 4)), rz; horizon =20)mult_bp =lpiv(@formula(cumul(y) ~ (cumul(g) ~ bp) +lags(y, 4) +lags(g, 4)), rz; horizon =20)s_news =summarize(mult_news, hac; level =0.90) # the multiplier is `shock`, no `term`s_bp =summarize(mult_bp, hac; level =0.90)[m.nobs for m in mult_bp.models][[1, 9, 21]] # h = 0, 8, 20: Stata's samples
3-element Vector{Int64}:
504
496
484
One-step cumulative multipliers. 90% bands, Newey–West automatic bandwidth; the dotted line is a multiplier of one.
The two-year multiplier is 0.67 for the news shock and 0.41 for Blanchard–Perotti. Both are below one, the paper’s headline result.
Against the published numbers
RZ is the authors’ published multiplier from Multiplier-Standard-Errors.xlsx in the replication package. Horizons 8 and 16 are the two-year and four-year multipliers in the paper.
h
news
s.e.
RZ
RZ s.e.
BP
s.e.
RZ
RZ s.e.
0
1.306
0.321
1.255
0.329
0.179
0.149
0.208
0.155
1
1.060
0.261
1.035
0.235
0.223
0.120
0.235
0.143
2
0.847
0.174
0.820
0.155
0.257
0.105
0.257
0.133
3
0.706
0.126
0.695
0.123
0.254
0.101
0.251
0.133
4
0.680
0.100
0.674
0.097
0.272
0.105
0.271
0.132
6
0.677
0.076
0.673
0.074
0.356
0.098
0.353
0.119
8
0.669
0.061
0.667
0.060
0.413
0.089
0.411
0.104
10
0.706
0.055
0.705
0.053
0.442
0.089
0.439
0.102
12
0.719
0.051
0.717
0.052
0.461
0.091
0.458
0.102
14
0.717
0.045
0.716
0.046
0.473
0.094
0.471
0.105
16
0.710
0.043
0.708
0.046
0.469
0.107
0.467
0.119
18
0.715
0.046
0.713
0.053
0.456
0.119
0.453
0.133
20
0.729
0.054
0.727
0.063
0.443
0.125
0.440
0.139
At these 13 horizons, for both shocks, every point estimate is within 0.19 published standard errors of the authors’ value, with a median gap of 0.03. The remaining difference is most likely the data vintage: the spreadsheet of published numbers predates the last revision of the shipped data by a week.
The standard errors differ more, 0.76 to 1.12 times the published ones, for two reasons:
Bandwidth. The paper uses ivreg2, robust bw(auto). Both sides implement the Newey–West (1994) automatic rule, but they select different bandwidths on the same regression. With a fixed bandwidth, Bartlett(b) here and bw(b) in Stata, coefficients and IV standard errors agree to about \(10^{-10}\).
Degrees of freedom.ivreg2 applies no small-sample factor, and neither do the IV standard errors here. The OLS impulse-response standard errors above carry \(n/(n-k)\) and are 1–3% larger than Stata’s ivreg2 equivalents.
The two-step multiplier
The paper’s alternative estimator is the ratio of the cumulative sums of the two impulse responses, \(\sum_{j \le h} \hat\beta^y_j \,/\, \sum_{j \le h} \hat\beta^g_j\):
twostep(ry, rg, shock) =cumsum(coefpath(ry; term = shock)) ./cumsum(coefpath(rg; term = shock))two_news =twostep(irf_y_news, irf_g_news, :newsy)two_bp =twostep(irf_y_bp, irf_g_bp, :bp)(news =maximum(abs.(two_news .- s_news.coef)), bp =maximum(abs.(two_bp .- s_bp.coef)))
(news = 0.003135302083235758, bp = 0.0003613935994520867)
The two estimators agree to within 0.0031 (news) and 0.0004 (Blanchard–Perotti) at every horizon, because here the one-step regressions use the same sample as the impulse-response regressions. Ramey and Zubairy note that the two can diverge in other specifications, precisely when the samples differ.
Is the instrument strong enough?
A multiplier estimated with a weak instrument is biased and its bands undercover. weakivtest computes the Montiel Olea–Pflueger effective F at each horizon, from the HAC covariance attached to the model:
mop(m) = [weakivtest(m +vcov(hac), h) for h in0:20]w_news, w_bp =mop(mult_news), mop(mult_bp)w_news[9] # two years out
Montiel-Pflueger robust weak instrument test
──────────────────────────────────────────────────────
TSLS coefficient (SE): 0.6690 ( 0.0625)
LIML coefficient (SE): 0.6690 ( 0.0625)
GMMf coefficient (SE): 0.6690 ( 0.0625)
kappa: 1.0000
F[nonrobust]: 71.863
F[effective]: 18.087
F[robust]: 18.087
Confidence level alpha: 5%
──────────────────────────────────────────────────────
Use:
Compare F[effective] or F[robust] to the common critical values below.
If F exceeds the critical value at tau, weak-IV bias is below that tau bound.
──────────────────────────────────────────────────────
tau Common
──────────────────────────────────────────────────────
tau=5% 37.418
tau=10% 23.109
tau=20% 15.062
tau=30% 12.045
──────────────────────────────────────────────────────
Montiel Olea–Pflueger effective F, log scale. The dashed line is the 5% critical value for a worst-case bias of 10%.
The news instrument never clears the critical value of 23.11: its effective F peaks at 19.57 at 5 quarters. Blanchard–Perotti is strong at short horizons and falls below the threshold from horizon 11 on. At \(h = 0\) its instrument is the regressor, so the first stage is exact and the figure omits that point. Ramey and Zubairy respond to weak instruments with Anderson–Rubin inference, which the package does not implement yet.