正在学习

20.3.1 Empirical Example

20.3.1 Empirical Example

As an empirical illustration, we consider long-run forecasts of the quarterly inflation rate, i.e., the log-first difference in the CPI, the unemployment rate, and the interest rate. These time series are quite persistent as shown in previous chapters. We show results for forecast horizons h = {1, 2, 3, 4, 8, 12, 16, 20} quarters and compare forecasts from a random walk, random walk with drift, an AR model with lag length selected by the AIC, and forecasts from AR(1) and AR(4) models. The out-ofsample results reported in table 20.1 are generated for the period 1970–2014 using an expanding estimation window.

First consider the inflation rate. Here the results show that the autoregressive models generally perform better than the random walk model for horizons of 1 through 3 quarters, but that the random walk model is better at the longest 20 quarter horizon. The random walk model with drift produces very poor forecasts, especially at long horizons. For the unemployment rate the AR(4) model dominates at all horizons and the random walk model produces RMSE values that are generally a bit higher than those from the AR models. Conversely, the random walk model produces the lowest RMSE values for the interest rate series at both short and long horizons.

In all cases the random walk with drift model performs very poorly because the series investigated here are not trending deterministically in any particular direction. Whether the simple random walk specification or the autoregressive specification is best varies, however, across variables and across sample periods. Which model produces the best forecasts will depend on the true, but unknown, degree of persistence of the underlying series; this is consistent with our theoretical results.

20.4 FORECASTING WITH PERSISTENT REGRESSORS

A common forecasting situation involves a regressor that displays persistence — the regressor is reverting to its mean only very slowly, if at all—but appears to have predictive power based on tests for the exclusion of that regressor in a linear regression. Examples include using the forward premium to predict changes in exchange rates (Bilson, 1981), short-run interest rates, consumption, and stock prices for forecasting changes in income (Chen, 1991; Hall, 1978), or the dividend–price ratio or earnings–price ratio for forecasting future stock returns (Stambaugh, 1999; Valkanov, 2003). In many cases, variation in the outcome variable captured by such predictor variables tends to be modest, although often economically interesting, as in the case of a small, but highly persistent risk premium on a financial asset.

TABLE 20.1:
Root mean squared forecast errors (RMSE) for different forecasting models. The table reports RMSE values for the quarterly inflation, unemployment and interest rate series at different forecast horizons ranging from 1 through 20 quarters. The forecast evaluation period is 1970– 2014; parameters are estimated using a recursively expanding estimation window for the models with unknown parameters.

StepsRWRW driftAR(AIC)AR(1)AR(4)
Inflation rate
1Q4.1444.5133.3163.6263.432
2Q4.5605.7763.4883.8033.477
3Q4.1966.7603.5433.7523.467
4Q3.6537.9123.7353.8253.487
8Q4.27714.5324.2434.2394.034
12Q4.22921.2104.3614.3844.286
16Q4.37027.9524.3654.4014.346
20Q4.11834.7774.1884.2214.180
Unemployment rate
1Q0.3580.4000.2790.3650.277
2Q0.6550.7440.5620.6630.557
3Q0.9151.0590.8430.9190.831
4Q1.1401.3491.1101.1381.089
8Q1.7592.3431.7251.7011.698
12Q2.0403.2131.8991.9231.872
16Q2.1074.0261.9561.9891.937
20Q2.0814.8721.9902.0061.985
Interest rate
1Q1.1641.2071.3981.2011.269
2Q1.5161.6521.6801.5641.649
3Q1.6321.9261.7621.6431.766
4Q1.9752.4032.2041.9822.108
8Q2.8763.9883.0822.8093.083
12Q3.3515.3023.6623.3813.731
16Q3.5066.4424.1223.6324.145
20Q3.4037.5214.4573.8984.565

To capture this setting, consider the model

TABLE 20.2:

Size properties for one-sided test.

8=-0.95-0.75-0.5
0.4160.2890.180
0.1390.1160.091
0.1050.0920.079

Note: Size is 5% for a one-sided upper-tail test.

The regressor, follows the same process as in (20.2). Assumptions on are the same as in the earlier section. Regressions such as (20.10) typically do not allow to be serially correlated. If were serially correlated, we would want to model this correlation to improve predictability. Moreover, in empirical applications there is usually little or no serial correlation to be found. The terms and are assumed to be deterministic, with

We further assume that the partial sums of the residuals converge to a bivariate Brownian motion process with variance–covariance matrix, , given by10

δ is the correlation between and . This is also the long-run correlation between these shocks when is serially correlated but globally covariance stationary.

Monte Carlo simulation results have lead a number of authors to note that ttests of the hypothesis versus are not well approximated by the standard normal distribution when is close to 1 and δ is far from 0. For , the test underrejects in the upper tail and overrejects in the lower tail; the opposite holds for see Mankiw and Shapiro (1986), Elliott and Stock (1994), and Stambaugh (1999). Table 20.2 shows upper-tail rejections for tests with a nominal 5% level for and . The values for are selected to be relevant choices for forecasting changes in stock returns with the dividend–price ratio (Stambaugh, 1999).

Econometric analysis has explained such findings either as a small sample bias problem that disappears asymptotically or by modeling as being local-to-unity and showing that there is always a region for in which such tests will overreject, regardless of the sample size. Under both approaches, we have a SUR model for , and we can rotate so , where is orthogonal to . If there are no deterministic terms in the model, this allows us to write the OLS estimate of as

where is the OLS estimate of obtained from an AR(1) regression on . The intuition behind the “small sample” approach becomes clear from. The second line has a sampling distribution centered on 0. The OLS coefficient is downward biased for so if then will be upwards biased, as found in the simulations discussed above. The extent of the bias depends on the nature of the deterministic terms and is greater when constants are estimated in the regressions. Stambaugh (1999) presents analytical results for the size of the bias when Amihud and Hurvich (2004) suggest a correction for the bias, and base tests on their corrected estimator. Each of these methods requires that . The problem with this approach is that for roots “close” to 1, the assumptions underlying the correction fail and the correction does not solve the overrejection problem and means that one needs to know whether is close to 1. In general, if is far enough away from 1, the original size distortion problem is small as is the correction. Hence, such methods seem applicable only for “moderate” values of and it is difficult to tell the precise region where is both close enough to 1 that we need to correct the statistic, yet far enough away from 1 that the correction solves the overrejection problem.

Asymptotic approaches allow for smaller than 1 (but close to 1) as well as uniformly. Following this approach and employing the expression, Elliott and Stock (1994) use the local-to-unity representation to show that when , t-statistics converge to the limiting distribution

where is a standard (mixed) normal random variable that is independent of and has a nonstandard distribution equivalent to the distribution of a t-test testing that has a coefficient when regressing on and lagged changes that correct for the serial correlation in . For , this is the well-known Dickey–Fuller distribution for a t-statistic in a unit root test. This nonstandard distribution is a function of the local-to-unity parameter. Thus, when δ is small or 0, there is little size distortion in testing whether or not is a good predictor of however, for values of δ further away from we expect size distortions even asymptotically if is sufficiently close to 1.

Understanding whether or not this problem arises empirically is relatively straightforward since the extent of the problem depends on the parameters and the sample size. We can estimate consistently from the data by using OLS on and estimating the spectral density matrix using standard methods outlined in Newey and West (1987) or Andrews (1991). Estimates of the correlation from this estimated matrix are consistent, though biased towards 0. No consistent estimator is available for the local-to-unity parameter, However, we can still understand empirically whether there is a problem because we can construct confidence intervals on that are informative about which values of appear plausible. For any value of size distortions in conventional t-tests are largest when is near 1. For example, if , a nominal 10% equal-tailed test leads to lower-tail rejections of 11% and upper-tail rejections around 2%. When , the lower-tail rejections increase to 38% while the upper-tail rejections almost completely disappear—they are about one-tenth of 1%. As γ gets larger and decreases, these size distortions disappear.

Different approaches have been suggested for constructing confidence intervals on in (20.10). Cavanagh, Elliott, and Stock (1995) suggest a Bonferroni approach, using confidence intervals for along with the distribution in (20.13) to provide intervals for Campbell and Yogo (2006) show gains over this method from a similar approach that is based on a rotated version of (20.10). Jansson and Moreira (2006) suggest a method that conditions on a sufficient statistic for which results in a test that maximizes power conditional on this statistic. Since this test is similar for all values of by construction it is unbiased. However, this property comes at the cost of reduced power—Jansson and Moriera show that for many alternatives the Campbell and Yogo (2006) approach has better power. Elliott (2011) shows that additional stationary covariates that help explain the simultaneity leading to a nonzero can be used to mitigate or remove size distortions. Elliott, Müller, and Watson (2015) develop a method for constructing hypothesis tests on that is optimal against a point in the alternative space and controls size for all values of δ. These latter two methods have the property that they do not necessarily require to be near 1.

练习题

For the inflation rate forecasts, which model generally performs better for horizons of 1 through 3 quarters?

A. Random walk model
B. Random walk with drift model
C. AR(4) model
D. Autoregressive models

Which model produces the lowest RMSE values for the interest rate series at both short and long horizons?

A. AR(1) model
B. Random walk model
C. AR(4) model
D. Random walk with drift model

For the unemployment rate, which models' performance can be described as follows? (Select all that apply)

A. The AR(4) model dominates at all horizons
B. The random walk model produces RMSE values generally a bit higher than those from the AR models
C. The random walk with drift model performs very well
D. The AR(1) model is the best at long horizons

The random walk with drift model performs well for the series investigated in the empirical example because they are trending deterministically.

The best - performing model for forecasts depends on the true, but unknown, degree of persistence of the underlying series, which is consistent with our ___ results.

Explain why the random walk with drift model performs poorly for the series in the empirical example.

In the empirical example, for the inflation rate, which model is better at the 20 - quarter horizon?

A. AR(1) model
B. Random walk model
C. AR(4) model
D. AR model with lag length selected by AIC

Which of the following statements about the empirical example are correct? (Select all that apply)

A. The out - of - sample results are generated for the period 1970 - 2014
B. An expanding estimation window is used for generating the results
C. The forecast horizons include h = {1, 2, 3, 4, 8, 12, 16, 20} quarters
D. The in - sample results are generated for the period 1970 - 2014

The AR(4) model produces RMSE values generally lower than those from the random walk model for the interest rate series at short horizons.

What factors determine which model produces the best forecasts in the empirical example?

When considering the comparison of models in the empirical example, which concept from prior knowledge is relevant? The decision to impose a unit root or estimate model parameters is similar to the trade - off in the ___ case.

A. Estimating model parameters () instead of imposing unit root (kp_202_001_012)
B. Pre - test for and its recommendation (kp_202_001_013)
C. Comparison of risks for three strategies (kp_202_001_014)
D. Long - run forecast for stationary models with short memory (kp_202_2_1)

Which prior knowledge points are related to the concept of long - horizon forecasts in the empirical example? (Select all that apply)

A. Long - run forecast for unit root processes (kp_202_2_2)
B. Forecast horizon and sample size relationship (kp_202_2_3)
C. Divergence of MSE and rescaling (kp_202_2_4)
D. Estimating model parameters () instead of imposing unit root (kp_202_001_012)

The Granger Representation Theorem implications (kp_20_3_005) are relevant when considering the difference in forecasts between VAR and ECM specifications in the empirical example, as differences arise due to imposing versus not imposing restrictions. This statement is True or False?

In the multivariate case, similar to the univariate case, there is a trade - off in imposing restrictions. Abadir, Hadri, and Tzavalis (1999) examine the impact of estimating the coefficients on estimation error. This is related to the prior knowledge point about the ___ in imposing restrictions (kp_20_3_006).

Explain how the value of cointegrating vectors for forecasting (kp_20_3_007) is related to the empirical example. Consider a situation where is large and is very persistent.

登录后解锁笔记、知识点解析、AI 问答

立即登录