正在学习
20.3.1 Empirical Example
20.3.1 Empirical Example
As an empirical illustration, we consider long-run forecasts of the quarterly inflation rate, i.e., the log-first difference in the CPI, the unemployment rate, and the interest rate. These time series are quite persistent as shown in previous chapters. We show results for forecast horizons h = {1, 2, 3, 4, 8, 12, 16, 20} quarters and compare forecasts from a random walk, random walk with drift, an AR model with lag length selected by the AIC, and forecasts from AR(1) and AR(4) models. The out-ofsample results reported in table 20.1 are generated for the period 1970–2014 using an expanding estimation window.
First consider the inflation rate. Here the results show that the autoregressive models generally perform better than the random walk model for horizons of 1 through 3 quarters, but that the random walk model is better at the longest 20 quarter horizon. The random walk model with drift produces very poor forecasts, especially at long horizons. For the unemployment rate the AR(4) model dominates at all horizons and the random walk model produces RMSE values that are generally a bit higher than those from the AR models. Conversely, the random walk model produces the lowest RMSE values for the interest rate series at both short and long horizons.
In all cases the random walk with drift model performs very poorly because the series investigated here are not trending deterministically in any particular direction. Whether the simple random walk specification or the autoregressive specification is best varies, however, across variables and across sample periods. Which model produces the best forecasts will depend on the true, but unknown, degree of persistence of the underlying series; this is consistent with our theoretical results.
20.4 FORECASTING WITH PERSISTENT REGRESSORS
A common forecasting situation involves a regressor that displays persistence — the regressor is reverting to its mean only very slowly, if at all—but appears to have predictive power based on tests for the exclusion of that regressor in a linear regression. Examples include using the forward premium to predict changes in exchange rates (Bilson, 1981), short-run interest rates, consumption, and stock prices for forecasting changes in income (Chen, 1991; Hall, 1978), or the dividend–price ratio or earnings–price ratio for forecasting future stock returns (Stambaugh, 1999; Valkanov, 2003). In many cases, variation in the outcome variable captured by such predictor variables tends to be modest, although often economically interesting, as in the case of a small, but highly persistent risk premium on a financial asset.
TABLE 20.1:
Root mean squared forecast errors (RMSE) for different forecasting models. The table reports RMSE values for the quarterly inflation, unemployment and interest rate series at different forecast horizons ranging from 1 through 20 quarters. The forecast evaluation period is 1970– 2014; parameters are estimated using a recursively expanding estimation window for the models with unknown parameters.
| Steps | RW | RW drift | AR(AIC) | AR(1) | AR(4) |
| Inflation rate | |||||
| 1Q | 4.144 | 4.513 | 3.316 | 3.626 | 3.432 |
| 2Q | 4.560 | 5.776 | 3.488 | 3.803 | 3.477 |
| 3Q | 4.196 | 6.760 | 3.543 | 3.752 | 3.467 |
| 4Q | 3.653 | 7.912 | 3.735 | 3.825 | 3.487 |
| 8Q | 4.277 | 14.532 | 4.243 | 4.239 | 4.034 |
| 12Q | 4.229 | 21.210 | 4.361 | 4.384 | 4.286 |
| 16Q | 4.370 | 27.952 | 4.365 | 4.401 | 4.346 |
| 20Q | 4.118 | 34.777 | 4.188 | 4.221 | 4.180 |
| Unemployment rate | |||||
| 1Q | 0.358 | 0.400 | 0.279 | 0.365 | 0.277 |
| 2Q | 0.655 | 0.744 | 0.562 | 0.663 | 0.557 |
| 3Q | 0.915 | 1.059 | 0.843 | 0.919 | 0.831 |
| 4Q | 1.140 | 1.349 | 1.110 | 1.138 | 1.089 |
| 8Q | 1.759 | 2.343 | 1.725 | 1.701 | 1.698 |
| 12Q | 2.040 | 3.213 | 1.899 | 1.923 | 1.872 |
| 16Q | 2.107 | 4.026 | 1.956 | 1.989 | 1.937 |
| 20Q | 2.081 | 4.872 | 1.990 | 2.006 | 1.985 |
| Interest rate | |||||
| 1Q | 1.164 | 1.207 | 1.398 | 1.201 | 1.269 |
| 2Q | 1.516 | 1.652 | 1.680 | 1.564 | 1.649 |
| 3Q | 1.632 | 1.926 | 1.762 | 1.643 | 1.766 |
| 4Q | 1.975 | 2.403 | 2.204 | 1.982 | 2.108 |
| 8Q | 2.876 | 3.988 | 3.082 | 2.809 | 3.083 |
| 12Q | 3.351 | 5.302 | 3.662 | 3.381 | 3.731 |
| 16Q | 3.506 | 6.442 | 4.122 | 3.632 | 4.145 |
| 20Q | 3.403 | 7.521 | 4.457 | 3.898 | 4.565 |
To capture this setting, consider the model
TABLE 20.2:
Size properties for one-sided test.
| 8= | -0.95 | -0.75 | -0.5 |
| 0.416 | 0.289 | 0.180 | |
| 0.139 | 0.116 | 0.091 | |
| 0.105 | 0.092 | 0.079 |
Note: Size is 5% for a one-sided upper-tail test.
The regressor, follows the same process as in (20.2). Assumptions on are the same as in the earlier section. Regressions such as (20.10) typically do not allow to be serially correlated. If were serially correlated, we would want to model this correlation to improve predictability. Moreover, in empirical applications there is usually little or no serial correlation to be found. The terms and are assumed to be deterministic, with
We further assume that the partial sums of the residuals converge to a bivariate Brownian motion process with variance–covariance matrix, , given by10
δ is the correlation between and . This is also the long-run correlation between these shocks when is serially correlated but globally covariance stationary.
Monte Carlo simulation results have lead a number of authors to note that ttests of the hypothesis versus are not well approximated by the standard normal distribution when is close to 1 and δ is far from 0. For , the test underrejects in the upper tail and overrejects in the lower tail; the opposite holds for see Mankiw and Shapiro (1986), Elliott and Stock (1994), and Stambaugh (1999). Table 20.2 shows upper-tail rejections for tests with a nominal 5% level for and . The values for are selected to be relevant choices for forecasting changes in stock returns with the dividend–price ratio (Stambaugh, 1999).
Econometric analysis has explained such findings either as a small sample bias problem that disappears asymptotically or by modeling as being local-to-unity and showing that there is always a region for in which such tests will overreject, regardless of the sample size. Under both approaches, we have a SUR model for , and we can rotate so , where is orthogonal to . If there are no deterministic terms in the model, this allows us to write the OLS estimate of as
where is the OLS estimate of obtained from an AR(1) regression on . The intuition behind the “small sample” approach becomes clear from. The second line has a sampling distribution centered on 0. The OLS coefficient is downward biased for so if then will be upwards biased, as found in the simulations discussed above. The extent of the bias depends on the nature of the deterministic terms and is greater when constants are estimated in the regressions. Stambaugh (1999) presents analytical results for the size of the bias when Amihud and Hurvich (2004) suggest a correction for the bias, and base tests on their corrected estimator. Each of these methods requires that . The problem with this approach is that for roots “close” to 1, the assumptions underlying the correction fail and the correction does not solve the overrejection problem and means that one needs to know whether is close to 1. In general, if is far enough away from 1, the original size distortion problem is small as is the correction. Hence, such methods seem applicable only for “moderate” values of and it is difficult to tell the precise region where is both close enough to 1 that we need to correct the statistic, yet far enough away from 1 that the correction solves the overrejection problem.
Asymptotic approaches allow for smaller than 1 (but close to 1) as well as uniformly. Following this approach and employing the expression, Elliott and Stock (1994) use the local-to-unity representation to show that when , t-statistics converge to the limiting distribution
where is a standard (mixed) normal random variable that is independent of and has a nonstandard distribution equivalent to the distribution of a t-test testing that has a coefficient when regressing on and lagged changes that correct for the serial correlation in . For , this is the well-known Dickey–Fuller distribution for a t-statistic in a unit root test. This nonstandard distribution is a function of the local-to-unity parameter. Thus, when δ is small or 0, there is little size distortion in testing whether or not is a good predictor of however, for values of δ further away from we expect size distortions even asymptotically if is sufficiently close to 1.
Understanding whether or not this problem arises empirically is relatively straightforward since the extent of the problem depends on the parameters and the sample size. We can estimate consistently from the data by using OLS on and estimating the spectral density matrix using standard methods outlined in Newey and West (1987) or Andrews (1991). Estimates of the correlation from this estimated matrix are consistent, though biased towards 0. No consistent estimator is available for the local-to-unity parameter, However, we can still understand empirically whether there is a problem because we can construct confidence intervals on that are informative about which values of appear plausible. For any value of size distortions in conventional t-tests are largest when is near 1. For example, if , a nominal 10% equal-tailed test leads to lower-tail rejections of 11% and upper-tail rejections around 2%. When , the lower-tail rejections increase to 38% while the upper-tail rejections almost completely disappear—they are about one-tenth of 1%. As γ gets larger and decreases, these size distortions disappear.
Different approaches have been suggested for constructing confidence intervals on in (20.10). Cavanagh, Elliott, and Stock (1995) suggest a Bonferroni approach, using confidence intervals for along with the distribution in (20.13) to provide intervals for Campbell and Yogo (2006) show gains over this method from a similar approach that is based on a rotated version of (20.10). Jansson and Moreira (2006) suggest a method that conditions on a sufficient statistic for which results in a test that maximizes power conditional on this statistic. Since this test is similar for all values of by construction it is unbiased. However, this property comes at the cost of reduced power—Jansson and Moriera show that for many alternatives the Campbell and Yogo (2006) approach has better power. Elliott (2011) shows that additional stationary covariates that help explain the simultaneity leading to a nonzero can be used to mitigate or remove size distortions. Elliott, Müller, and Watson (2015) develop a method for constructing hypothesis tests on that is optimal against a point in the alternative space and controls size for all values of δ. These latter two methods have the property that they do not necessarily require to be near 1.
练习题
For the inflation rate forecasts, which model generally performs better for horizons of 1 through 3 quarters?
Which model produces the lowest RMSE values for the interest rate series at both short and long horizons?
For the unemployment rate, which models' performance can be described as follows? (Select all that apply)
The random walk with drift model performs well for the series investigated in the empirical example because they are trending deterministically.
The best - performing model for forecasts depends on the true, but unknown, degree of persistence of the underlying series, which is consistent with our ___ results.
Explain why the random walk with drift model performs poorly for the series in the empirical example.
In the empirical example, for the inflation rate, which model is better at the 20 - quarter horizon?
Which of the following statements about the empirical example are correct? (Select all that apply)
The AR(4) model produces RMSE values generally lower than those from the random walk model for the interest rate series at short horizons.
What factors determine which model produces the best forecasts in the empirical example?
When considering the comparison of models in the empirical example, which concept from prior knowledge is relevant? The decision to impose a unit root or estimate model parameters is similar to the trade - off in the ___ case.
Which prior knowledge points are related to the concept of long - horizon forecasts in the empirical example? (Select all that apply)
The Granger Representation Theorem implications (kp_20_3_005) are relevant when considering the difference in forecasts between VAR and ECM specifications in the empirical example, as differences arise due to imposing versus not imposing restrictions. This statement is True or False?
In the multivariate case, similar to the univariate case, there is a trade - off in imposing restrictions. Abadir, Hadri, and Tzavalis (1999) examine the impact of estimating the coefficients on estimation error. This is related to the prior knowledge point about the ___ in imposing restrictions (kp_20_3_006).
Explain how the value of cointegrating vectors for forecasting (kp_20_3_007) is related to the empirical example. Consider a situation where is large and is very persistent.
登录后解锁笔记、知识点解析、AI 问答
立即登录