正在学习
17.4.3 Finite Sample Behavior of Tests
17.4.3 Finite Sample Behavior of Tests
The results described thus far are asymptotic, but the size of the forecast evaluation tests in finite samples has been examined in various Monte Carlo experiments. For situations where the forecast errors are given, differences across various test statistics have been analyzed in Harvey, Leybourne, and Newbold (1998). These authors do not replicate the forecast method but instead draw forecast errors directly as coming from a joint t-distribution with few degrees of freedom. For small samples they find that size distortions in (17.7) can be severe, although distributions must be quite fat tailed for this to hold.
West (2001) generates data from the model and compares this model (with an estimated effect of on against a model that uses to predict with independent of . The variables are drawn from an independent normal distribution using a variance–covariance matrix with diagonal elements (1, 1, 2). West examines the test statistic (17.8) as well as this same statistic with its standard error, S, replaced by the adjusted standard error indicated by the results of West and McCracken (1998). Standard errors change with the relative size of the evaluation sample to the estimation sample, , and depend on the method used to construct the forecasts. For this model (17.8) is theoretically oversized since the estimation error in constructing the forecasts adds a nonnegative component to the standard error. This is borne out by Monte Carlo results. The size of the corrected test statistic is well controlled as long as the evaluation sample is not too small; it is undersized for large estimation samples and oversized for small estimation samples.
For nested models, Clark and McCracken (2001) examine data-generating processes where are generated by a VAR(1) or VAR(2) with independent (across time and variables) residuals. The forecasting models are an autoregression (of correct order) in or the correctly specified To examine the size of the test, the VAR is chosen such that lags of have no predictive power. Using tests based on the regression (17.7), the ENC-T and ENC-NEW statistics with critical values constructed from their asymptotic results, Clark and McCracken find that the size of these tests is well controlled.
Clark and McCracken (2005a) propose a restricted VAR bootstrap to deal with cases where serial correlation and conditional heteroskedasticity in the forecast errors introduce nuisance parameters in the asymptotic distribution of the test statistic. The bootstrap uses a VAR specification for whose parameters are estimated using full-sample OLS on the restricted model. Resampled residuals from this model are used to recursively generate bootstrapped time series of . In turn these bootstrapped series are used to estimate the restricted and unrestricted forecasting models and the process is repeated a large number of times to compute the full distribution of test statistics bootstrapped under the null of no predictive power for the variables included only in the large model. Tests based on the (relative) forecast performance of the restricted and unrestricted models in the actual data are then compared with the percentiles of this distribution to compute -values.
The power of various forecast evaluation tests has also been examined in Monte Carlo experiments. For the nested case, Clark and McCracken (2001) use the Monte Carlo design above but allow lags of to affect . They find that the ENC-NEW test has higher power than the tests based on the regression (17.7) and the ENC-T statistics. They also show that a full sample Granger causality test has considerably more power than any of the out-of-sample tests. This is not surprising since we are examining tests statistics for which classical assumptions apply so full sample tests will generally have higher power.
Hansen and Timmermann (2015) establish analytical results for power in the case with two nested linear regression models whose parameters are recursively estimated. Their results show that, for a fixed length of the total sample size (T ), the power of out-of-sample tests decreases in the size of the data sample used for initial parameter estimation and conversely increases, the more data are used for out-of-sample forecast evaluation.
These out-of-sample tests have been extensively used to compare forecasts of stock returns (Lettau and Ludvigson, 2001; Hansen, Lunde, and Nason, 2011) and inflation and exchange rates (Rossi, 2013b).
17.4.4 Empirical Application
Table 17.2 reports the RMSE values associated with a rolling estimation window along with p-values for encompassing, the Diebold–Mariano and Giacomini–White tests for 11 univariate prediction models fitted to stock returns. Each model includes the predictor listed in the rows in the table, while the benchmark prevailing mean only includes an intercept. For the DM tests the parameters are estimated using an expanding window, while for the Giacomini–White test we use a rolling window with 20 years of monthly observations (120 data points). The out-of-sample evaluation period is 1970–2013.
Many of the models generate RMSE values comparable to those of the benchmark prevailing mean model (whose RMSE value is 4.4871) with the models based on the default spread (dfr), long term return (ltr), inflation (infl), and the term spread (tms) performing slightly better. However, neither the Diebold–Mariano test, nor the HLN or Giacomini–White test reject the null of equal predictive accuracy for any of the predictor variables.
练习题
In West's (2001) model , what is the relationship between the variables and ?
What is the effect of small evaluation samples on the corrected test statistic in West's model?
In Clark and McCracken's (2001) study, what is the predictive power of lags of in the VAR model?
Which of the following are true about the restricted VAR bootstrap proposed by Clark and McCracken (2005a)? (Select all that apply)
What are the implications of the adjusted MSE test proposed by Clark and West (2007)? (Select all that apply)
The size of the ENC-T and ENC-NEW statistics is well controlled when the VAR model is chosen such that lags of have no predictive power.
The adjusted MSE test statistic proposed by Clark and West (2007) is asymptotically normal under certain conditions.
In West's model, the standard errors change with the relative size of the evaluation sample to the estimation sample, denoted as , and depend on the method used to construct the forecasts. For this model, (17.8) is theoretically __________ since the estimation error in constructing the forecasts adds a nonnegative component to the standard error.
Clark and McCracken (2001) use tests based on the regression (17.7), the ENC-T and ENC-NEW statistics with critical values constructed from their __________ results.
Explain the purpose of the restricted VAR bootstrap proposed by Clark and McCracken (2005a).
登录后解锁笔记、知识点解析、AI 问答
立即登录