正在学习
17.8 CHOICE OF SAMPLE SPLIT
17.8 CHOICE OF SAMPLE SPLIT
White’s Reality Check bootstrap approach can be used to address the effect on inference of searching for a superior model across many different model specifications. Another dimension along which forecasters could have searched for superior performance is the sample split, i.e., the decision how to split a sample with observations into an initial estimation and model selection period that uses the first observations, and a forecast evaluation period that uses the remaining observations. Instead of studying the specific starting date for the forecast evaluation period, , we can equivalently consider which is the fraction of the sample used for initial estimation. To obtain well-behaved test statistics, it is common to assume that the choice of sample split is limited to an interval , typically leaving out 10–15% of the sample on each side. This ensures sufficient data for initial model estimation and for subsequent forecast evaluation.
Hansen and Timmermann (2012) study the size of tests such as (17.28) when the -value is chosen as the smallest value associated with all possible split points, i.e.,
where is defined in (17.25), is a consistent estimate of the variance of the forecast errors, is the inverse of the c.d.f. of the test statistic . This is a quantile function and so gives the p-value of the test for a given value of λ. A single randomly selected -value will follow a uniform distribution but this is no longer the case when we consider the smallest value selected from a potentially large number of test statistics. Overlaps in the estimation and evaluation windows used by the tests, , introduce further complications since the individual -values will be strongly correlated. Thus, if the reported -value is the smallest value chosen from many possible sample splits, conventional inference will no longer be valid.
Hansen and Timmermann show that conventional tests of the predictive accuracy of nested models that account for nonstandard limiting distributions but ignore the effect of sample split mining can be greatly oversized. For example, with one additional predictor variable in the large model, a test with a nominal size of 5% can actually reject close to 15% of the time as a result of such sample split mining. Inflation in the size of the test grows even larger when the number of additional predictor variables in the large forecasting model is increased.
As a constructive way to account for the effect of sample split mining on inference, Hansen and Timmermann (2012) suggest an adjusted minimum -value test corresponding to the statistic in (17.46). The test is based on the statistic whose critical values can be derived using Monte Carlo simulations of the asymptotic distribution. Hansen and Timmermann compute adjusted -values from the quantiles of the sorted distribution of -values simulated from a large number of sample paths.
For example, for a test with a critical level of 5% and one additional predictor variable, the -value adjusted for sample split mining in the interval would have to fall below 1.3% for there to be significant evidence that the larger model dominates the small model. This result is intuitive: if only a single -value had been inspected, there would be no distortion to inference, and we could simply use the critical values reported by McCracken (2007). On the other hand, if the reported p-value is pre-selected from a large set of candidate p-values, then we require stronger evidence—a larger test statistic and a smaller p-value—to be confident in the strength of the evidence.
A similar message comes out of Rossi and Inoue (2012) who consider test statistics computed either from averages or from the maximum of test statistics across different sample splits in a manner analogous to Hansen and Timmermann (2012). Rossi and Inoue also provide sample split mining-robust test statistics and show through Monte Carlo simulations that conventional tests that ignore such mining can lead to large size distortions.
练习题
What is the purpose of White’s Reality Check bootstrap approach?
What does represent in the context of sample split?
Which of the following are reasons for limiting the choice of sample split to an interval ?
Conventional inference remains valid if the reported -value is the smallest value chosen from many possible sample splits.
Hansen and Timmermann show that conventional tests ignoring the effect of sample split mining can be greatly oversized.
The test statistic used by Hansen and Timmermann to account for the effect of sample split mining is based on the statistic whose critical values can be derived using __________ simulations.
For a test with a critical level of 5% and one additional predictor variable, the -value adjusted for sample split mining in the interval would have to fall below __________% for there to be significant evidence that the larger model dominates the small model.
Explain the impact of sample split mining on the size of the test according to Hansen and Timmermann.
How do Rossi and Inoue's findings relate to Hansen and Timmermann's study on test statistics?
Which of the following statements are true regarding the adjusted minimum -value test suggested by Hansen and Timmermann?
When conducting a sample split for forecast evaluation, what is the primary reason for limiting the choice of to an interval that typically leaves out 10-15% of the sample on each side?
Which of the following statements are true regarding the impact of sample split mining on test size and inference?
The adjusted minimum -value test statistic's critical values can be derived using Monte Carlo simulations of the asymptotic distribution, and this test is robust to the effect of sample split mining on inference.
Hansen and Timmermann show that conventional tests of the predictive accuracy of nested models that account for nonstandard limiting distributions but ignore the effect of ___, can be greatly oversized.
登录后解锁笔记、知识点解析、AI 问答
立即登录