正在学习

17.10 IN-SAMPLE VERSUS OUT-OF-SAMPLE FORECAST COMPARISON

17.10 IN-SAMPLE VERSUS OUT-OF-SAMPLE FORECAST COMPARISON

When we are interested in evaluating the expected loss or some other statistic that involves data generated by any of the above estimation schemes, the sampling behavior of the resulting test statistics must account for which estimation scheme was used by the forecasting model. It is less clear whether the “pseudo” out-of-sample approach based on any of the recursive schemes is better than simply basing the forecast evaluation on models that are estimated on the full data sample.

It might seem natural to mimic the forecaster’s problem when testing whether one model provides better forecasts than another model. However, a well-understood set of optimality results establishes that under quite general conditions, the optimal model comparison test should use the entire data set for model estimation. Typically the conditions required for this result are satisfied by the assumptions made to establish the sampling properties of out-of-sample test statistics. Hence, if the objective is to examine which forecasting model is better, the power of in-sample tests for the difference in the models’ performance is higher than the power of outof-sample tests. This point has been made in the forecasting context by Inoue and Kilian (2005) and, more recently, by Hansen and Timmermann (2015). Hansen and Timmermann (2015) demonstrate that commonly used out-of-sample test statistics for relative MSE performance are equivalent to the difference between two Wald statistics, one based on the full sample, the other based on only part of the sample. One implication of this is that out-of-sample forecast evaluation tests can imply a substantial loss of power compared to full-sample tests that do not split the sample in this manner.

When the forecast evaluation tests are used to choose between different methods for constructing a forecasting model, the problem becomes essentially a variation on model selection. As we saw in chapter 6, there is no optimal approach to model selection, so it is possible that different methods have different, possibly desirable, properties in different circumstances. However, this issue has not yet been examined in the literature.

Tests for equally precise forecasts may be better examined by means of fullsample regressions rather than using the pseudo out-of-sample experiments that underline many of the methods detailed above. The null and alternative models are presented in terms of model parameters with a large number of regularity conditions on the data and models which are often tested more powerfully using the full sample.

Why are out-of-sample tests then so popular in the forecasting literature? One reason is the (healthy) skepticism among economists when presented with a forecast model that provides a suspiciously impressive in-sample fit. As we saw in the previous chapter, in-sample results tend to exaggerate a model’s likely future forecasting performance. This holds even if only a single forecast model is considered.

If, in fact, multiple models were considered and only the best model’s performance gets presented, such data mining can greatly exaggerate that model’s true expected forecasting performance. In some situations, out-of-sample forecasts can be used to address this data-mining concern. Specifically, if it can be certified that the forecaster did not in any way use the out-of-sample data to select and estimate a model, this evaluation sample is not contaminated and so can be used as a relatively “clean” sample on which to evaluate the model. There are, however, many ways in which the evaluation sample can get contaminated. For example, suppose we estimate a prediction model for US T-bill rates based on data up to 2000, but we include variables that we know (based on subsequent information) were correlated with monetary policy during the financial crisis of 2008–2009. This benefit of hindsight could in turn favor the selected model in a way that is unlikely to reflect its future predictive performance. Moreover, if a researcher splits the data into separate estimation and evaluation samples, but studies different models’ performance over the evaluation sample, this would again mean that the best model’s forecasting performance is no longer a single independent draw from a fresh data sample.

Such practice is a concern for any pseudo out of-sample test and blurs the line between in-sample and out-of-sample forecast evaluation tests. However, it does not remove the practical point that the effect of data-mining can be far stronger for in-sample tests—since the model parameters are being tailored to the evaluation sample—than for out-of-sample tests.

A second reason for continuing with the out-of-sample approach is the belief that the conditions under which the in-sample tests are optimal are invalid. For example, the process generating the outcome might not be stationary and hence the correct forecasting model might not be constant over time. Examining the out-of-sample performance of the models might suggest a model that performs better more often than an in-sample test that—due to the failure of assumptions underlying the construction of the null distribution—might not control the test’s size and hence performs poorly. While this argument has its merits, unfortunately the construction of null distributions for the out-of-sample tests also requires fairly strong assumptions and hence these tests could also fail to properly control size in such situations.

A third reason for using out-of-sample tests is that these allow us to address whether a forecasting model is practically useful even after accounting for recursive estimation error. Suppose the question is not whether the parameters of a large prediction model are different from 0 in the limit, but whether the associated forecasts are better than those generated by another, possibly nested, model. This is the question that is relevant in many forecasting situations and one that can be addressed using the out-of-sample forecast evaluation framework of Giacomini and White (2006).

练习题

According to the optimality results, under general conditions, which data set should be used for model estimation in an optimal model comparison test?

A. A subset of the full data set
B. The full data set
C. Only out-of-sample data
D. Randomly selected data points

What is the relationship between the power of in-sample tests and out-of-sample tests for model comparison?

A. In-sample tests have lower power than out-of-sample tests
B. In-sample tests have equal power to out-of-sample tests
C. In-sample tests have higher power than out-of-sample tests
D. The power relationship depends on the model complexity

What are commonly used out-of-sample test statistics for relative MSE performance equivalent to?

A. The sum of two Wald statistics
B. The difference between two Wald statistics
C. The product of two Wald statistics
D. The ratio of two Wald statistics

Out-of-sample forecast evaluation tests imply a substantial gain of power compared to full-sample tests.

When forecast evaluation tests are used for model selection, it becomes a variation on model selection.

Tests for equally precise forecasts are better examined using pseudo out-of-sample experiments.

In-sample results tend to exaggerate a model’s likely ___ forecasting performance.

If multiple models were considered and only the best model’s performance gets presented, such data mining can greatly exaggerate that model’s true expected forecasting performance, which is known as the effect of ___.

Explain why out-of-sample forecasts address the data-mining concern.

What is one reason for the popularity of out-of-sample tests in the forecasting literature despite their potential power loss?

Which of the following are implications of data-mining on in-sample tests? (Select all that apply)

A. Data-mining can exaggerate true expected forecasting performance
B. The effect of data-mining is weaker for in-sample tests than for out-of-sample tests
C. The effect of data-mining can be far stronger for in-sample tests
D. In-sample tests are immune to data-mining effects

What are some reasons for continuing with the out-of-sample approach despite its drawbacks? (Select all that apply)

A. Belief that in-sample test conditions are invalid
B. Out-of-sample tests are easier to implement
C. The process generating the outcome might not be stationary
D. Out-of-sample tests always control size properly

How do White’s Reality Check and the Bonferroni bound address the issue of multiple model comparison? Provide a brief explanation for each.

When comparing forecasting models, which approach is generally more powerful for detecting differences in model performance according to optimality results?

A. Out-of-sample tests using recursive schemes
B. In-sample tests using the full data set
C. Pseudo out-of-sample experiments
D. Tests based on overlapping windows

Which statement best describes the relationship between out-of-sample forecast evaluation tests and full-sample tests in terms of statistical power?

A. Out-of-sample tests generally have equal power to full-sample tests
B. Out-of-sample tests have higher power when the process is stationary
C. Out-of-sample tests imply a substantial loss of power compared to full-sample tests
D. Full-sample tests are only valid for linear models

Tests for equally precise forecasts are better conducted using full-sample regressions rather than pseudo out-of-sample experiments because the latter may not properly control for model search effects.

The ___ approach is used to correct for the effect of searching across multiple models when selecting the best forecasting model, taking into account that the best model's performance might be exaggerated due to data mining.

登录后解锁笔记、知识点解析、AI 问答

立即登录