正在学习
16.6 CONCLUSION
16.6 CONCLUSION
It is common to evaluate sequences of forecasts generated either from survey data or from econometric models. Such evaluations differ from conventional specification analysis in econometrics in that they are typically conducted on out-of-sample forecasts. Out-of-sample forecasts restrict the information set—including the data used to estimate model parameters on—to the data available at the point in time where the forecasts are generated. As new data arrives, typically the parameters get reestimated using the most recent data. This out-of-sample approach is meant to simulate the forecasts generated by a forecaster in “real time” and to avoid look-ahead biases in so far as possible.
In situations where learning effects are important because the parameters of the forecasting model are updated recursively through time and the estimation window is not very large, care needs to be exercised when conducting inference about rationality for the underlying forecasting model. Results by West (1996) and West and McCracken (1998) address these issues, although they also impose tight restrictions on the types of models used to generate the forecasts.
Evaluation and Comparison of Multiple Forecasts
In many situations two or more competing forecasts of the same outcome are available and we are interested in comparing the forecasts. We could be interested in asking whether one forecast is better than the other forecast or whether they result in similar losses. If more than two forecasts are available, we can ask whether one particular forecast dominates all other forecasts.
In fact, we often have an embarrassment of riches when it comes to forecasts of the same outcome. Many of the forecast methods discussed in the previous chapters can be used to construct different forecasts of the same outcome even with the same data. Using different data sets further enriches the set of forecasts available for predicting an outcome. For example, ARIMA models or naive “no change” models are routinely used to generate baseline forecasts against which more sophisticated forecast methods can be compared.
Consider the case with m competing forecasts, each of which gives rise to a risk function which we can stack in an vector,
A common approach to choosing between forecasts is to select the method with the smallest expected loss. Other criteria could also be examined, such as stochastic dominance of one forecast over another. These questions can be examined using a sample analog of the risk, previously defined by
Since we have a vector of forecasts, this is now an vector of estimated risks and is a vector of the same dimension as the vector of forecasts, f .
Suppose we are comparing the performance of two or multiple forecasting methods and are interested in finding out whether one of the methods performed better than the other over some sample period. Provided that certain stationarity conditions hold for the sequence of loss differentials, we can use results in Diebold and Mariano (1995) which develop simple regression-based tests of equal predictive accuracy based on the mean difference in sample average losses.
In contrast, suppose we are interested in using the sequence of forecasts generated by two or multiple forecasting models to answer which model provides a better fit to the data. This requires taking into account that the forecasts in (17.2) are constructed from estimates, and hence estimation error in the construction of the forecasts may contribute to the sampling distribution of the average risk estimates. This is of course parallel to the case with a single forecast. When out-of-sample methods are used to generate the forecasts, the results of West (1996) and West and McCracken (1998) can then be applied since they were derived for vectors of losses. However, it is possible that two or more forecasts become asymptotically equivalent, resulting in a variance–covariance matrix of the limiting risks that is not of full rank. This happens if two or more models nest the true model. In this case the two models’ parameters converge to the same point and hence their forecasts become asymptotically equivalent. A second case arises with two linear forecasting models where the first model is the best predictor and the second model contains the same predictors plus additional variables that are irrelevant. The coefficient estimates on these irrelevant variables will, asymptotically, converge to 0, making the two models identical. West (1996) and West and McCracken (1998) rule out such possibilities and their results cannot be employed in this situation. Alternative approximations to the test statistics are instead required.
Comparisons of the predictive accuracy of two forecasting methods can lead to three outcomes. First, one forecast method may completely dominate the other method, in which case we would choose the dominant method. The notion that one method dominates another is typically labeled “encompassing.” Hoel (1947) seems to have first considered this problem and Chong and Hendry (1986) made significant advances to understanding encompassing tests. Second, one forecast may be best, but does not contain all useful (forecast-relevant) information available from the second forecast. In this case, we may not wish to discard the second forecast. Third, it is possible that the two forecast methods achieve the same expected loss, in which case we would be indifferent between the two methods and might want to use a combined forecast, as discussed in chapter 14.
This chapter first examines methods for comparing two forecasts, for which there are many results available in the literature. The two basic questions asked in such comparisons are, first, does one forecast dominate—or encompass, in the forecasting terminology—the other forecast? Second, are the losses generated by the two sequences of forecasts the same on average? These questions are covered in section 17.1 (forecast encompassing tests) and section 17.2 (tests of equivalent loss using the Diebold–Mariano test), respectively. Section 17.3 introduces an approach to comparing forecast methods suggested by Giacomini and White (2006) while section 17.4 discusses comparisons of forecasting performance with nested models. Section 17.5 extends the analysis to comparisons involving multiple forecasting models. Section 17.6 discusses ways to address data mining when multiple forecasting models are being compared, while section 17.7 covers recent stepwise methods for identifying superior forecasting models. Choice of sample split in outof-sample forecast evaluation experiments is discussed in section 17.8, while section 17.9 links forecast evaluation issues to questions from forecast combination and model selection and section 17.10 compares in-sample to out-of-sample methods. Section 17.11 concludes.
Throughout the chapter we will mostly assume a single-period forecast horizon and so set . For direct forecasting models the horizon makes little difference to most results and this assumption helps us keep notation simple.
练习题
What is the main purpose of out-of-sample forecasts in econometrics?
Which of the following is a key consideration when conducting inference about rationality in forecasting models with learning effects?
What are valid approaches to comparing multiple forecasts? (Select all that apply)
Out-of-sample forecasts use all available data, including future data, for parameter estimation.
The risk function for competing forecasts can be stacked in an vector represented as What does represent?
Explain the concept of stochastic dominance in the context of comparing forecasts.
Which of the following is a potential outcome when comparing the predictive accuracy of two forecasting methods?
What are sources of multiple forecasts for the same outcome? (Select all that apply)
Regression-based tests of equal predictive accuracy are based on the mean difference in sample average losses.
What is the sample analog of the risk function used for choosing between forecasts?
What does the Diebold and Mariano (1995) test compare?
What are the implications of two forecasts becoming asymptotically equivalent? (Select all that apply)
West and McCracken's results can be applied when two forecasts become asymptotically equivalent.
The notion that one forecast method completely dominates another is typically labeled as ___.
Explain the role of stationarity conditions in regression-based tests of equal predictive accuracy.
Which of the following is a reason for using out-of-sample methods in forecast evaluation?
What are the assumptions for West and McCracken's result? (Select all that apply)
The SUR estimator for assumes no serial correlation in forecast errors beyond that accounted for in .
The panel regression model for individual forecast errors is represented as where and . What does represent?
What is the role of the covariance matrix in the SUR estimator for ?
When comparing the predictive accuracy of two forecasting methods using out-of-sample data, which of the following is a necessary condition for applying the Diebold and Mariano (1995) test?
Which of the following statements are true regarding the comparison of multiple forecasts?
When forecasting models are updated recursively through time with a small estimation window, the results of West (1996) and West and McCracken (1998) can be directly applied without any modifications.
In comparing the predictive accuracy of two forecasting methods, if the sequence of loss differentials does not satisfy the stationarity conditions required by the Diebold and Mariano test, an alternative approach would be to use ___.
登录后解锁笔记、知识点解析、AI 问答
立即登录