正在学习
1.1.3 Part III
1.1.3 Part III
The third part of the book deals with forecast evaluation methods. Evaluation of forecast methods is central to the forecasting problem and the difficulties involved in this step explain both the plethora of methods suggested for forecasting any particular outcome and the need for careful evaluation of forecasting methods.
To see the central issue, consider the simple problem of forecasting the next outcome, , in a sequence of independently and identically distributed data with mean , variance , and no explanatory variables. It is well known that under mean squared error (MSE) loss the best forecast is an estimate of the mean, such as the sample mean . Since the outcome is a random variable whose distribution is centered on , the forecast is typically different from the outcome even if we had a perfect estimate of , if we knew , as long as . Observing a single outcome far away from the forecast is therefore not necessarily indicative of a poor forecast. More generally, methods for forecast evaluation have to deal with the fact that (in expectation) the average in-sample loss and the average out-of-sample loss differ. To see this, suppose we use the sample mean as our forecast. For any in-sample observation, , the MSE of the forecast (or fitted value) is
Here the third term in the second line comes from the cross product when we compute the squared terms in the first line.
In contrast, the MSE of out-of-sample forecasts of is
Here there is no cross-product term. Comparing these two expressions, we see that estimation error reduces the in-sample MSE but increases the out-of-sample MSE. In both cases the terms are of order and so the difference disappears asymptotically. However, in many forecasting problems this smaller-order term is important both statistically and economically. When we consider many different models of the outcome, differences in the MSE across models are of the same order as the effects on estimation error. This makes it difficult to distinguish between models and is one reason why model selection is so difficult. The insight that the in-sample fit improves by using overparameterized models, whereas out-of-sample predictive accuracy can be reduced by using such models, strongly motivates the use of out-of-sample evaluation methods, although caveats apply as we discuss in part III of the book.
In the past 20 years many new forecast evaluation methods have been developed. Prior to this development, most academic work on evaluation and ranking of forecasting performance paid very little attention to the consideration that forecasts were obtained from recursively estimated models. Thus, often studies used the sample mean squared forecast error, computed for a particular empirical data set, to give an estimate of a model’s performance without accompanying standard errors. An obvious limitation of this approach is that such averages often are averages over very complicated functions of the data. Through their dependence on estimated parameters these averages are also typically correlated across time in ways that give rise to quite complicated distributions for standard test statistics. For some of the simpler ways that forecasts could have been generated recursively, recent papers derive the resulting standard errors, although much more work remains to be done to extend results to many of the popular forecasting methods used in practice.
Chapter 15 first establishes the properties that a good forecast should have in the context of the underlying loss function and discusses how these properties can be tested in practice. The chapter goes from the case where very little structure can be imposed on the loss function to cases where the loss function is known up to a small set of parameters. In the latter case it can be tested that the derivative of the loss with respect to the forecast, the so-called generalized forecast error, is unpredictable given current information. The chapter also shows how assumptions about the loss function can be traded off against testable assumptions on the underlying datagenerating process.
Chapter 16 gives an overview of basic issues in evaluating forecasts, along with a description of informal methods. This chapter examines the evaluation of a sequence of forecasts from a single model. Critical values for the tests of forecast efficiency depend on how the forecast was constructed, specifically whether a fixed, rolling, or expanding estimation window was used.
Chapter 17 extends the assessment of the predictive performance of a single model to the situation with more than one forecast to examine and so addresses the issue of which, if any, forecasting method is best. We review ways to compare the forecasting methods and strategies for testing hypotheses useful to identifying methods that work well in practice. Special attention is paid to the case with nested forecasting models, i.e., cases where one model includes all the terms of another benchmark model plus some additional information. We distinguish between tests of equal predictive accuracy and tests of forecast encompassing, the latter case referring to situations where one forecast dominates another. We also discuss how to test whether the best among many (possibly thousands) of forecasts is genuinely better than some benchmark.
Chapter 18 examines the evaluation of distributional forecasts. A complication that arises is that we never observe the density of the outcome; only a single draw from the distribution gets observed. Various approaches have been suggested to deal with this issue, including logarithmic scores and probability integral transforms. We discuss these as well as ways to evaluate whether the basic features of a density forecast match the data.
练习题
Under mean squared error (MSE) loss, what is the best forecast for independently and identically distributed (i.i.d.) data with mean ?
Why can the forecast be different from the outcome even with a perfect estimate of for i.i.d. data?
Which of the following statements about in - sample and out - of - sample MSE are correct?
The difference between in - sample and out - of - sample MSE due to estimation error disappears asymptotically as the sample size gets larger.
Using overparameterized models always improves the in - sample fit and out - of - sample predictive accuracy.
The in - sample MSE formula simplifies to ___.
The out - of - sample MSE formula simplifies to ___.
Explain why model selection is difficult in forecasting problems.
What motivates the use of out - of - sample evaluation methods in forecasting?
Which of the following are limitations of past forecast evaluation methods? (Select all that apply)
When evaluating forecasts using mean squared error (MSE) loss, which of the following statements is true about the in-sample and out-of-sample MSE for i.i.d. data with mean and variance ? Assume the forecast is the sample mean .
Which of the following are reasons why out-of-sample evaluation methods are strongly motivated in forecasting? (Select all that apply)
True or False: The difference between in-sample and out-of-sample MSE disappears asymptotically as , but the smaller-order terms (of order ) can still be important in practice.
Under mean squared error (MSE) loss, the best forecast for i.i.d. data with mean is an estimate of the mean, such as the ___.
登录后解锁笔记、知识点解析、AI 问答
立即登录