正在学习
17.11 CONCLUSION
17.11 CONCLUSION
Forecasters often work with multiple competing models or observe multiple forecasts from survey data. In such a situation it is not only of interest to ask whether the individual forecasts are “optimal,” using the methods from chapters 15 and 16, but also to ask whether there exists a single dominant forecasting model, or perhaps a set of models that are better than the other ones (as in the model confidence set).
Evaluating Density Forecasts
Chapter 12 and 13 considered the situation where a predictive distribution, , rather than a point forecast, was provided. This chapter examines how we can evaluate such predictive densities. For a single density forecast this means evaluating whether the density is correctly specified and leads to good forecasting performance. If there are multiple density forecasts, we may be interested in comparing them and perhaps selecting the best. Although the ideas of density forecast evaluation mirror those for point forecasts, the literature on density forecast evaluation is less well developed.
The primary difficulty that arises when evaluating density forecasts is that we never actually observe the density of an outcome; only a single draw from the distribution is observed. Further complicating matters, when we observe outcomes over a period of time the predictor variables change and so we obtain density forecasts and outcomes conditional on different values of .
One way to evaluate density forecasts is to explicitly rely on a loss function that maps the density forecast and outcome to a single number that can be the basis for an examination of average loss. Such loss functions are typically known as scoring rules, especially in the statistics literature, and were discussed in chapters 2 and 13. Loss functions should be chosen to address the particular forecaster’s problem as different forecasters may use different loss functions. In practice, however, most studies rely on a small set of loss functions, the most popular being the log score, i.e., the average likelihood given certain distributional assumptions. Other methods that map the density forecast and outcome to a single number do not have an explicit loss function in mind but result in statistics that trade off different types of errors.
A second approach looks for desirable properties of the density forecast such as calibration or sharpness or matching the distribution function of the outcome. These features are often examined separately, with “good” density forecasts possessing as many desirable properties as possible.
Section 18.1 discusses density evaluation methods based on loss functions. Density forecasts are sometimes evaluated using features such as calibration, resolution, and sharpness. We cover this in section 18.2. Probability integral transform tests use information on the full density and are covered in section 18.3. Section 18.4 covers multicategory forecasts, while section 18.5 discusses how to evaluate interval forecasts, and section 18.6 concludes. Again, we focus on the case with a single-period forecast horizon, , to keep notation simple.
练习题
When evaluating density forecasts, what is the primary difficulty that arises?
Which of the following is a popular loss function used for evaluating density forecasts?
What are some desirable properties of density forecasts that can be examined?
The literature on density forecast evaluation is as well developed as the literature on point forecast evaluation.
When evaluating density forecasts, one approach is to rely on a loss function that maps the density forecast and outcome to a single number. Such loss functions are typically known as ___.
Explain why it is difficult to compare multiple density forecasts.
Which section discusses density evaluation methods based on loss functions?
What are some methods or features used to evaluate density forecasts, according to the text?
Probability integral transform tests use information on the full density to evaluate density forecasts.
What is the focus of section 18.5 in the context of density forecast evaluation?
When evaluating density forecasts, what does it mean for a forecast to be 'well-calibrated'?
The ___ is a popular loss function that represents the average likelihood given certain distributional assumptions.
How does the choice of loss function affect the evaluation of density forecasts?
Which of the following statements are true about evaluating density forecasts? (Select all that apply)
In-sample tests for model comparison generally have higher power than out-of-sample tests.
When evaluating multiple competing forecasting models, which of the following is a valid approach to determine if there exists a single dominant model? Assume that the models are being compared using out-of-sample forecasts to avoid data-mining concerns.
Which of the following are valid considerations when evaluating density forecasts? Select all that apply.
When comparing multiple density forecasts, it is sufficient to rely solely on the log score as a loss function to determine which forecast is the best.
The primary difficulty in evaluating density forecasts is that we never observe the full density of an outcome; instead, we only observe ___.
登录后解锁笔记、知识点解析、AI 问答
立即登录