正在学习
13.1 ROLE OF THE LOSS FUNCTION
13.1 ROLE OF THE LOSS FUNCTION
The conditional density in (13.2) is a population object and is constructed without reference to the loss function. It would seem that provision of a predictive density is superior to reporting a point forecast since it both (a) can be combined with a loss function to produce any point forecast; and (b) is independent of the loss function. In classical estimation of the predictive density, neither of these points really holds up in practice. First, given the predictive density it is possible to generate point forecasts. Unless the predictive distribution is parametric, however, there are practical issues with the presentation and communication of predictive densities to facilitate such calculations. Moreover, in the classical setting the estimated predictive distributions depend on the loss function. All parameters of the predictive density need to be estimated and these estimates require some loss function, so loss functions are thrown back into the mix. The catch here is that the loss functions that are often employed in density estimation do not line up with those employed for point forecasting which can lead to inferior point forecasts.
Even when decisions are ultimately based on a point forecast, the conditional density is still worth reporting if the forecast user’s loss function is unknown. Alternatively, the forecast may be intended for multiple users whose losses are known and could differ across end users. In these situations, the provision of the conditional density forecast can be viewed as the step before the loss function necessarily gets involved. Such a two-step procedure has obvious limitations, however. For example, the variables in the information set that are relevant could be different for different loss functions.
Example 13.1.1 (Optimal forecasts with dynamics in first and second moments under different loss functions). Suppose the density forecast is fully characterized by the conditional mean and volatility given current information, denoted by and so that
Under MSE loss, the optimal forecast would be the conditional mean
In contrast, under an asymmetric lin-lin loss function of the type
the optimal forecast takes the form
where is the quantile function of the innovation, η. The forecast under lin-lin loss generally differs from , assuming that α even when has a symmetric distribution. Hence a forecaster with squared error loss will only be interested in information that helps predict the conditional mean, whereas a forecaster with asymmetric lin-lin loss will want to make use of information on variables that help forecast the conditional volatility,
Moreover, conditional distributions are difficult to estimate well, and so point forecasts based on estimates of the conditional density may be highly suboptimal from an estimation perspective. It is also unclear what constitutes a “good” density forecast when we abstract from loss functions. Any estimate of the conditional density will have errors, and errors that are innocuous for some applications may be very important in other applications. Even if the conditional density avoids reference to the loss function, it is hard to imagine that optimal estimation of the density will be independent of the loss function. The loss function (or decision problem) will generally be central to the metric used to evaluate the estimator of the density.
From the perspective of generating point forecasts for a user with squared error loss, seemingly not much is lost by ignoring heteroskedasticity in the residuals. This is not quite true, however, since more efficient estimates of the parameters of the conditional mean can be obtained by accounting for time-varying heteroskedasticity. Moreover, in cases where forecasts require iterating multiple steps ahead on a model that involves nonlinear dynamics in the conditional mean, heteroskedasticity will directly affect the point forecast; see the discussion in chapter 8. In this chapter we mostly ignore such issues.
13.2 VOLATILITY MODELS
Often density modeling is separated into two steps: first construct a model for the conditional mean and conditional volatility. Then, in a second step, model the density of the residuals obtained after subtracting the conditional mean (or predicted value) and dividing by the conditional volatility. Different approaches for modeling the conditional mean have been described in chapters 6–11. This section describes a variety of approaches to modeling volatility dynamics, before approaches to modeling the density of the normalized residuals are covered in subsequent sections.1
练习题
Which of the following statements about predictive densities is correct?
Under MSE loss, what is the optimal forecast for ?
What are the implications of using different loss functions for forecasting? (Select all that apply)
The estimated predictive distributions in the classical setting are independent of the loss function.
The optimal forecast under asymmetric lin-lin loss is the same as the conditional mean when .
Under MSE loss, the optimal forecast is the ___.
The quantile function of the innovation, , is denoted as ___.
Why might a forecaster with asymmetric lin-lin loss be interested in information that helps predict the conditional volatility?
Which of the following is a reason for reporting the conditional density forecast even when decisions are based on a point forecast?
What are the advantages of density forecasts over point forecasts? (Select all that apply)
Point forecasts obtained as the solution to are independent of the loss function .
The basic density forecasting problem involves a single outcome variable, , and conditioning variables, ___.
Why might a portfolio manager be interested in the degree of uncertainty surrounding a point forecast of stock returns?
Which of the following statements about loss functions in binary forecasting are correct? (Select all that apply)
登录后解锁笔记、知识点解析、AI 问答
立即登录