正在学习

13.1 ROLE OF THE LOSS FUNCTION

13.1 ROLE OF THE LOSS FUNCTION

The conditional density in (13.2) is a population object and is constructed without reference to the loss function. It would seem that provision of a predictive density is superior to reporting a point forecast since it both (a) can be combined with a loss function to produce any point forecast; and (b) is independent of the loss function. In classical estimation of the predictive density, neither of these points really holds up in practice. First, given the predictive density it is possible to generate point forecasts. Unless the predictive distribution is parametric, however, there are practical issues with the presentation and communication of predictive densities to facilitate such calculations. Moreover, in the classical setting the estimated predictive distributions depend on the loss function. All parameters of the predictive density need to be estimated and these estimates require some loss function, so loss functions are thrown back into the mix. The catch here is that the loss functions that are often employed in density estimation do not line up with those employed for point forecasting which can lead to inferior point forecasts.

Even when decisions are ultimately based on a point forecast, the conditional density is still worth reporting if the forecast user’s loss function is unknown. Alternatively, the forecast may be intended for multiple users whose losses are known and could differ across end users. In these situations, the provision of the conditional density forecast can be viewed as the step before the loss function necessarily gets involved. Such a two-step procedure has obvious limitations, however. For example, the variables in the information set that are relevant could be different for different loss functions.

Example 13.1.1 (Optimal forecasts with dynamics in first and second moments under different loss functions). Suppose the density forecast is fully characterized by the conditional mean and volatility given current information, denoted by and so that

Under MSE loss, the optimal forecast would be the conditional mean

In contrast, under an asymmetric lin-lin loss function of the type

the optimal forecast takes the form

where is the quantile function of the innovation, η. The forecast under lin-lin loss generally differs from , assuming that α even when  has a symmetric distribution. Hence a forecaster with squared error loss will only be interested in information that helps predict the conditional mean, whereas a forecaster with asymmetric lin-lin loss will want to make use of information on variables that help forecast the conditional volatility,

Moreover, conditional distributions are difficult to estimate well, and so point forecasts based on estimates of the conditional density may be highly suboptimal from an estimation perspective. It is also unclear what constitutes a “good” density forecast when we abstract from loss functions. Any estimate of the conditional density will have errors, and errors that are innocuous for some applications may be very important in other applications. Even if the conditional density avoids reference to the loss function, it is hard to imagine that optimal estimation of the density will be independent of the loss function. The loss function (or decision problem) will generally be central to the metric used to evaluate the estimator of the density.

From the perspective of generating point forecasts for a user with squared error loss, seemingly not much is lost by ignoring heteroskedasticity in the residuals. This is not quite true, however, since more efficient estimates of the parameters of the conditional mean can be obtained by accounting for time-varying heteroskedasticity. Moreover, in cases where forecasts require iterating multiple steps ahead on a model that involves nonlinear dynamics in the conditional mean, heteroskedasticity will directly affect the point forecast; see the discussion in chapter 8. In this chapter we mostly ignore such issues.

13.2 VOLATILITY MODELS

Often density modeling is separated into two steps: first construct a model for the conditional mean and conditional volatility. Then, in a second step, model the density of the residuals obtained after subtracting the conditional mean (or predicted value) and dividing by the conditional volatility. Different approaches for modeling the conditional mean have been described in chapters 6–11. This section describes a variety of approaches to modeling volatility dynamics, before approaches to modeling the density of the normalized residuals are covered in subsequent sections.1

练习题

Which of the following statements about predictive densities is correct?

A. Predictive densities are always parametric and easy to communicate.
B. Predictive densities can be combined with any loss function to produce a point forecast.
C. Classical estimation of predictive densities does not depend on the loss function.
D. The optimal forecast under MSE loss is the conditional volatility.

Under MSE loss, what is the optimal forecast for ?

A.
B.
C.
D.

What are the implications of using different loss functions for forecasting? (Select all that apply)

A. The optimal forecast under asymmetric lin-lin loss differs from the conditional mean.
B. The optimal forecast under MSE loss is the conditional volatility.
C. A forecaster with squared error loss will focus on information that helps predict the conditional mean.
D. A forecaster with asymmetric lin-lin loss will focus on information that helps predict the conditional volatility.

The estimated predictive distributions in the classical setting are independent of the loss function.

The optimal forecast under asymmetric lin-lin loss is the same as the conditional mean when .

Under MSE loss, the optimal forecast is the ___.

The quantile function of the innovation, , is denoted as ___.

Why might a forecaster with asymmetric lin-lin loss be interested in information that helps predict the conditional volatility?

Which of the following is a reason for reporting the conditional density forecast even when decisions are based on a point forecast?

A. The forecast user’s loss function is known.
B. The forecast is intended for a single user.
C. The forecast user’s loss function is unknown.
D. The predictive distribution is parametric.

What are the advantages of density forecasts over point forecasts? (Select all that apply)

A. Density forecasts convey the precision of the forecast.
B. Density forecasts are sufficient for all users with different loss functions.
C. Policy makers need full distribution forecasts to consider the distribution of possible outcomes.
D. Density forecasts are easier to communicate than point forecasts.

Point forecasts obtained as the solution to are independent of the loss function .

The basic density forecasting problem involves a single outcome variable, , and conditioning variables, ___.

Why might a portfolio manager be interested in the degree of uncertainty surrounding a point forecast of stock returns?

Which of the following statements about loss functions in binary forecasting are correct? (Select all that apply)

A. Loss functions in binary forecasting are very simple because there are only four outcomes.
B. Under binary loss, the conditional mean estimate is a point forecast.
C. Transforming a distributional forecast to a point forecast depends on whether the probability of an outcome is above a cutoff determined by the loss function.
D. Loss functions for point and distributional forecasts have no relationship.

登录后解锁笔记、知识点解析、AI 问答

立即登录