正在学习

2.2.1.6 Binary Loss

2.2.1.6 Binary Loss

When the space of outcomes Y is discrete, the forecast errors typically take on only a small number of possible values. Hence in constructing a loss function for such problems, all that is required is to evaluate each of a small number of possibilities. The simplest case arises when forecasting a binary outcome so that or . In this case there are only four possible pairings of the point forecast and outcome: two where the forecast gives the correct outcome and two errors. If we restrict the loss function to not depend on Z (this case is examined below) and also restrict the problem so that a correct forecast has the same value regardless of the value for Y, then the binary loss function can be written

Here we have set the loss from a correct prediction to 0 and normalized the losses from an incorrect forecast to sum to 1 by dividing by their sum; see Schervish (1989), Boyes, Hoffman, and Low (1989), Granger and Pesaran (2000), and Elliott and Lieli (2013).

For (2.18) to be a valid loss function, we require that . This ensures that the properties of the loss function listed in (2.3) hold. Notice that the binary loss function can be written as , since the loss is equal to

2.2.2 Level- and Forecast-Dependent Loss Functions

Economic loss is mostly assumed to depend on only the forecast error, This is too restrictive an assumption for situations in which the forecaster’s objective function depends on state variables such as the level of the outcome variable Y. More generally, we can consider loss functions of the form . The most common level-dependent loss function is the mean absolute percentage error (MAPE), given by

Since the forecast and forecast error have the same units as the outcome, the MAPE is a unitless loss function. This is considered to be an advantage when constructing the sample analog of this loss function and employing it to evaluate forecast methods across outcomes measured in different units. If the loss function is well grounded in terms of the actual costs arising from the forecasting problem, dependence on units does not seem to be an important issue—comparisons across different forecasts with different units should be related not through some arbitrary adjustment but instead in a way that trades off the costs associated with the forecast errors for each of the outcomes. This is achieved by the multivariate loss functions examined in the next section.

Scaling the forecast error by the outcome in (2.19) has the effect of weighting forecast errors more heavily when y is near 0 than when y is far from 0. This is difficult to justify in many applications. Moreover, if the predictive density for Y has nontrivial mass at 0, then the expected loss is unlikely to exist, hence invalidating many of the results from decision theory for this case. Nonetheless, MAPE loss remains popular in many practical forecast evaluation experiments.

More generally, level- and forecast-dependent loss functions can be written as but do not reduce to or . Although loss functions in this class are not particularly common, there are examples of their use. For example, Bregman (1967) suggested loss functions of the form

where is a strictly convex function, so . Squared error loss is nested as a special case of 2.20.

Differentiating (2.20) with respect to the forecast, f , we get

which generally depends on both y and f . This, along with the assumption that , ensures that the conditional mean is the optimal forecast. Bregman loss is further discussed in Patton (2015).

In an empirical application of level-dependent loss, Patton and Timmermann (2007b) find that the Federal Reserve’s forecasts of output growth fail to be optimal if their loss is restricted to depend only on the forecast error. Rationalizing the Federal Reserve’s forecasts requires not only that overpredictions of output growth are costlier than underpredictions, but also that overpredictions of output are particularly costly during periods of low economic growth. This finding can be justified if the cost of an overly tight monetary policy is particularly high during periods with low economic growth when such a policy may cause or extend a recession.9

练习题

Which of the following is NOT a valid pairing in the binary loss function when and ?

A.
B.
C.
D.

What condition must be satisfied for the binary loss function to be valid?

A.
B.
C.
D.

Select all correct statements about the binary loss function when .

A. if
B. if
C. if
D. if

The binary loss function can be written as where .

The binary loss function is normalized so that the sum of losses from incorrect forecasts equals ___.

Explain why the condition is necessary for the binary loss function.

What is the formula for the Mean Absolute Percentage Error (MAPE)?

A.
B.
C.
D.

Select all advantages of using MAPE as a loss function.

A. It is unitless
B. It depends on the units of the outcome
C. It is useful for comparing forecasts across different units
D. It heavily weights errors when is near 0

MAPE is considered advantageous when evaluating forecast methods across outcomes measured in different units.

The MAPE loss function scales the forecast error by ___.

Why can MAPE be problematic when is near 0?

What is the general form of the Bregman loss function?

A.
B.
C.
D.

Select all properties of the Bregman loss function.

A. It depends on both and
B. It reduces to
C. It ensures the conditional mean is the optimal forecast
D. It is only valid for linear functions

The Bregman loss function's optimal forecast is the conditional mean.

The Bregman loss function is valid for any strictly convex function where ___.

How does the Bregman loss function generalize other loss functions?

Which of the following statements are true about the binary loss function? (Select all that apply)

A. It is valid only if
B. It assigns a loss of 0 to correct forecasts
C. It can be written as
D. It is unbounded

Which of the following are true about level- and forecast-dependent loss functions? (Select all that apply)

A. They depend on both the forecast and the outcome
B. MAPE is an example of such a loss function
C. They always reduce to
D. Bregman loss is an example of such a loss function

Explain the relationship between the binary loss function and the forecast error .

What is the significance of the parameter in the binary loss function?

Which of the following statements correctly describes the relationship between binary loss functions and level-dependent loss functions like MAPE?

A. Binary loss functions depend on the forecast error scaled by the outcome level, while MAPE depends only on the forecast error.
B. Binary loss functions are used for continuous outcomes, while MAPE is used for binary outcomes.
C. Binary loss functions are defined for discrete outcomes and do not scale by outcome level, while MAPE scales the forecast error by the outcome level.
D. Binary loss functions and MAPE both require the outcome level to be non-zero to avoid division by zero.

Which of the following are valid conditions or properties for the binary loss function and MAPE?

A. For binary loss, ensures it is a valid loss function.
B. MAPE is unitless because the forecast error is scaled by the outcome level.
C. Binary loss functions can be written as since they depend only on the sign of the forecast error.
D. MAPE is always well-defined regardless of the predictive density for .
E. Binary loss functions assign the same loss to correct predictions regardless of the outcome value.

The Bregman loss function can reduce to squared error loss under specific conditions, while the binary loss function cannot be expressed as a function of the forecast error alone unless restricted to the sign of the error.

The MAPE loss function is given by . This scaling by means that forecast errors are weighted more heavily when is ___.

登录后解锁笔记、知识点解析、AI 问答

立即登录