正在学习

15.2 LOSS DECOMPOSITION METHODS

15.2 LOSS DECOMPOSITION METHODS

Methods for decomposing the loss are popular in the applied forecasting literature, particularly under squared error loss. Theil (1961) suggested the following decomposition:


Figure 15.5: Density plots for the one-quarter-ahead forecast errors from the Federal Reserve’s Greenbook forecasts of GDP growth and the inflation rate.

where is the correlation between y and The sample analog to this expression is

Here and are sample means and , where are sample estimates of , and Cov(y, f ), respectively.

The first term in (15.2) measures the difference in the means of the forecast and outcome and so captures whether the average forecast is close to the average outcome. This term is 0 when the average outcome and forecast lie on the line in the scatterplot of forecasts versus outcomes. A large value of this term indicates that average forecasts are off center, or biased. For forecasts constructed using lagged values of the outcome, this term is unlikely to be large.

The second term in (15.2) measures differences between the standard deviations of and y and captures whether the variation in the forecasts is similar to that of the outcomes. For hard-to-predict variables such as stock returns, this term tends to be very large since the fluctuation in the outcome is an order of magnitude greater than that of the forecasts. Conversely, if forecasts are generated by a random walk model that sets , this term will be close to 0 since the forecast and outcome display the same amount of variability.

The last term in (15.2) depends on the correlation between the forecast and the outcome and is smaller, the larger their correlation. This term will be large for hardto-predict variables, although the scaling by may reduce this term.

An alternative approach, more commonly seen in binary prediction for which it was originally developed, decomposes mean squared errors into calibration and resolution terms:

Here unconditional expectations are taken over both y and The first term in this decomposition is the unconditional variance of the outcome variable. This does not depend on the forecast and so can be taken as given for the remainder of the analysis. Calibration is defined as . For any forecast, , this measures the difference between the expected value of the outcome given the forecast and the forecast itself. Hence, the second term in (15.3) equals the average squared calibration error, where the average is taken over f. A well-calibrated model has a calibration measure of 0 for all .

The last term in (15.3) is known as the resolution term and enters with a minus sign. Resolution measures the squared average (computed over all forecasts) of the difference between the average outcome conditional on the forecast and the unconditional average outcome. This term will be small, indicating poor resolution, if, conditional on the forecast, the average value of the outcome is equal to the unconditional average value of the outcome for a wide range of forecasts. In such cases, the forecast is not picking up much variation in the outcome variable apart from its mean. Conversely, high resolution is achieved when the forecast picks up a lot of the variation in the outcome variable and tends to be large.

Estimation of the terms in the decomposition in (15.3) requires estimates of the distribution of the outcome variable conditional on the forecast. A common approach is to discretize the forecast to enable such calculations. Galbraith and van Norden (2011) propose to use nonparametric methods for the case with continuous forecasts.

The decomposition can be modified for data that appear nonstationary or show trending behavior by setting . The interpretation of the various components in the decompositions now applies to the actual and predicted change. This ensures that the variance and covariance terms can be estimated consistently when samples become large and is similar to first-differencing data to induce stationarity for variables with a unit root.

An important limitation of these decompositions is that there are no objective measures for how large or small the terms should be for any particular forecasting problem. Moreover, the terms are typically correlated so that one model’s improvement in a specific term may result in a deterioration in the other terms. In this case, a reasonable forecast improvement would require that the improvement in one of the terms more than offsets any reductions in performance in the other terms.

练习题

In Theil's loss decomposition under squared error loss, which term represents the difference in the means of the forecast and outcome?

A.
B.
C.
D.

Which of the following is the final form of Theil's decomposition under squared error loss?

A.
B.
C.
D.

Which of the following are components of the sample analog of Theil's decomposition? (Select all that apply)

A.
B.
C.
D.
E.

The first term in (15.2) is 0 when the average outcome and forecast lie on the line in the scatterplot of forecasts versus outcomes.

The second term in (15.2) is likely to be very large for hard-to-predict variables such as stock returns because the fluctuation in the outcome is much greater than that of the forecasts.

The last term in (15.2) depends on the correlation between the forecast and the outcome and is smaller, the larger their correlation. This term is represented as , where . The sample analog of this term is ___.

In the alternative decomposition for binary prediction, the term represents the ___.

Explain the interpretation of the second term in the alternative decomposition for binary prediction.

Which of the following statements are true about the interpretation of terms in (15.2)? (Select all that apply)

A. The first term measures the bias in the forecasts.
B. The second term measures the difference in variability between the forecasts and outcomes.
C. The last term measures the impact of the correlation between forecasts and outcomes.
D. The first term measures the variability of the outcomes.

How does the sample analog of Theil's decomposition relate to the population version? Provide a brief explanation.

In Theil's loss decomposition, the term measures the difference in the means of the forecast and outcome. If a scatterplot of forecasts versus outcomes shows that the average forecast and average outcome lie on the line, what can be concluded about this term?

A. It is positive and large.
B. It is negative and large.
C. It is zero.
D. It cannot be determined from the given information.

Which of the following statements are true regarding the terms in Theil's loss decomposition under squared error loss? Select all that apply.

A. The term measures differences between the standard deviations of and .
B. The term depends on the correlation between the forecast and the outcome.
C. The term is always positive.
D. The term is larger when the correlation between the forecast and the outcome is smaller.

In Theil's loss decomposition, the term will be close to zero if the forecasts are generated by a random walk model that sets .

In Theil's loss decomposition, the term that captures whether the variation in the forecasts is similar to that of the outcomes is ___.

登录后解锁笔记、知识点解析、AI 问答

立即登录