正在学习

12.3 CONSTRUCTING POINT FORECASTS FOR BINARY OUTCOMES

12.3 CONSTRUCTING POINT FORECASTS FOR BINARY OUTCOMES

Estimation problems involving point forecasts generally require that the forecasts have the same support as the outcome. For the binary variable this suggests forecasts that predict either −1 or 1, while for Y the two values are {0, 1}. This makes it very simple to set up the forecasting problem. The loss function need cover only four possible situations—two where the forecast is correct and two where it is false. This, as we saw in section 12.1, permits a general setting in which the universe of loss functions can easily be characterized. Only loss functions of this form have the properties discussed in chapter 2. For example MSE and MAE loss are applicable when and

Utility functions with binary outcomes complicate estimation because they are discontinuous in the forecast. The forecasting problem reduces to finding a function that tells us for which values of Z we have . This again brings out the direct relation between density forecasting and point forecasting, since an estimate of is a density forecast. This relationship motivates the most common approach which we next describe.

12.3.1 Forecasts via

Estimation of the density forecast along with the form of equation (12.2) suggest that a forecast of can be constructed by choosing a forecasting model through the indicator variable

This is an indirect approach since is estimated based on loss functions such as MSE or log score and not on the loss function in (12.1). Hence, there is no guarantee that the estimated coefficients are chosen optimally.

There are numerous examples of this approach in the literature. Dimitras, Zanakis, and Zopounidis (1996) survey business failure prediction models and find that 17 of 66 of the studies surveyed employed this method. Boyes, Hoffman, and Low (1989) use logit models to predict credit default, while Leung, Daouk, and Chen (2000) use both probit and logit models to predict the direction of the stock market. Martin (1977) and Ohlson (1980) use this approach to predict corporate bankruptcies. The choice of a cutoff probability, above which we assign a forecast of 1, tends to be arbitrary in these studies. Leung, Daouk, and Chen (2000) and Qi and Yang (2003) both choose 0.5 as their cutoff, while Boyes, Hoffman, and Low (1989) suggest a loss-based cutoff.

12.3.2 Discriminant Analysis

Discriminant analysis methods are popular in statistics and applied fields other than economics, and date back to work by Fisher (1936). These methods split the joint distributions of the covariates into two different groups, one for , the other for , and arise from considering a hypothesis test between these two populations of covariates. Modeling the population joint densities parametrically, we can compute the likelihood that any subsequent observed set of covariates are generated from the density associated with the group or the other group. A likelihood ratio test between the two groups yields the forecasting rule that if the likelihood evaluated for this group at exceeds that of the other group. If Z is joint normally distributed with common covariance matrices across groups, the prediction rule is linear in the observed z-variables and hence leads to decisions based on linear “scores” that are linked one-to-one with decisions based on the underlying likelihoods. Rather than weighting both likelihoods evenly, one can be weighted higher than the other to yield a procedure that alters the balance between false positive and false negative predictions. The loss function can be brought to bear on the decision through such weights.

Discriminant analysis is often criticized because of its parametric assumptions of joint normality of the covariates. However, this really amounts to the same thing as choosing a parametric model for the conditional probability as becomes clear through the relationship with the logit model. Amemiya (1985, section 9.2.8) shows that under normality of conditional on discriminant analysis is equivalent to using a logit model that includes the covariates linearly as well as quadratically. Under the assumption that the variance–covariance matrices are identical across groups, the quadratic terms drop out. When the group variances are identical, there is a one-to-one (population) relation between the logit specification based on a linear index and the discriminant analysis method. Dimitras, Zanakis, and Zopounidis (1996) note the similarity of predictions using these two methods, where differences arise due to differences in estimation methods for the unknown parameter weights.

练习题

What is the required support for forecasts of a binary variable ?

A.
B.
C.
D.

Which loss function is applicable when and ?

A. Cross-entropy loss
B. Mean Absolute Error (MAE)
C. Mean Squared Error (MSE)
D. Both B and C

Which of the following statements about utility functions with binary outcomes are true?

A. They are continuous in the forecast.
B. They are discontinuous in the forecast.
C. The forecasting problem involves finding .
D. The forecasting problem involves finding .

The forecasting problem for binary outcomes is directly related to density forecasting because an estimate of is a density forecast.

The forecast of can be constructed using the indicator variable . What does represent?

Explain why the approach of estimating based on loss functions such as MSE or log score is considered indirect.

Which of the following studies used a cutoff probability of 0.5 for their forecasts?

A. Boyes, Hoffman, and Low (1989)
B. Leung, Daouk, and Chen (2000)
C. Martin (1977)
D. Ohlson (1980)

Which of the following are true about discriminant analysis methods?

A. They are popular in economics.
B. They date back to work by Fisher (1936).
C. They split the joint distributions of covariates into two groups.
D. They are only used for continuous outcomes.

Discriminant analysis yields the forecasting rule that if the likelihood evaluated for this group at is less than that of the other group.

If is joint normally distributed with common covariance matrices across groups, the prediction rule is linear in the observed ___.

How can the balance between false positive and false negative predictions be altered in discriminant analysis?

What is a common criticism of discriminant analysis?

A. It is too computationally intensive.
B. It assumes joint normality of the covariates.
C. It cannot handle binary outcomes.
D. It requires large sample sizes.

Under normality of conditional on , which of the following statements about discriminant analysis and the logit model are true?

A. Discriminant analysis is equivalent to using a logit model with covariates linearly.
B. Discriminant analysis is equivalent to using a logit model with covariates quadratically.
C. If variance-covariance matrices are identical, quadratic terms drop out.
D. When group variances are identical, there is a one-to-one relation between the logit specification and discriminant analysis.

Which of the following are true about the relationship between density forecasting and point forecasting for binary outcomes?

A. Density forecasting is unrelated to point forecasting.
B. An estimate of is a density forecast.
C. Density forecasting and point forecasting are inversely related.
D. The forecasting problem involves finding .

Which of the following statements are true regarding the relationship between density forecasting and point forecasting for binary outcomes?

A. Density forecasting is unrelated to point forecasting.
B. An estimate of is a density forecast.
C. The forecasting problem reduces to finding a function that tells us for which values of we have .
D. The relationship between density forecasting and point forecasting is indirect.

Explain how the logit model specification is derived and its relationship to the conditional probability .

登录后解锁笔记、知识点解析、AI 问答

立即登录