正在学习
12.3 CONSTRUCTING POINT FORECASTS FOR BINARY OUTCOMES
12.3 CONSTRUCTING POINT FORECASTS FOR BINARY OUTCOMES
Estimation problems involving point forecasts generally require that the forecasts have the same support as the outcome. For the binary variable this suggests forecasts that predict either −1 or 1, while for Y the two values are {0, 1}. This makes it very simple to set up the forecasting problem. The loss function need cover only four possible situations—two where the forecast is correct and two where it is false. This, as we saw in section 12.1, permits a general setting in which the universe of loss functions can easily be characterized. Only loss functions of this form have the properties discussed in chapter 2. For example MSE and MAE loss are applicable when and
Utility functions with binary outcomes complicate estimation because they are discontinuous in the forecast. The forecasting problem reduces to finding a function that tells us for which values of Z we have . This again brings out the direct relation between density forecasting and point forecasting, since an estimate of is a density forecast. This relationship motivates the most common approach which we next describe.
12.3.1 Forecasts via
Estimation of the density forecast along with the form of equation (12.2) suggest that a forecast of can be constructed by choosing a forecasting model through the indicator variable
This is an indirect approach since is estimated based on loss functions such as MSE or log score and not on the loss function in (12.1). Hence, there is no guarantee that the estimated coefficients are chosen optimally.
There are numerous examples of this approach in the literature. Dimitras, Zanakis, and Zopounidis (1996) survey business failure prediction models and find that 17 of 66 of the studies surveyed employed this method. Boyes, Hoffman, and Low (1989) use logit models to predict credit default, while Leung, Daouk, and Chen (2000) use both probit and logit models to predict the direction of the stock market. Martin (1977) and Ohlson (1980) use this approach to predict corporate bankruptcies. The choice of a cutoff probability, above which we assign a forecast of 1, tends to be arbitrary in these studies. Leung, Daouk, and Chen (2000) and Qi and Yang (2003) both choose 0.5 as their cutoff, while Boyes, Hoffman, and Low (1989) suggest a loss-based cutoff.
12.3.2 Discriminant Analysis
Discriminant analysis methods are popular in statistics and applied fields other than economics, and date back to work by Fisher (1936). These methods split the joint distributions of the covariates into two different groups, one for , the other for , and arise from considering a hypothesis test between these two populations of covariates. Modeling the population joint densities parametrically, we can compute the likelihood that any subsequent observed set of covariates are generated from the density associated with the group or the other group. A likelihood ratio test between the two groups yields the forecasting rule that if the likelihood evaluated for this group at exceeds that of the other group. If Z is joint normally distributed with common covariance matrices across groups, the prediction rule is linear in the observed z-variables and hence leads to decisions based on linear “scores” that are linked one-to-one with decisions based on the underlying likelihoods. Rather than weighting both likelihoods evenly, one can be weighted higher than the other to yield a procedure that alters the balance between false positive and false negative predictions. The loss function can be brought to bear on the decision through such weights.
Discriminant analysis is often criticized because of its parametric assumptions of joint normality of the covariates. However, this really amounts to the same thing as choosing a parametric model for the conditional probability as becomes clear through the relationship with the logit model. Amemiya (1985, section 9.2.8) shows that under normality of conditional on discriminant analysis is equivalent to using a logit model that includes the covariates linearly as well as quadratically. Under the assumption that the variance–covariance matrices are identical across groups, the quadratic terms drop out. When the group variances are identical, there is a one-to-one (population) relation between the logit specification based on a linear index and the discriminant analysis method. Dimitras, Zanakis, and Zopounidis (1996) note the similarity of predictions using these two methods, where differences arise due to differences in estimation methods for the unknown parameter weights.
练习题
What is the required support for forecasts of a binary variable ?
Which loss function is applicable when and ?
Which of the following statements about utility functions with binary outcomes are true?
The forecasting problem for binary outcomes is directly related to density forecasting because an estimate of is a density forecast.
The forecast of can be constructed using the indicator variable . What does represent?
Explain why the approach of estimating based on loss functions such as MSE or log score is considered indirect.
Which of the following studies used a cutoff probability of 0.5 for their forecasts?
Which of the following are true about discriminant analysis methods?
Discriminant analysis yields the forecasting rule that if the likelihood evaluated for this group at is less than that of the other group.
If is joint normally distributed with common covariance matrices across groups, the prediction rule is linear in the observed ___.
How can the balance between false positive and false negative predictions be altered in discriminant analysis?
What is a common criticism of discriminant analysis?
Under normality of conditional on , which of the following statements about discriminant analysis and the logit model are true?
Which of the following are true about the relationship between density forecasting and point forecasting for binary outcomes?
Which of the following statements are true regarding the relationship between density forecasting and point forecasting for binary outcomes?
Explain how the logit model specification is derived and its relationship to the conditional probability .
登录后解锁笔记、知识点解析、AI 问答
立即登录