正在学习
12.1 POINT AND DENSITY FORECASTS FOR BINARY OUTCOMES
12.1 POINT AND DENSITY FORECASTS FOR BINARY OUTCOMES
We first introduce some notation needed in the analysis. Let if some event occurs, while if the event does not occur, and let the random variable denote the probability that given Z. The choice of how we map the two outcomes to specific numbers is arbitrary, some methods in the literature use the normalization that the outcome equals −1 when the event does not occur. To cover this case we also define which transforms the outcome space to {−1, 1}. The literature on binary forecasting uses both of these codings; either leads to the same results, although some of the expressions change. We focus on but also provide some results for to map back to the original papers.
Typically the conditional mean provides a point forecast for the outcomes of whereas gives the density forecast. For binary outcomes, however, this is not the case. Since Y takes one of only two values, then and so estimating one is the same as estimating the other. Point forecasts on the other hand should have the same support as the outcome, so point forecasts for forecasting Y should be either 0 or 1. This will never be the conditional mean for the Bernoulli distribution except in trivial cases where all outcomes are identical or the outcome is known with certainty, i.e., we have , and hence , excluding trivial cases.
A density forecast along with a utility function can be used to generate a point forecast, so the standard relation between these two objects holds despite the different relations they individually bear with the conditional expectation. As usual, we are interested in predicting the outcome with a forecast function . The utility function of the decision maker, , depends on the action, the outcome, and potentially all or some subset of the covariates, Z. If the forecast and outcome are both binary, utility can take only four possible values for any z:

Figure 12.1: Probability forecast versus cutoff.
Sensible loss functions satisfy that and for all so that a correct forecast is always superior to an incorrect one. The utility from correctly predicting a negative outcome need not be equal to the utility from correctly predicting a positive outcome, so we could have . For example, if the predicted outcome is whether a stock price goes up or down , we would have if being correct when the market falls is better than being correct when it rises. This simply says that avoiding losses is better for investors than making gains and might reflect a concave utility function.
Recalling from chapter 3 that risk can be defined as the negative of expected utility, , a simple result lets us characterize the optimal density forecast with binary outcomes. Conditioning on the risk is minimized by setting , i.e., forecasting a positive outcome, when
We refer to as the cutoff point at which the forecast switches between its two possible outcomes. When the probability that is greater than some cutoff, , determined by the utility function, the optimal forecast is
A common choice of cutoff is . This choice is illustrated in figure 12.1, which shows for a univariate Z the situation where is monotonically increasing in Z with . The model forecasts an outcome of for . This choice is optimal when , i.e., when the relative utilities from a correct forecast versus an incorrect one are equal2 regardless of the direction of Y. Choices of cutoff other than 0.5 are relevant in situations where these marginal gains differ. For example, if the gain for is bigger than that for we could set the cutoff below 0.5 so as to “overweight” the more valuable forecasts,
When the utility function does not depend on , and the utility- maximization problem can be reduced through a variety of renormalizations. Without loss of generality, we can subtract from the utility for , so that the utility equals 0 for a correct positive forecast, while for a false negative it is equal to . Similarly, for we can subtract from utility, thereby setting the utility for a correct negative forecast to 0 and for a false positive. This change leaves the numerator and denominator of c in (12.2) unchanged and allows us to write
In this case we can use that utility has no natural units to normalize This leaves us with , and the simplified representation3
where and the utility function is expressed directly as a function of the cutoff value, c .
If we construct distributional forecasts of the outcome , we must employ a utility function on the probability for both estimation and forecast evaluation purposes. One probability forecast is better than another if it results in a higher utility for all possible cutoffs, c , i.e., is the dominant rule regardless of the utility function. If were known, this would be the dominant rule. More likely, however, has to be estimated and the model is misspecified and so is not equal to the true conditional probability. In this situation, some models could be best for certain regions of while other models could be best for other regions of c.
The scoring rule used for estimating a loss function, if proper, can be related directly to a weighting function over the possible cutoffs, c, which range from 0 to 1.
Proper scoring rules take the form
where for and

Figure 12.2: Weights for different scoring rules.
This relation was first shown in Shuford Jr, Albert, and Massengill (1966) and restated in Schervish (1989, Theorem 4.1). Schervish (1989) further showed that strictly proper scoring rules can be written as
where for . This equation interprets proper scoring rules as weighted averages of utility functions. Different choices of yield different weighting functions. For example, setting results in the log score function— better known as the maximum likelihood objective function given a model for the conditional probability—whereas yields the quadratic score. Notice the very different weights on these two scoring rules. The log score puts very high weights on individuals whose loss functions yield cutoffs close to 0 or 1 relative to those with loss functions that result in cutoffs near 0.5. The quadratic or Brier score (Brier, 1950) weighs all individuals evenly, while the spherical score weighs cutoffs close to 0.5 more heavily than cutoffs near the tail. The weights for the log score, Brier, and spherical scoring rules are shown in figure 12.2.
The relationship between the weights v(c) and the scoring rule can be shown to be
These results hold for the conventional scoring rules applied to binary problems such as the log score, Brier score, and spherical score, all of which weigh utility functions symmetrically around . Moreover, the results can be employed to construct a variety of proper scoring rules that need not treat costs symmetrically.
练习题
In binary outcome analysis, what does the random variable represent?
For binary outcomes, which of the following is true about the conditional mean and the probability ?
Which of the following are properties of sensible loss functions for binary forecasts?
The utility from correctly predicting a negative outcome must be equal to the utility from correctly predicting a positive outcome.
The optimal density forecast with binary outcomes is characterized by setting when , where is defined as . The value is called the ___.
Explain why point forecasts for binary outcomes should be either 0 or 1.
Which of the following is a common choice for the cutoff point ?
Which of the following statements are true about the utility function for binary forecasts?
The conditional mean provides a point forecast for binary outcomes.
What is the relationship between the density forecast and the point forecast for binary outcomes?
In binary outcome forecasting, what is the relationship between and when takes values in ?
Which of the following statements is true regarding the utility function in binary forecasting?
What are the characteristics of a sensible loss function in binary forecasting? (Select all that apply)
In binary forecasting, setting the cutoff point is always the optimal choice regardless of the utility function.
登录后解锁笔记、知识点解析、AI 问答
立即登录