正在学习

12.2 DENSITY FORECASTS FOR BINARY OUTCOMES

12.2 DENSITY FORECASTS FOR BINARY OUTCOMES

As noted in the previous section, the distribution of a binary outcome is described entirely by its mean, so density estimation and estimation of the mean amount to the same estimation problem. Hence, the most common form of forecast reported for 0–1 outcomes is an estimate of the conditional probability that i.e.,

Forecasters with MSE loss could also be interested in reporting the conditional probability that constructed by choosing as

whose solution is the conditional mean, . This is simply a special case of the result in chapter 3 pertaining to the Bernoulli distribution. For the 0–1 case, the density forecast and the conditional mean forecast are identical.

12.2.1 Parametric Density Forecasting Models

Parametric models of the conditional mean of Y are popular since they can be constructed to ensure that the probability lies in the 0–1 interval provided that they are nonlinear functions of the data. The most popular methods are the probit and logit models. Linear models truncated to fall between 0 and 1 can also be employed.

Estimation of the parameters of a model for the conditional distribution requires some form of loss function. Typically we choose a strictly proper scoring rule as reviewed in chapter 2. Recall that scoring rules take the form , where is finite and is maximized when it is set equal to the true conditional probability that generates y. The sample analog is

Sample estimates of the parameters, are then given by

Logit and probit models are typically estimated by maximum likelihood (log scoring rule) methods using an index model. Given a parametric function for the probability that and a one-step-ahead forecasting model, this approach chooses to maximize

The probit specification for the linear index model sets

where is the c.d.f. of a normal distribution; the logit model sets

Choices between these specifications are usually motivated by some background model , where is unobserved, while the indicator is observed. The probit or logit model (or other specifications) then follow from distributional assumptions on see Maddala (1986) for an overview.

Under assumptions on the data-generating process, general properties can be established for estimates of the conditional probability (12.7) using different models and scoring rules. If the data are strictly stationary and sufficiently well behaved, . Such results rely on different proofs of laws of large numbers, an example of which is given in Elliott, Ghanem, and Krüger (2014). If the data are covariance stationary rather than strictly stationary, it is still possible that the estimator converges to the value that minimizes average expected loss, . Use of a strictly proper scoring rule is necessary to ensure identification. de Jong and Woutersen (2011) present consistency results for the probit model with lagged values for y (which are binary) added.

For all strictly proper scoring rules, is the true conditional probability provided that the model is correctly specified. Regardless of the scoring rule, it follows that the forecasts should be similar in sufficiently large samples. This gives little reason to choose between scoring rules as long as they are strictly proper. However, the log scoring rule yields an efficient estimate for since it is the maximum likelihood objective and so would be the preferred scoring rule in finite samples.

Correct specification of the conditional probability is a strong assumption and is difficult to defend unless there is good economic theory guiding the choice of the model. If the model is not correctly specified, depends on the chosen scoring rule. Section 12.1 examined how scoring rules are derived from weighting different utility functions. As the weighted average utility functions change, so do the approximate models that maximize them. Hence, the choice of scoring rule depends on which of these weighting schemes over utility functions is most relevant to the problem at hand.

Multistep forecasts of binary variables engender the same choice between using either a direct or an iterated method for constructing forecasts as discussed in chapter 7. The direct approach uses the likelihood

and hence computation of forecasts is straightforward. Iterated forecasts require taking expectations over multiple paths of . Kauppi and Saikkonen (2008) note that because of the binary nature of such calculations are simpler than in the general case and provide (model-specific) formulas for constructing iterated forecasts. As noted in chapter 9, it is more likely that the direct method yields better forecasts when the model is misspecified.

When parametric specifications such as the probit or logit are employed to model the conditional probability, maximum likelihood (log scoring) is typically chosen as the objective function. Examples include Estrella and Mishkin (1998) and Wright (2006) who use probit models to forecast recessions with various financial variables.

练习题

What is the most common form of forecast reported for 0–1 outcomes?

A. The variance of
B. The median of
C. The conditional probability that
D. The mode of

For forecasters with MSE loss, what is the solution to minimizing ?

A. The conditional variance of
B. The conditional median of
C. The conditional mean of
D. The mode of

Which of the following are popular parametric models for the conditional mean of ?

A. Probit model
B. Logit model
C. Linear regression model
D. Truncated linear model

For binary outcomes, the density forecast and the conditional mean forecast are identical.

The sample analog for estimating the parameters of a model for the conditional distribution is given by , where is a ___.

Explain why the logit and probit models are typically estimated by maximum likelihood methods.

What is the probit specification for the linear index model?

A.
B.
C.
D.

Which of the following statements are true about the logit model specification?

A. It uses the c.d.f. of a normal distribution.
B. It sets .
C. It is a linear model.
D. It ensures probabilities lie in the 0–1 interval.

The choice between probit and logit models is usually motivated by the distributional assumptions on the error term in the background model .

Under assumptions on the data-generating process, if the data are strictly stationary and sufficiently well behaved, . This result relies on ___.

Why is the log scoring rule preferred in finite samples for estimating ?

Which of the following statements are true about strictly proper scoring rules?

A. They ensure the forecasts are similar in large samples.
B. They yield inefficient estimates for .
C. They are necessary for identification.
D. They depend on the chosen model specification.

Which of the following statements about binary outcome density estimation is correct?

A. The density of a binary outcome is described by both its mean and variance.
B. The most common form of forecast for binary outcomes is an estimate of the conditional variance.
C. The density estimation and estimation of the mean are the same problem for binary outcomes, and the forecast is an estimate of the conditional probability that .
D. For binary outcomes, density estimation is independent of the conditional mean estimation.

Which of the following are true about parametric density forecasting models for binary outcomes?

A. Linear models can be used without any modifications for binary outcomes.
B. Probit and logit models are popular parametric models for binary outcomes.
C. Parametric models must ensure that the probability lies in the 0–1 interval.
D. Linear models truncated to fall between 0 and 1 can also be employed.

For binary outcomes, the log scoring rule used in maximum likelihood estimation of logit and probit models yields an efficient estimate for and is the preferred scoring rule in finite samples.

The probit specification for the linear index model sets ___ , where is the c.d.f. of a normal distribution.

登录后解锁笔记、知识点解析、AI 问答

立即登录