正在学习
3.1.2 Interpretation of Forecast Optimality
3.1.2 Interpretation of Forecast Optimality
The notion of an optimal forecast that minimizes the expected loss in (3.2) has several implications. First, taking expectations over Y in (3.2), it follows that optimal point forecasts aim to work well on average rather than for a particular value (single draw) of the outcome. In this sense the forecast attempts to estimate not the realization of the outcome, but rather a function of its predictive or conditional distribution. The optimal forecast itself is a function of z and can be considered a parameter of the conditional distribution, . Under MSE loss the relevant function is the conditional mean; under MAE loss the relevant function is the conditional median; for Linex loss it is a feature of the moment generating function for the conditional random variable. The optimal forecast in this sense is unique. However, as we show below, when we move to estimating this function (or parameter), there will be no unique optimal forecast since there are many estimators for these features of the conditional distribution for Y, none of which are the single optimal estimator.
That the optimal forecast minimizes the average conditional loss also highlights the distinction between statistically bad forecasts and economically bad forecasts. A statistically bad forecast comes from constructing a poor model given the available data, while an economically bad forecast provides forecasts that are too imprecise for useful decision making. Forecasts are often viewed as “poor” if they are far from the observed realization of the outcome variable. However, in practice even good forecast models occasionally produce forecasts that are far from the outcome. For example, this will happen when we use an estimate of the conditional mean as our forecast but observe an outcome, y, drawn from the tails of its conditional distribution . We do not necessarily view this as a bad forecast in the sense that such outcomes can occur even when the model is correctly specified.
Second, the optimal forecast in all these examples depends on the conditioning variables used to produce the forecast, Z. Conditioning on different predictor variables generally leads to different optimal forecasts given those predictor variables. This is true regardless of the forecasting problem. Any claims of optimality for a particular forecast model should be restricted to optimality for the given set of predictor variables used to construct the forecast.
Third, the optimization in (3.3) is with respect to a class of functions The set of possible functions is often too wide to suggest the form of a reasonable forecast. Consider, for example, MSE loss for which we know that the optimal forecast takes the form of the conditional mean. However, the class of possible models for the conditional mean is immensely large, and hence does not actually define the form of the conditional expectation, (3.5), even up to a finite set of unknown parameters. Results can sometimes be established when the search over functions is restricted, but the notion of forecast optimality is clearly weaker.
Example 3.1.5 (MSE loss and linear projections). Under MSE loss the optimal forecast takes the form . Computing the conditional mean requires knowledge about the true relationship between y and z. Forecasts commonly use simple regression models based on linear projection of y , where are the parameters from the linear projection. Such projections are a subset of all possible prediction models so MSE and linear projections are generally suboptimal in population. By construction, forecast errors and projections are orthogonal within the class of linear projections:
Within the class of linear models, linear projection is therefore optimal.
We might limit the set of forecast models to specific parametric models rather than extending it to a wide class of functions. As in the previous example, this set may or may not include the optimal model when all possible functions are considered. When this restricted class can be written in parametric form as a function of the unknown parameters and the data the optimal forecast in this restricted class is given by where
We refer to as the pseudo-true value for However, need not be unique because the expected value of the loss function may obtain a minimum at a number of different values for Note that the value of the pseudo-true parameter, , depends on , the parameters of the predictive distribution , so the expected loss is still a function of rather than simply .
While reduced-form expressions for the optimal forecast are often unavailable, some results can be shown for restricted classes of loss functions and conditional distributions for given . For error-based loss functions Granger (1969b) establishes that the conditional mean is the optimal forecast if the loss function is symmetric about 0 and the conditional distribution of the outcome is symmetric about so that is symmetric about 0. This result also requires that either (i) the derivative of the loss function, is strictly monotonically increasing or (ii) the conditional density is continuous and unimodal. We show the result in the first case. The expected loss is given by
where and is the forecast bias. Differentiating (3.11) with respect to we have
Using the assumed symmetry of and the antisymmetry of we have
This means that is a solution to (3.12). As pointed out by Granger, the solution to (3.12) is a unique minimum. To see this, suppose that there is another value of α that satisfies (3.13), i.e.,
Because is increasing, we have , while conversely . Because and integrates to 1, this means that (3.14) cannot hold unless . The second-order condition for the optimum forecast guarantees that is a minimum.
练习题
Under MSE loss, what is the optimal forecast function?
What is the key difference between a statistically bad forecast and an economically bad forecast?
Which of the following statements are true about the optimal forecast?
Under MAE loss, the optimal forecast is the conditional median.
The optimal forecast is always unique, regardless of the loss function used.
Under Linex loss, the optimal forecast is a feature of the ___ generating function for the conditional random variable.
The optimal forecast under MSE loss is given by , which is the ___.
Explain why the optimal forecast depends on the conditioning variables used.
What is the relationship between the optimal forecast and the class of functions used for optimization?
登录后解锁笔记、知识点解析、AI 问答
立即登录