正在学习
3.2.1 Loss-Based versus Two-Step Approaches
3.2.1 Loss-Based versus Two-Step Approaches
The classical (or frequentist) approach to construction of forecasting models falls into two broad areas: either working directly with the sample analog of the loss function or using a two-step approach that first finds the form of the optimal forecast, then uses plug-in estimates for the unknown parameters. This section examines general issues that arise under these approaches. Much of the remainder of the book examines forecast construction in the context of more specific forecasting problems.
The sample analog approach requires finding the sample analog to (3.15) and constructing a forecast model , where are unknown parameters with values that depend on , although as we discuss in chapters 6 and 11 the forecast model is typically also unknown. If is identical to the optimal forecast equals the conditional mean of under MSE loss), then the forecast model is correctly specified. In practice, this should be considered an unlikely event. More likely, is an approximation to the correct specification.
The two-step approach approximates the relevant feature of the conditional distribution of the predicted variable—e.g., under MSE loss this feature is the conditional mean of given a forecast model . Again, if this model is identical to the relevant feature of , then the model is correctly specified, although more likely is an approximation. In the second step, an estimate for , denoted is used to generate a forecast . The two-step approach therefore requires that the relevant feature of can be determined whereas the sample analog approach can always be used. Nevertheless, the two-step approach is more ubiquitous in practice. This is largely because the majority of forecasting studies is undertaken under MSE loss, reducing the forecast analysis to attempts at estimating the conditional mean of the outcome.
Example 3.2.1 (Linex loss, continued). Let , so the data only comprise past values of the dependent variable. we assume that ind , then
and we have . A plug-in estimator might use the forecast
where and is some estimator of the variance,
Since we are estimating a feature of a conditional distribution such as its mean, it follows that point forecasting is really an estimation problem. Most of the results on risk and forecast optimality are analogous to those in parameter estimation theory. For the remainder of this section we go through some of these basic results, adapted to the forecasting problem.
First, consider the case where the forecast model is correctly specified, so is equal to the correct feature of . The analog procedure or the second step of the plug-in approach requires that be estimated from data, so the actual forecast model is not but . Calculation of the risk for a forecasting method involves taking the expectation over z and so gets affected by estimation of the unknown parameters. Because the additional randomness induced by parameter estimation increases the risk, the estimated forecast model will never obtain the optimality properties discussed in the previous section.
Example 3.2.2 (MSE loss given the observed history of the outcome). Consider forecasting under MSE loss using , where ind so that the data comprise past observations drawn independently from the same distribution. The population optimal forecast is . This produces a risk of . When is unknown, we might instead use the sample mean computed over observations as our forecast, where . The risk of this forecast is
This is strictly greater than the risk of the population forecast,
The risk measure in (3.15) typically differs from the conditional expected loss measure in (3.2) since varies with t. We illustrate this point in the next example.
Example 3.2.3 (MSE loss for linear forecasting model). Consider a linear forecasting model , where is independent of . Under squared error loss the expected loss is
Conditional on the data, , and a forecast , for the least squares estimator this becomes
which depends on all values of up to time The unconditional expected loss, obtained by integrating over all -values, becomes
where
Risk functions that assume a correct specification, , and a known value of (and hence can be considered infeasible lower bounds on risk or “limiting” approximations to risk in large samples. At the same time when considering the optimality of a forecast procedure, parameter estimation error should be taken into account. A direct consequence of this is that even in the simplest forecast environments—and even with a correct model specification—there is no single optimal procedure that minimizes risk for all possible values of Instead, the risk functions of different procedures will typically cross for different regions of . The next example illustrates this point.
Example 3.2.4 (MSE loss for shrinkage estimator). Consider the forecasting problem in Example 3.2.2. Rather than use the sample mean estimator, we could consider the shrinkage estimator where . In this case the risk is
For in the range the forecast has lower risk than the sample mean outside this range has lower risk.
The results above were all predicated on the forecast model being correctly specified up to a set of unknown parameters, However, as discussed in the previous subsection, we may restrict our search to a subset of possible models that does not include the optimal model. For example, a forecaster with MSE loss might restrict models to linear projections even though the conditional mean is not linear in the conditioning variables. In this case, optimality is established with reference to this limited set of models.
Standard notions of risk for estimators extend to forecasts, including issues related to optimality of estimation methods. In parallel to estimation methods, a forecast method belongs to a family of optimal procedures if there exists some region of the parameter space, , for which no other forecast method has lower risk. It should be clear from the above examples that this concept of optimality generally defines a set of forecast methods as opposed to a single forecast method. Hence it is generally not possible to claim that there is a single optimal forecast method for any problem, unless is restricted to a point, in which case it must be known.
These considerations also raise the possibility that some forecast procedures are dominated by other forecasts for every value of and so are inferior and should never be used. This notion corresponds to the idea of inadmissible estimators.
Example 3.2.5 (MSE loss for shrinkage estimator, continued). Allowing the shrinkage factor in Example 3.2.4 to take on any value in a set such as the real line, R, generates a family of estimators which we denote . For , it follows from (3.20) that there is no range of for which dominates the sample mean , since the associated risk is strictly larger. Such estimators should thus never be employed. However, there is no such ranking when , in which case stronger shrinkage leads to lower risk for near 0.
练习题
Which of the following best describes the sample analog approach in forecasting?
In the two-step approach, what is the second step after approximating the relevant feature of the conditional distribution?
Which of the following statements about the two-step approach are correct?
The sample analog approach can always be used, whereas the two-step approach requires that the relevant feature of can be determined.
Under MSE loss, the optimal forecast is the conditional mean of given , which is denoted as ___.
Explain why point forecasting can be considered an estimation problem.
In the Linex loss example, what is the form of the optimal forecast ?
Which of the following statements about the risk of an estimated forecast model are correct?
The risk measure in (3.15) is typically the same as the conditional expected loss measure in (3.2).
What is the expected loss under squared error loss for a linear forecasting model ?
Which of the following are key differences between the sample analog approach and the two-step approach in forecasting?
In the context of the MSE loss example with observed history, the population optimal forecast is ___.
Under MSE loss, the optimal forecast is the conditional mean of given . If a two-step approach is used to approximate this feature with a forecast model , and is estimated using data, which of the following statements is correct?
Which of the following statements are true regarding the sample analog approach and the two-step approach in forecasting?
The risk of the estimated forecast model will be equal to the risk of the population optimal forecast when the forecast model is correctly specified.
登录后解锁笔记、知识点解析、AI 问答
立即登录