正在学习

16.1 THE SAMPLING DISTRIBUTION OF AVERAGE LOSSES

16.1 THE SAMPLING DISTRIBUTION OF AVERAGE LOSSES

Good forecasting models produce “small” expected losses, while bad models produce “large” expected losses, where “small” and “large” loss are naturally defined with respect to the loss that would be incurred using an optimal forecast.1 To evaluate a forecasting model, we therefore need to estimate its expected loss. The measure used for this purpose is the sample average loss. This, however, is a random variable that will fluctuate across different samples. Small average losses in a given sample could just be due to luck or could, alternatively, reflect the performance of a genuinely good model. To distinguish between these alternatives, we need to know more about the sampling distribution of average losses.

Risk was defined as

in the notation of chapter 3. Clearly it is useful to understand the risk associated with the forecast at the time this is computed. While risk is unobserved, we can construct an estimate of it through the average loss from a sequence of forecasts and outcomes

Typical examples of the types of expected losses that are estimated empirically include the mean squared error (MSE) or root mean squared error (RMSE), which preserves the unit of the outcome variable, or mean absolute error (MAE). Such measures are sample averages based on forecast errors, , measured over a sample, :

Other loss functions have similar sample analogs.

In addition to average losses, other measures are sometimes reported even though they do not pertain to any particular loss function. Sometimes the forecast error is scaled by the outcome variable which gives rise to the percentage error , and the associated root mean squared percentage error (RMSPE),

Unless based on the forecaster’s loss function, these measures can be viewed more as informal summary statistics that can be useful for giving some sense of the degree of predictability of the outcome variable and the magnitude of forecast errors in the same sense as the standard error from a regression model. Great care must be exercised, however, in interpreting and basing statistical inference on such sample averages.

16.1.1 In-Sample Forecast Evaluation

In-sample, or full-sample, forecasts use the entire data from to estimate model parameters, , and evaluate forecasts constructed off such estimates:

Because data at time were not available in real time to construct the forecast at time , such forecasts could not have been constructed in real time and thus are infeasible. Nevertheless, by making use of the full data set, these forecasts make efficient use of all data that is available at time and eliminate the variation in the sequence of forecasts arising from revisions to model parameters. Hence, these forecasts represent the best forecasts available to a forecaster with access to the full information set, which we denote by

Chapter 6 showed that in-sample evaluation of forecasting models—the practice of using the same sample for model estimation and forecast evaluation—results in a bias towards overparameterized models. By virtue of the properties of minimization, the in-sample estimate, , minimizes the sample estimate of the average loss, i.e.,

for any , including the limiting value . It follows that in-sample evaluation measures typically underestimate the unconditional expected loss.

In many situations provides a consistent estimate of the unconditional expected loss, as the effect of parameter estimation error through the variance term disappears asymptotically. Indeed, as discussed in chapter 4, estimation error disappearing asymptotically is precisely what justifies estimation of the parameters of the forecasting model by minimizing the forecast loss. If the in-sample average loss failed to be consistent for the unconditional loss, this would suggest using a different estimator.

The general problem with in-sample forecast evaluation methods is that they underestimate the true unconditional expected loss and hence yield an inaccurate picture of the likely future loss that a given forecasting model produces. This has led most researchers to consider out-of-sample forecasting methods for evaluating the expected loss, which we next turn to.

练习题

Which of the following statements correctly describes the relationship between in-sample forecast evaluation and the estimation of model parameters?

A. In-sample forecast evaluation uses a subset of data to estimate model parameters and evaluate forecasts.
B. In-sample forecast evaluation uses the entire data set to estimate model parameters and evaluate forecasts, which can lead to overparameterization bias.
C. In-sample forecast evaluation is not affected by the choice of estimation window.
D. In-sample forecast evaluation always provides an unbiased estimate of the unconditional expected loss.

Which of the following are valid concerns when using in-sample forecast evaluation methods? Select all that apply.

A. In-sample evaluation can underestimate the true unconditional expected loss.
B. In-sample evaluation is not affected by small sample sizes.
C. In-sample evaluation may lead to overparameterization bias.
D. In-sample evaluation provides a consistent estimate of the unconditional expected loss as the sample size grows.

The in-sample estimate of the average loss, , is given by . This estimate is ___ by the properties of minimization, meaning it is less than or equal to the average loss for any other parameter vector .

登录后解锁笔记、知识点解析、AI 问答

立即登录