正在学习
14.2.1 Estimation of Forecast Combination Weights under MSE Loss
14.2.1 Estimation of Forecast Combination Weights under MSE Loss
Many different ways exist for estimating the combination weights under MSE loss. The existence of such an extensive list of methods boils down to a number of previously discussed issues in constructing forecasts, namely, the role of estimation error, the lack of a single optimal estimation scheme, along with the empirical observation that simple methods are difficult to beat in practice.
A common baseline is to use a simple equal-weighted average of the forecasts:
An advantage of this approach is that it generates no estimation error since the combination weights are imposed rather than estimated. This has obvious intuitive appeal given the restrictions usually imposed on the combination weights. The method also has appeal if, when constructing forecast combinations over time, the panel of forecasts being combined changes dimension or if only a very short time series of the forecasts is available. Concerns about outlier forecasts can be addressed by using the median forecast or the trimmed mean—assuming that m is sufficiently large—instead of the mean.
The trimmed mean works as follows. Suppose the forecasts have been ranked from smallest to largest . Trimming a proportion of the smallest and largest forecasts, the trimmed mean is computed as
where rounds to the nearest (smaller) integer.
A natural alternative to these simple robust estimation methods is to use sample estimates for the population-optimal weights in (14.9). Bates and Granger (1969) suggest simply replacing the unknown variances and covariance in the formula for the optimal weights (14.9) with the equivalent sample estimates. More generally, for an vector of forecasts, the estimated variance–covariance matrix could be used in the formula for the combination weights (14.9). This plug-in solution turns out to be numerically identical to the restricted least squares estimator of the weights from a regression of the outcome on the forecasts and no intercept subject to the restriction that the coefficients sum to 1 :
where is the vector of forecasts and is the usual sample estimator of the covariance matrix,
To see the equivalence between the Bates and Granger approach and restricted least squares, consider the linear regression of the outcome on the forecasts:
Rearranging, we have
where e is the vector of forecast errors. Under the restriction that , it follows that . Minimizing the sum of squared residuals subject to the restriction that the weights sum to 1 is therefore the same as minimizing , and thus gives the same result as Bates and Granger (1969); see Granger and Ramanathan (1984).3 If there is suspicion that the individual forecasts are biased, a constant can be included in the regression.
Variants of least squares combination weights, much like variants of linear forecasting models with general data, revolve around attempts to reduce risk over some parts of the space of models, where the risk arises from the estimation error involved in constructing the weights. One possibility is to use Bayesian or empirical Bayes methods to estimate the combination weights. Clemen and Winkler (1999) propose one such approach that relies on the use of an inverse Wishart distribution on
Different assumptions on priors and model parameters yield different estimators. Diebold and Pauly (1990) suggest both Bayesian and empirical Bayes shrinkage-type methods. Using the results in chapter it follows that for a Gaussian model and Gaussian priors the Bayesian estimator yields forecast combination weights formed as a weighted average of the prior, , and the least squares estimates, . As noted above, the natural uninformative prior is to assign equal weights to the forecasts, so suppose that the prior weights on the m forecasts are . Under the g -prior of Zellner (1986b) discussed in chapter 5, where is a scalar prior that controls the degree of shrinkage with large values of implying high precision and strong shrinkage towards the prior mean of , the posterior mean of the combination weights becomes
where is the OLS estimate of ω in (14.14). More generally, assuming normal priors on ω, , where is the matrix of regressors in (14.14), Diebold and Pauly (1990) use empirical Bayes combination weights given by
where
where is an estimator of the variance of the residuals in the regression model, . Equation (14.17) is similar to a shrinkage-type estimator for the combination weights.
Between data-free schemes such as equal weighting and methods attempting to approximate the optimal combination weights are methods that utilize data in the forecast combination scheme without attempting to asymptotically obtain the optimal combination weights. The simplest example is the trimmed mean in (14.13) which discards some fraction of the largest and smallest forecasts before taking the mean of the remaining forecasts.
Another frequently used approach simply ignores correlations across forecast errors and uses weights that are proportional to the inverse of the individual models’ MSE values, MSEi :
A variant of this method is considered by Aiolfi and Timmermann (2006) who propose a robust weighting scheme that lets the combination weights be inversely proportional to the forecast models’ rank, Ranki :
Here the model with the lowest MSE value gets a rank of 1, the model with the second lowest MSE performance gets a rank of 2, and so forth. This combination scheme again ignores correlations across forecast errors.
Aiolfi and Timmermann (2006) also consider a factor-based clustering approach. If the outcome variable has a factor structure and the individual forecasts can be clustered according to which factors they track, in some cases little is lost by pooling forecasts within clusters. Their approach is to identify a small set of clusters, form equal-weighted forecasts within each cluster, and then apply least squares combination methods such as (14.14) to these pooled forecasts.
练习题
What is the primary advantage of using an equal-weighted average of forecasts?
Which of the following is the formula for the trimmed mean of forecasts?
Which of the following are valid methods for addressing concerns about outlier forecasts?
The Bates and Granger approach for combination weights is numerically identical to the restricted least squares estimator of the weights from a regression of the outcome on the forecasts with no intercept and the restriction that the coefficients sum to 1.
The formula for the estimated combination weights using the Bates and Granger approach is , where is the sample estimator of the ___.
Explain why the equal-weighted average of forecasts generates no estimation error.
Which of the following is a key assumption in the equivalence between the Bates and Granger approach and restricted least squares?
Which of the following statements about the trimmed mean are correct?
The Bayesian estimator for a Gaussian model and Gaussian priors yields forecast combination weights that are a weighted average of the prior weights and the least squares estimates.
Under the g-prior of Zellner, the posterior mean of the combination weights is given by , where controls the degree of ___.
What is the role of the scalar in the g-prior of Zellner?
Which of the following are true about the Bates and Granger approach and restricted least squares?
Which of the following is a key difference between the equal-weighted average and the Bates and Granger approach?
The trimmed mean is calculated by removing the smallest and largest forecasts and averaging the remaining forecasts without any adjustment for the number of forecasts removed.
When the individual forecast errors have identical variance and identical pairwise correlations , which of the following is true about the optimal combination weights under MSE loss?
Which of the following statements are correct regarding the estimation of forecast combination weights under MSE loss?
The trimmed mean can be used to address concerns about outlier forecasts, and it is calculated by trimming a proportion of the smallest and largest forecasts and then taking the average of the remaining forecasts.
If the forecast errors have equal variance, , under MSE loss, it is optimal to assign ___ weights to the forecasts, regardless of their correlation.
登录后解锁笔记、知识点解析、AI 问答
立即登录