正在学习
16.3.1 West’s Result under Squared Error Loss
16.3.1 West’s Result under Squared Error Loss
Conditions (1)–(5) in West’s theorem may appear somewhat abstract and hard to verify. To make the result more accessible, we next discuss West’s theorem for the common case with squared error loss and univariate forecasts based on a linear regression model, , whose parameters are estimated by recursive OLS. Assumption (1) in West’s theorem involves both the loss function and the method for constructing the forecast. Under squared error loss, and assuming ,
and so
showing that the loss is twice differentiable.
Turning to the second assumption in West’s theorem, the parameter estimate, , obtained under an expanding window takes the form
\hat { \beta } _ { t } - \beta ^ { * } = \mathopen { } \mathclose \bgroup \left( ( t - 1 ) ^ { - 1 } \sum _ { s = 1 } ^ { t - 1 } x _ { s } x _ { s } ^ { \prime } \aftergroup \egroup \right) ^ { - 1 } \mathopen { } \mathclose \bgroup \left( ( t - 1 ) ^ { - 1 } \sum _ { s = 1 } ^ { t - 1 } x _ { s } (y _ { s + 1} - \beta ^ {* \prime} x _ { s}) \aftergroup \egroup \right) .
This case is covered by assumption (2) with ; B (t) is the inverse of the variance–covariance matrix of and so has limit B under relatively mild conditions on . This leaves the definition of to be , where the dependence on is explicit. We define as the linear projection of on , so by construction. If the regressors were nonstationary, would not converge to B and so this case is ruled out by assumption (2).
Assumption (3) in West’s theorem is a smoothness condition. In the example here, the second derivative does not depend on the parameters and so this assumption follows directly given assumptions on the regressors. Assumption (4) places limited dependence and moment restrictions on the regressors and forecast errors, which again follows under weak assumptions on the DGPs. For example, in the very common case where are covariance stationary, these conditions will hold and the theorem applies.
West’s theorem shows that in the special case where , the covariance matrix of the sample mean can be approximated by the usual spectral density estimate, ignoring sampling variation when we construct the estimates. One such special case arises when the training (estimation) sample used to construct the forecasts overwhelms the out-of-sample period so . Intuitively it makes sense that parameter estimation error is not relevant in this situation.
A second special case is less obvious. Under squared error loss, the forecasts are generated so that the prediction error is uncorrelated with the predictors and so
and so the usual asymptotic results apply. This combination of mean squared error loss and forecasts based on a linear regression model is very popular in applied work. For this pairing it is straightforward to compute standard errors for the expected loss, even though this is uncommon to do in practice. One only need treat the time series of squared error losses as raw data, compute the mean and spectral density at zero frequency in the usual way (e.g., by applying the methods of Newey and West (1987) or Andrews (1991)), and report the associated standard error.
In other cases, sampling error in the parameter estimates must be taken into account. Indeed, when the second and third terms in (16.15) are nonzero, the usual variance estimator is not consistent as it omits the extra two terms in the result. For such cases, West (1996) suggests estimators for the additional terms.
We next present a simple example that illustrates the individual terms in West’s theorem.
Example 16.3.1 (Forecasting a univariate i.i.d. variable by its mean under squared error loss). Let ind and suppose the one-step-ahead forecast model is based on the sample mean,
with . Under squared error loss, . The model is correctly specified in this case and so Under squared error loss the mean value expansion is exact. Using that
the three terms in (16.15) are therefore given by
Averaging over the out-of-sample observations, we get
The first of these terms is the average loss with no uncertainty about . Contributions from parameter uncertainty come from the remaining terms and will disappear asymptotically.
West’s result excludes some interesting cases where the assumptions of the theorem fail to hold. One such situation arises when a vector of forecasts is considered, one forecasting model nests other models, and . Suppose the training sample used for parameter estimation is large so the asymptotic approximation sets the coefficients in the model used to generate the forecasts at their pseudo-true values, . For nested models where the small model is the true one, these parameter values are the same and hence the two sets of forecasts are identical and the covariance matrix will be singular. This is ruled out by assumption (4) in West’s theorem since is assumed to be nonsingular. This situation is not relevant when evaluating a single forecast, but becomes important in the comparison of multiple forecasts described in the next chapter.
For some cases the loss function is not twice differentiable. Examples of loss functions that do not satisfy this property are the MAE, lin-lin, and asymmetric quadratic loss functions, that are not differentiable at a single point. To some extent this requirement reflects the proof in West (1996) which assumes a twice continuously differentiable loss function in order to apply a second-order mean value theorem. Indeed, other methods of proof are available. McCracken (2000) relaxes the differentiability condition with a smoothness condition that requires the expected loss (or difference in losses) to be continuously differentiable around the probability limits of the parameter estimates, while maintaining other conditions similar to those in West (1996). Other loss functions, such as the binary forecasting loss function (see chapter 2) and profit functions that rely on discrete actions cannot be handled in this way, and require a separate approach to establishing the consistency of expected loss, its rate of convergence, and also conditions under which expected loss is asymptotically normal.
练习题
Under squared error loss, what is the first derivative of the loss function with respect to ?
What is the form of the parameter estimate obtained under an expanding window?
Which of the following are true about in West's theorem?
The definition of is , and by construction.
Assumption (3) in West’s theorem requires that the loss function is not differentiable.
Under squared error loss, the second derivative of the loss function with respect to is ___.
In the special case where , the covariance matrix of the sample mean can be approximated by the ___.
Explain why the prediction error is uncorrelated with the predictors under squared error loss.
Which of the following are assumptions in West’s theorem?
What is the implication of the special case where the training sample overwhelms the out-of-sample period ()?
In the context of West's theorem under squared error loss, which of the following statements about the parameter estimate obtained under an expanding window is correct?
Which of the following are assumptions in West's theorem that are relevant when considering squared error loss and a linear regression model with recursive OLS parameter estimation?
Under squared error loss, if the forecasts are generated such that the prediction error is uncorrelated with the predictors, then .
In West's theorem, when , one special case is when the training (estimation) sample used to construct the forecasts overwhelms the out - of - sample period so ___.
登录后解锁笔记、知识点解析、AI 问答
立即登录