正在学习

16.3 CONDUCTING INFERENCE ON THE OUT-OF-SAMPLE AVERAGE LOSS

16.3 CONDUCTING INFERENCE ON THE OUT-OF-SAMPLE AVERAGE LOSS

The simulated (or “pseudo”) out-of-sample (OoS) approach described in the previous section seeks to mimic the real-time updating scheme underlying most forecasts. The basic idea is to split the data into an initial estimation sample (the in-sample period) and a subsequent evaluation sample (the out-of-sample period). Forecasts are based on parameter estimates that use data only up to the date when the forecast is computed. As the sample expands, the model parameters get updated, resulting in a sequence of forecasts. Comparing realized values to these forecasts, we can compute the average loss.

West (1996) provides the main set of results that establishes conditions under which the estimated loss is both consistent and asymptotically normal under expanding (recursive), rolling, and fixed estimation windows. The assumptions involve conventional restrictions on the data-generating process, but impose further limitations on the form of the forecasting model. The analysis assumes that the initial sample for estimating has observations, and that this allows the construction of out-of-sample forecast errors, where is the forecast horizon. If the full sample is used for parameter estimation, each forecast is computed using

Let be a vector of losses whose individual elements are denoted . The observed losses depend on the parameter estimates, and are denoted . The analysis requires that there exists a pseudo-true vector, , for the model used to generate the forecasts; is the limiting value of as the sample size gets very large. Define as the loss function evaluated at this pseudo-true value, while is the derivative evaluated at Also, let be the orthogonality conditions used to estimate the model parameters so that is a zero-mean process. The following assumptions summarize the technical conditions needed for West’s result (West, 1996, Theorem 4.1).

  1. The function is twice differentiable in a neighborhood around

  2. The estimate can be written as a moment-type estimator of full rank. Specifically, , where is is with , for B of rank k; ; and

  3. The second derivative, , is bounded.

  4. The derivative, , when evaluated at the pseudo-true value satisfies certain mixing, moment, and stationarity conditions. Moreover, letting and is assumed to be positive definite.

  5. We have and . Define if and = 0 if and

Under assumptions (1)–(5), West (1996) establishes that

where, in general,

and , and is the limiting variance–covariance matrix of the estimators, i.e., asymptotically . If either or , then

The result establishes that asymptotically, we can conduct inference about forecast accuracy using standard normal distributions provided that the correct (asymptotically valid) estimate of is used. From (16.14), consists of three components: is the long-run variance of the loss under known model parameters, the third term reflects how parameter estimation affects the variance of the loss, while the second term reflects the covariance between the first and third terms. Since the second and third terms reflect estimation error, they will depend on the underlying estimation scheme—fixed, rolling, or expanding—via .

To understand the result, it is useful to go through the heuristics of the proof. The proof works by expanding the average OoS loss around the pseudo-true parameter value, , at each data point. Assumption (1) in West’s theorem allows a mean value theorem expansion around so that, for some intermediate value, ,

The first term in (16.15) is the contribution to the average loss evaluated at the pseudo-true value, i.e., the average loss when the parameters are known. Under West’s conditions, this converges to its expected value and contributes the term to the covariance matrix. This is similar to the usual covariance for the mean of a stationary series. The second term in (16.15) involves the term and need not be identically equal to 0 since parameter estimates are periodically updated.4 In some cases this term is not only nonzero but does not vanish asymptotically when scaled by , resulting in the additional term in the variance in (16.14). The third term in (16.15) involves a quadratic in and also does not disappear asymptotically.

In cases where the second and third terms in equal , estimation error does not contribute to the variance of the average out-of-sample loss and so the result captures only first-order effects, while ignoring terms of order such as parameter estimation error. In such cases, in-sample and out-of-sample average loss have identical means and approximately identical variances, which makes it difficult to determine what is a better way to evaluate predictive performance in general.

One case of particular interest arises when the same loss function (e.g., squared error loss) is used for parameter estimation as well as forecast evaluation. In this case the extra terms in West’s variance drop out and so the expression simplifies a great deal.

When West’s theorem applies, , and so the result establishes that the estimated average loss is a consistent estimator for the actual average loss.

West’s result deals with inference about the expected loss evaluated at the population value of the model parameters, , i.e., . Complications arise because the inference cannot be based on the unobserved population value of the parameters but instead is based on finite sample information for which parameter estimation error can play an important role.

West’s result does not cover inference about the finite-sample predictive ability of a model which is related to the model’s expected loss evaluated at the current parameter values, . This case is considered by Giacomini and White (2006).

West (1996, table 2) presents Monte Carlo results for a dynamic bivariate model with exogenous regressors. For his design , and so the variance–covariance matrix should account for parameter estimation error in the construction of the forecasts. Across a range of cases West shows that without this correction, the size of the test can be far greater than its nominal size, e.g., with rejection rates of 20–50% instead of 5%. With the correction for estimation error, size is well controlled unless the number of observations in the initial regression model is small. Specifically, the test is undersized if is small and oversized if is relatively large.

练习题

What is the primary purpose of the simulated out-of-sample approach?

A. To minimize the sample size for estimation
B. To mimic the real-time updating scheme underlying most forecasts
C. To increase the number of forecasts
D. To use the full data sample for estimation and forecast evaluation

What does West's result establish about the estimated loss?

A. It is always unbiased
B. It is consistent and asymptotically normal under certain conditions
C. It is independent of the sample size
D. It is only valid for fixed estimation windows

What is the pseudo-true vector in the context of forecasting models?

A. The initial estimate of
B. The limiting value of as the sample size gets very large
C. The observed losses depending on the parameter estimates
D. The vector of losses whose individual elements are denoted

Which of the following are assumptions required for West's result?

A. The function is twice differentiable in a neighborhood around
B. The estimate can be written as a moment-type estimator of full rank
C. The second derivative is unbounded
D. The derivative satisfies certain mixing, moment, and stationarity conditions when evaluated at

The second derivative must be bounded for West's result to hold.

The number of out-of-sample forecast errors, , is given by , where is the number of observations in the initial sample and is the ___.

Explain the significance of the pseudo-true vector in the context of West's result.

Which of the following statements are true regarding the orthogonality conditions used in West's result?

A. is a zero-mean process
B. are used to estimate the model parameters
C. are independent of the sample size
D. are used to define the pseudo-true vector

West's result implies that the estimated loss is asymptotically normal regardless of the estimation window used (expanding, rolling, or fixed).

What is the role of the matrix in West's asymptotic normality result?

Which of the following best describes the relationship between the sample size conditions and West's result?

A. and must be finite
B. and must tend to infinity and must tend to a constant
C. must be zero
D. must be zero

The matrix in West's asymptotic normality result consists of three components: , a term reflecting how parameter estimation affects the variance of the loss, and a term reflecting the covariance between the first and third terms. The second term is given by , where is defined based on the ratio as if , if , and if .

Which of the following are true about the interpretation of West's asymptotic normality result? (Select all that apply)

A. It allows for inference about forecast accuracy using standard normal distributions
B. The correct estimate of must be used
C. The result is only valid for in-sample forecasts
D. The components of include the long-run variance of the loss, the effect of parameter estimation on the variance, and their covariance

The term in the matrix reflects the effect of parameter estimation on the variance of the loss.

How does the choice of estimation window (expanding, rolling, or fixed) affect the components of the matrix in West's result?

What is the primary advantage of using the simulated out-of-sample approach over in-sample evaluation?

A. It is computationally simpler
B. It provides a consistent estimate of the unconditional expected loss
C. It eliminates the need for parameter estimation
D. It always results in lower average loss

In the context of West's result, the term is defined as , where are the ___.

Which of the following are implications of West's asymptotic normality result for forecast evaluation? (Select all that apply)

A. It supports the use of standard normal distributions for hypothesis testing
B. It requires the estimation of the matrix
C. It is only applicable to linear regression models
D. It highlights the importance of the choice of estimation window

Which of the following is NOT a condition required for West's result on the asymptotic normality of the estimated loss?

A. The function is twice differentiable in a neighborhood around
B. The estimate can be written as a moment-type estimator of full rank
C. The second derivative, , is unbounded
D. The derivative, , when evaluated at the pseudo-true value satisfies certain mixing, moment, and stationarity conditions

登录后解锁笔记、知识点解析、AI 问答

立即登录