正在学习

3.2.2 Risk and the Information Set

3.2.2 Risk and the Information Set

The way risk is defined in equation (3.15) means that its precise form depends on the specification of the information set. How the information set is specified is thus important for risk, just as it is in the evaluation and comparison of forecasts, as we shall see in the third part of the book.

The specification of the information set is particularly important to the notion of a trade-off between bias and variance that is key to understanding many forecasting problems. We next use the special case with a linear forecasting model and MSE loss to illustrate this point. We assume that not all possible predictor variables have been included in the model.

Suppose there are K potential predictor variables and let be the random variable that captures current and past values of all the K potential predictor variables, including past and present values of the outcome, Similarly, let be the random variable generated by a subset of predictors. Finally, define as the random variable generated by the subset of omitted variables, so . Note that different expressions are obtained in equation (3.15), depending on whether we integrate over all of or over only a subset of it.

Example 3.2.6 (Risk for subset regressions). Consider the same problem as in Example 3.2.3. However, suppose that the forecasting model includes only regressors, including a constant unless all variables (including y) have mean 0.

The least squares estimators for β can be written as , where has the same dimension as with 0s in place of estimates for the omitted variables and the least squares estimators from a regression on the included regressors in the remaining places partitioned according to . Accounting for estimation, the forecast error is given by

where is the expected value of conditional on the omitted regressors. This is the standard omitted variable bias formula, where the mean of the ordinary least squares (OLS) estimates is a function of the omitted variables.

Integrating the squared forecast error only over conditional on the expected loss becomes

The second term is a variance term related to estimation error. Using the same math as in Example 3.2.3, this second term is approximately equal to . The third term is a bias term.

We refer to the second and third terms in (3.21) (the terms inside the large square brackets) collectively as the “estimation effect terms.” Notice that nothing (apart from collecting more data and hence further enlarging can be done to improve on the first term, although this avenue for improving forecasts should never be neglected.

This formulation of the problem results in the usual bias–variance trade-off that is common in calculations of the MSE. When forecasts are built from linear regressions and a constant is included, one can consider the notion of a biased forecast only when conditioning on potential omitted predictors. Nearly all of the choices that arise in the construction of forecasting models, regardless of the loss function, revolve around the cost and benefit from adding predictor variables. Under MSE loss we see that there is a trade-off between the variance (the second term in (3.21)) and the squared bias (the third term).

Adding variables increases the loss through the variance term, which increases by . However, the potential benefit is that the bias term could well be smaller as a result of the estimation of an extra parameter of .

The variance term disappears at rate and so is of a smaller order (as a function of the sample size) than the first term in (3.21). For egregiously misspecified regressions, the bias term could be of the same order as the first term, and hence for large enough biases this term can dominate the variance term by an order of magnitude. Note, however, that sensibly applied statistical tests can usually be used to rule out such situations. Recall that under standard assumptions (and certainly for our i.i.d. examples) the power of hypothesis tests for the significance of the√ regression coefficients grows at rate . This means that statistical tests for model misspecification have good power in situations where the bias term is of a greater order than the variance term. When statistical tests find it difficult to distinguish between models, the bias–variance trade-off becomes more complicated.

Notice that the risk in (3.21) is a function of . Alternatively, we could have considered integrating over all of , to obtain a different expression for the risk, as we show in the next example.

Example 3.2.7 (Risk for subset regression, continued). Consider the same problem as in Example 3.2.6. We can write the full regression as

where and minimizes , where the expectation is computed over all random variables. The variance of is , where . Now the regression to be run is seen to be a “short” regression with a larger variance. It follows directly from the result in Example 3.2.3 that the risk is

Hence the scaling variable is larger than for the expression in Example 3.2.6.

The expression in (3.23) is the average of (3.21), averaged over the omitted variables, . In this version where we integrate out the omitted variables, the “estimation effect terms” are convoluted with the unpredictable component. The first approach is often more useful for comparing models. For methods where the variance term vanishes asymptotically as the sample size becomes large, often the result amounts to a comparison of versus . This holds, for example, for many approaches to out-of-sample forecast evaluation, covered in the third part of the book.

练习题

What does the specification of the information set affect in the context of risk?

A. The number of predictor variables
B. The precision of the forecasting model
C. The precise form of risk
D. The mean of the outcome variable

If represents the random variable capturing all potential predictor variables, what does represent?

A. The random variable generated by all predictors
B. The random variable generated by a subset of predictors
C. The random variable of the outcome only
D. The random variable generated by omitted predictors

In the forecast error formula, what does the term represent?

A. The estimation error
B. The omitted variable bias
C. The variance term
D. The conditional mean of the outcome

Which of the following are components of the expected loss formula?

A.
B.
C.
D. The number of predictors

The second term in the expected loss formula is related to the bias of the forecast.

Adding more predictor variables to a forecasting model always decreases the expected loss.

The term approximates the __________ term in the expected loss formula.

The trade-off between bias and variance is a common consideration in calculations of the __________.

Explain the relationship between the information set and risk in forecasting.

How does the inclusion of additional predictor variables affect the expected loss in a forecasting model?

Which of the following statements are true about the bias-variance trade-off in forecasting? (Select all that apply)

A. The bias-variance trade-off is only relevant under MSE loss.
B. Adding more predictors always reduces the bias term.
C. The trade-off involves balancing the variance term and the squared bias term.
D. The trade-off is a consideration in both linear and non-linear forecasting models.

In the context of risk and information sets, what is the primary challenge when extending results to conditional forecasts?

A. Determining the correct form of the information set
B. Constructing the distribution of the outcome given future values of conditioning variables
C. Estimating the parameters of the forecasting model
D. Choosing the appropriate loss function

When considering the risk for subset regressions, which of the following statements correctly describes the relationship between the variance term and the number of included predictors and sample size ?

A. The variance term is approximately equal to and increases with and decreases with .
B. The variance term is approximately equal to and increases with and decreases with .
C. The variance term is approximately equal to and decreases with and increases with .
D. The variance term is approximately equal to and decreases with and increases with .

Which of the following statements are true regarding the effect of adding variables to a forecasting model under MSE loss? Select all that apply.

A. Adding variables increases the variance term of the expected loss.
B. Adding variables decreases the bias term of the expected loss.
C. The increase in the variance term is .
D. The potential benefit of adding variables is that the bias term could be smaller.

In the context of risk for subset regressions, the mean of the ordinary least squares (OLS) estimates is a function of the omitted variables only when considering the unconditional forecast error.

The second term in the expected loss formula (3.21) is a ___ term related to estimation error, and using the same math as in the linear forecasting model example, it is approximately equal to .

登录后解锁笔记、知识点解析、AI 问答

立即登录