正在学习

14.4.1 Complete Subset Regressions

14.4.1 Complete Subset Regressions

Elliott, Gargano, and Timmermann (2013) propose an estimation method that uses equal-weighted combinations of forecasts based on all possible models restricted to include a fixed (given) number of regressors, . The approach, which they name complete subset regressions, first regresses the outcome, on a given subset of the regressors, then averages the results across all k-dimensional subsets of the regressors.

Assuming K regressors in the full model and k regressors chosen for each of the subset models, there will be subset regressions to average over. For example, the univariate case, , yields regressions, each of which includes a single variable. All elements of are 0 except for the least squares estimate of on in the ith row. The equal-weighted combination of forecasts from the individual univariate regression models becomes

Subset regression coefficients can be computed as averages over least squares estimates of the subset regressions. When the covariates are correlated, the individual regressions will be affected by omitted variable bias, but Elliott, Gargano, and

Timmermann (2013) show that the subset regression estimators are themselves approximately a weighted average of the full regression OLS estimator, . Assume that K is fixed and let be a matrix with 0s everywhere except for 1s in the diagonal cells corresponding to included variables, so that if the element of is 1, the j th regressor is included, while if this element is 0, the j th regressor is excluded.

Provided that as the sample size gets large, for some and , the estimator for the complete subset regression, , can be written as

where

In the special case where the covariates are orthonormal and with being a scalar, subset regression reduces to Ridge regression.

The amount of shrinkage implied by depends on both k and K and the smaller is k relative to , the greater the amount of shrinkage. This result relates shrinkage provided by model averaging to shrinkage on the individual coefficients whereas a typical Bayesian approach would separate the two. In general reduces to the Ridge estimator, either approximately or exactly, only if the regressors are uncorrelated. If this does not hold, subset regression coefficients cannot be interpreted as the outcome of simple regressor-by-regressor shrinkage of the OLS estimates, and instead depend on the full covariance matrix of all regressors.

Figure 14.8 illustrates the complete subset regression approach for our earlier application to US stock returns. We use a list of 11 possible predictor variables, so in this application. Along the horizontal axis we plot k, the number of included predictors which varies from (the prevailing mean) to (kitchen sink model). The root mean squared error is plotted on the vertical axis. For each value of k, the column of circles shows the RMSE values of all models with exactly k predictors, whereas the solid triangle shows the average RMSE value computed across these models. For example, there are 11 different RMSE values for and for , and 55 different RMSE values for and for . The range of values is larger in the middle, reflecting in part the bigger set of models being compared for these middle values of k. Note, however, that the average RMSE performance for a given value of k is trending upwards as k grows larger. This reflects the increased importance of estimation error for larger sets of regressors.

The curved line underneath the circles gives the RMSE value for the complete subset regressions corresponding to a given value of k. Notably, this line is lower than the lowest RMSE value of any individual model with the same number of regressors. The minimum value of the complete subset regression line is reached for . At this point (as well as for , the RMSE of the complete subset regression forecast is lower than that of the prevailing mean. It is also lower than the RMSE associated with the equal-weighted average forecast computed across all models, which tends to put too much weight on models with large numbers of regressors and poor forecasting performance.


Figure 14.8: Complete subset regression combination forecasts of stock returns using 11 predictors.

In settings with large-dimensional sets of predictors with weak power, Elliott, Gargano, and Timmermann (2015) show analytically that the complete subset regression approach achieves variance reduction and find empirically in applications to US unemployment, GDP growth, and inflation (as well as in Monte Carlo simulations) that the approach establishes a more favorable bias–variance trade-off than univariate models or dynamic factor models.

练习题

What is the key characteristic of complete subset regressions as proposed by Elliott, Gargano, and Timmermann (2013)?

A. They use only the best-fitting model for forecasting.
B. They use equal-weighted combinations of forecasts from all possible models with a fixed number of regressors.
C. They focus on models with the maximum number of regressors.
D. They ignore the correlation between covariates.

If there are regressors in the full model and regressors chosen for each subset model, how many subset regressions will there be to average over?

A.
B.
C.
D.

In the univariate case (), how many regressions will there be if there are regressors in the full model?

A.
B.
C.
D.

Which of the following statements about subset regression coefficients are correct?

A. They are computed as averages over least squares estimates of the subset regressions.
B. They are unaffected by omitted variable bias when covariates are correlated.
C. They can be interpreted as the outcome of simple regressor-by-regressor shrinkage of the OLS estimates when regressors are uncorrelated.
D. They are always equal to the full regression OLS estimator.

Which of the following are true about the complete subset regression estimator formula?

A. It is written as .
B. is a scalar value.
C. depends on the covariance matrix of the regressors.
D. The formula assumes that the sample size is small.

In the special case where the covariates are orthonormal, subset regression reduces to Ridge regression.

The amount of shrinkage implied by is independent of the values of and .

The equal-weighted combination of forecasts from the individual univariate regression models is given by the formula . Here, represents the ___.

The matrix in the complete subset regression estimator formula is a matrix with 0s everywhere except for 1s in the diagonal cells corresponding to ___.

Explain why the subset regression estimators are approximately a weighted average of the full regression OLS estimator when covariates are correlated.

In complete subset regressions, when and , how many subset regressions are there to average over?

A.
B.
C.
D.

Which of the following statements about complete subset regressions is correct when the covariates are orthonormal?

A. Subset regression has no relation to Ridge regression.
B. Subset regression reduces to Ridge regression with being a scalar.
C. The subset regression coefficients cannot be computed as averages over least squares estimates.
D. The amount of shrinkage is independent of and .

In complete subset regressions, when the covariates are correlated, the individual regressions are not affected by omitted variable bias.

In complete subset regressions, the equal - weighted combination of forecasts from the individual univariate regression models is given by , where is the number of ___.

登录后解锁笔记、知识点解析、AI 问答

立即登录