正在学习
9.7 CONCLUSION
9.7 CONCLUSION
Vector autoregressions provide a coherent way to generate internally consistent multiperiod forecasts that account for concurrent and dynamic correlations across the included variables. While the basic VAR model is a powerful tool in the forecasting literature, once more than a handful of variables are included, the number of parameters that require estimation increases rapidly. This can sometimes cause parameter estimation error to become larger. Different methodologies have evolved to handle issues related to such estimation errors, most notably Bayesian estimation methods. Baillie (1979) derives expressions for the asymptotic mean squared error of multistep forecasts generated by VARs and ongoing research continues to address the importance of controlling the risk of multivariate prediction models.
Important progress has been made in recent years in the areas of computational methods and estimation techniques. Bayesian methods as well as factor augmentation methods are now available to handle large conditioning information sets. The use of economically motivated constraints through DSGE models is another area where important progress has been made which can facilitate the use of economic theory. Such models can be used to conduct conditional forecast evaluation experiments and also have the potential for reducing the effect of estimation error.
Forecasting in a Data-Rich Environment
n many forecasting exercises, the forecaster has access to a wealth of potentially Irelevant variables that could be used for constructing a forecasting model. Just as often the forecaster has no strong theoretical reasons for excluding many of these variables from the model. This plethora of information is the result of different ways of measuring particular variables as well as different levels of information aggregation. Many economic variables, e.g., interest rates, output, or employment, lend themselves to numerous ways of measurement and reporting. For example, as a measure of interest rates one could use overnight bank rates, three-month Libor, T-bill rates at various points of the term structure, commercial paper rates, forward or swap rates. We often have no particular reason for choosing one measure over another. Turning to the question of aggregation, one could use a single measure for aggregate output such as real GDP or decompose this into sectors, subsectors, or geographical regions to obtain a finer-grained measure of output. The result is that often the dimension of reasonable predictor variables to include can be very large, even outstripping the sample size available for constructing the forecasting model.
When the dimension of the conditioning information variables is large relative to the sample size, most of the methods from previous chapters cannot be directly applied. Linear regression, VAR, or VARMA models and conventional covariancebased estimation methods are not feasible when there are more parameters to be estimated than observations available. Less parametric approaches are even less viable. And even if this constraint does not bind, when the number of estimated parameters is large relative to the sample size (but still smaller) the parameters are likely to be estimated too imprecisely for the construction of forecasting models that perform well. Many of the model selection methods also become infeasible to implement, as there are so many potential models to be considered.
Faced with these challenges, we need methods for constructing more parsimonious models. In situations where a linear model is entertained and it is expected that the best forecasting model is sparse—by which we mean that most of the variables are not useful for forecasting—methods such as the Lasso (chapter 6) can be used. These methods work best when the variables are not highly correlated, which in practice rules out many economic applications. Another approach attempts, on a variableby-variable basis, to select a subset of predictor variables deemed to be of particular relevance to the target variable, usually by searching over a subset of the potentially relevant variables. This is part of the broader area of regularization methods surveyed by Ng (2013).
The factor methods described in this chapter take a different approach, one that attempts to extract the salient information in using dimensionality reduction techniques. Such methods are particularly useful when the variables in are collinear. Factor models take linear combinations of the original choosing a small set of linear combinations that capture most of the variation in . Such linear combinations, or factors, are then employed to forecast . Specifically, suppose the individual x-variables are linearly related to a set of common factors, so that
where is a vector of factor loadings (constants), is a vector of factors, and are idiosyncratic shocks. Each of the right-hand-side objects is unobserved and the two stochastic components are uncorrelated so they represent a decomposition of . Taking the expectation of the square of each side (after removing means) results in the variation in being decomposed into a factor component and an idiosyncratic component. The hope in using factor analysis for forecasting is that there exists a decomposition with small enough that the factors explain a large amount of the variation in , but with far fewer variables when the factors are used in the forecasting equation rather that the original data.
Factor modeling in macroeconomic analysis dates back to Geweke (1976) and Sargent and Sims (1977). Early work on latent factor dynamics includes Sargent and Sims (1977), Engle and Watson (1981), Sargent (1989), Stock and Watson (1989, 1991), and Quah and Sargent (1993). Many studies in the early literature keep N fixed and let . There is also a closely related literature in finance that investigates the presence of common factors in cross sections of stock returns; see Chamberlain and Rothschild (1983) and Connor and Korajczyk (1986). This setting is better represented by fixing T and letting . More recent studies on dynamic factor models, such as Stock and Watson (2002a) and Bai and (2002, 2006), let both N and T go to infinity.
The literature varies in the strictness of the assumptions imposed on the factor structure. Assuming that a small number of factors account for most of the common variation across economic variables, the remaining idiosyncratic variation is either uncorrelated—the case of an exact factor structure (Chamberlain and Rothschild, 1983)—or display substantially weaker correlation cross-sectionally than in the original data, which gives rise to an approximate factor structure.
From a forecasting perspective, factor analysis is a relatively model-free approach that does not require making strong assumptions about which variables matter; see, , Breitung and Eickmeier (2006). Indeed, the researcher can take a relatively agnostic view on which variables to include, although in practice the design of the initial set of predictor variables from which the common factors are extracted can be quite important. The approach can be expected to perform well if the predictive content of “most” variables is captured by the common factor whereas the variablespecific information has predictive value only for a few select variables, most notably lags of the predicted variable itself.
This chapter goes through methods for incorporating common factors in prediction models. Section 10.1 discusses dynamic and static factor models along with the role played by factors in prediction models (as conditioning information) before section 10.2 turns to estimation of factor models and section 10.3 discusses methods for determining the number of factors required to characterize a largedimensional data set. Section 10.4 discusses practical issues in factor construction and interpretation of factors while section 10.5 discusses empirical evidence. Finally, section 10.6 briefly covers how panel methods can be used to construct forecasts for data with a large cross-sectional dimension. Section 10.7 concludes.
练习题
Which of the following best describes the primary advantage of Vector Autoregressions (VAR) in forecasting?
What is a key challenge when the dimension of conditioning information variables is large relative to the sample size?
Which of the following are methods that have been developed to handle large conditioning information sets? (Select all that apply)
What are the potential consequences of having a large number of estimated parameters relative to the sample size in forecasting models? (Select all that apply)
Factor models are particularly useful when the variables in are collinear.
The Lasso method is most effective when the variables are highly correlated.
The _______ method attempts to extract the salient information in using dimensionality reduction techniques.
In situations where the best forecasting model is sparse, meaning most variables are not useful for forecasting, the _______ method can be used.
Explain why Bayesian methods are useful in handling large conditioning information sets.
What is the primary goal of using factor models in forecasting, and how do they achieve this?
Which of the following is a reason why traditional methods like linear regression, VAR, or VARMA models may not be feasible in high-dimensional settings?
Which of the following statements about DSGE models are correct? (Select all that apply)
When dealing with a data-rich environment where the number of predictor variables exceeds the sample size, which of the following approaches is most appropriate for constructing a forecasting model?
Which of the following statements correctly describe the advantages of using Bayesian methods in forecasting models? (Select all that apply)
登录后解锁笔记、知识点解析、AI 问答
立即登录