正在学习

1.1.2 Part II

1.1.2 Part II

Part II of the book provides an overview of the various approaches to forecasting that have become standard in many areas, including the economic and finance forecasting literature. Chapters are based around either the amount of information available—from only the past history of the predicted variable through very large panels of variables—or the general estimation approach, principally parametric or nonparametric methods. To the extent possible, we provide details of how to go about constructing forecasts from the various methods, or alternatively direct readers to explanations available in the literature. We also discuss the trade-offs between different methods. In this sense we endeavor to provide a “first stop” for practitioners wishing to apply the methods covered in this section.

An important insight that arises from the decision-theoretic approach is that there is no single best or dominant approach to constructing a forecast for all possible forecasting situations. We discuss the types of forecasting situations where each individual method is likely to be a reasonable approach and also highlight situations where other approaches should be considered.

Throughout the book, we use a variety of empirical applications to illustrate how different approaches work. In most applications we use so-called pseudo out-of-sample forecasts which simulate the forecast as it could have been generated using data only up to the date of the prediction. This method restricts both model selection and parameter estimation to rely on data available at the point of the forecast. As time progresses and more data become available, the forecasting method, including the parameter estimates, are updated recursively. Such methods are commonly used to evaluate the usefulness of forecasts; a critical discussion of such out-of-sample forecasting methods versus in-sample methods is provided in part three of the book.

When building a forecasting model for an economic variable, the simplest specification of the conditioning information set is the variable’s own past history. This leads to univariate autoregressive moving average, or ARMA, models. Since Box and Jenkins (1970) these models have been extensively used and often provide benchmarks that are difficult to beat using more complicated forecasting methods. Linear ARMA models are also easy to estimate and a large literature has evolved on how best to cover issues in implementation such as lag length selection, generation of multiperiod forecasts, and parameter estimation. We discuss these issues in chapter 7. The chapter also covers exponential smoothing, unobserved components models, and other ways to account for trends when forecasting economic variables.

Chapter 8 continues under the assumption that the information set is limited to the predicted variable’s own past, but focuses on nonlinear parametric models. Examples include threshold autoregressions, smooth threshold autoregressions, and Markov switching models. These models have been used to capture evidence of nonlinear dynamics in many macroeconomic and financial time series. Unlike nonparametric models they do not, however, have the ability to provide a global approximation to general data-generating processes of unknown form.

Chapter 9 expands the information set to include multivariate information by considering a natural extension to univariate autoregressive models, namely vector autoregressions, or VARs. VARs provide a framework for producing internally consistent multiperiod forecasts of all the included variables. As used in macroeconomic forecasting VARs typically include a relatively small set of variables, often less than 10, but they still require a large number of parameters to be estimated if the number of included lags is high. To deal with the resulting negative effects of estimation errors on forecasting performance, a large literature has developed Bayesian methods for estimating and forecasting with VARs. Both classical and Bayesian estimation of VARs is covered in the chapter which also deals with forecasting when the future paths of some variables are specified, a common practice in scenario analysis or contingent forecasting.

The emergence of very large data sets has given rise to a wealth of information becoming readily available to forecasters. This poses both a unique opportunity— the potential for identifying new informative predictor variables—but also some real challenges given the limitations to most economic data. Suppose that N potential predictor variables are available, and that N is a large number, i.e., in the hundreds or thousands. Including all variables in the forecasting model—the so-called kitchen sink approach—is generally not feasible or desirable even for linear models since parameter estimation error becomes too large, unless the length of the estimation sample, T , is very large relative to N. Standard forecasting methods that conduct comprehensive model selection searches are also not feasible in this situation. If the true model is sparse, i.e., includes only few variables, one possibility is to use algorithms such as the Lasso, covered in chapter 6, to identify a few key predictors. Another strategy is to develop a few key summary measures that aggregate information from a large cross section of variables. This is the approach used by common factor models. Chapter 10 describes how these methods can be used in forecasting, including in factor-augmented VAR models that include both univariate autoregressive terms along with information in the factors. Finally, we discuss the possibility of using methods from panel data estimation to generate forecasts.

While chapters 7–10 focus on parametric estimation methods and so assume that a certain amount of structure can be imposed on the forecasting model, chapter 11 considers nonparametric forecasting strategies. These include kernel regressions and sieve estimators such as polynomials and spline expansions, artificial neural networks, along with more recent techniques from the machine-learning literature such as boosted regression trees. Although these methods have powerful abilities to approximate many data-generating processes as the number of terms included by the approach gets large, in practice any given estimated nonparametric model is itself an approximation to this approximation. Notably, the number of terms that can be successfully included in empirical applications will often be severely restricted by the available data sample. These approximate models thus do not have the same approximation ability as the models and thus themselves are approximations. Once again, the algorithm used to fit these forecasting models—along with the loss function used to guide the estimation—become key to their forecasting performance and to avoiding issues related to overfitting.

Forecasts of binary variables, i.e., variables that are restricted to take only two possible values, play a special role in decisions such as households’ choice on whether or not to buy a car, the decision on whether to pursue a particular education, or banks’ decisions on whether to change interest rates for short-term deposits. Restricting the outcome to only two possible values has the advantage that it crystallizes the costs of making wrong forecasts, i.e., false positives or false negatives. Chapter 12 takes advantage of these simplifications to cover point and probability forecasts of binary outcomes and discusses both statistical and utilitybased estimators for such data.

The decision-theoretic approach embodies a loss function that is appropriate for the decision to be made and not, as is so often the case, chosen for convenience. It results in a decision, i.e., a choice of an action to be made. This directs itself to basing estimation on an objective of providing the best decision. Alternatively, we might consider provision of a predictive distribution (density forecast) for an outcome as the objective of the forecasting problem. In chapter 13 we see that this perspective is useful for a wide range of decisions. Distribution forecasts also serve the important role of quantifying the degree of uncertainty surrounding point forecasts.

Distributional forecasting fills an important place in any forecaster’s toolbox but it does not replace point forecasting. First, although density forecasts can be used to construct point forecasts, typically it is the point forecast or decision that is required. Second, distributional forecasts rely on the distribution being estimated from data. This brings the loss function or scoring rule—the loss function used to estimate the density—back into the problem. Often ad hoc loss functions are employed to estimate the distributional forecast, leading to problems when the distributional forecast is subsequently used to construct the point forecast.

Given the plethora of different modeling approaches for construction of forecasts throughout chapters 7–13, it is not surprising that forecasters frequently have access to multiple predictions of the same outcome. Instead of aiming to identify a single best forecast, another strategy is to combine the information in the individual forecasts. This is the topic of forecast combinations covered in chapter 14. If the information used to generate the underlying forecasts is not available, forecast combination reduces to a simple estimation problem that basically treats the individual forecasts as predictors that could be part of a larger conditioning information set. Special restrictions on the forecast combination weights are sometimes imposed if it can be assumed that the individual forecasts are unbiased. If more information is available on the models underlying the individual forecasts, model combination methods can be used. These weight the individual forecasts based on their marginal likelihood or some such performance measure. Bayesian model averaging is a key example of such methods and is also covered in this chapter.

练习题

What is the main insight from the decision-theoretic approach regarding forecasting methods?

A. There is a single best approach for all forecasting situations.
B. The best approach depends on the amount of information available.
C. There is no single best approach for all forecasting situations.
D. Nonparametric methods are always superior to parametric methods.

What are pseudo out-of-sample forecasts used for?

A. To generate forecasts using all available data.
B. To simulate forecasts using data only up to the date of the prediction.
C. To evaluate in-sample forecasting methods.
D. To compare parametric and nonparametric methods.

Which of the following are characteristics of univariate ARMA models? (Select all that apply)

A. They use only the past history of the predicted variable.
B. They are difficult to estimate.
C. They often provide benchmarks that are hard to beat.
D. They are nonparametric models.

Which of the following are examples of nonlinear parametric models? (Select all that apply)

A. Linear regression models.
B. Threshold autoregressions.
C. Smooth threshold autoregressions.
D. Markov switching models.

Vector autoregressions (VARs) are limited to including only a few variables.

Including all potential predictor variables in a forecasting model is generally feasible and desirable.

The simplest specification of the conditioning information set for building a forecasting model for an economic variable is the variable’s own ___.

The Lasso algorithm can be used to identify a few key predictors when the true model is ___.

Explain the trade-off between using a large number of predictor variables in a forecasting model and the accuracy of parameter estimation.

What is the main advantage of using pseudo out-of-sample forecasts in evaluating forecasting methods?

Which of the following statements are true regarding the use of pseudo out-of-sample forecasts and the decision-theoretic framework in forecasting?

A. Pseudo out-of-sample forecasts simulate the forecast as it could have been generated using data only up to the date of the prediction.
B. The decision-theoretic framework assumes that the forecaster's loss function is irrelevant to the forecasting process.
C. Both pseudo out-of-sample forecasts and the decision-theoretic framework are used to evaluate the usefulness of forecasts.
D. The decision-theoretic framework emphasizes that there is no single best forecasting approach for all situations.
E. Pseudo out-of-sample forecasts rely on data that becomes available after the forecast date.

When building a forecasting model for an economic variable, the simplest specification of the conditioning information set is the variable's own past history, leading to ___ models.

登录后解锁笔记、知识点解析、AI 问答

立即登录