正在学习
4.3 PARAMETRIC VERSUS NONPARAMETRIC ESTIMATION APPROACHES
4.3 PARAMETRIC VERSUS NONPARAMETRIC ESTIMATION APPROACHES
Restricting the set of forecast models to a specific parametric form results in a weaker notion of forecast optimality. This follows directly from the notion of maximization, since maximizing over a wider set of models cannot result in a worse choice in sample. In practice, forecasters often choose between alternate parametric models in some reduced set, M. Limiting M to include only a finite number of models yields what is known as a model selection problem. When the set of models is infinite-dimensional, we refer to the model choice as a semi-nonparametric problem.
Chapter 6 examines the problem of model selection in far more detail, including discussions of properties of different methods and conditions under which these properties hold. Here we make only a few points that are pertinent to the current discussion. First, although there are often methods that can consistently select the correct model—assuming that it is included in the set under consideration—it does not follow that the risk function is the same as that obtained if the model were truly known. Notions of forecast optimality under model selection are therefore again much weaker, if they can even be established.
When the joint distributions of the outcome Y and predictors Z are unknown, it is popular to use sieve M-estimation to obtain a forecast model. Such methods generalize the estimation procedure in (4.2) by replacing a known or hypothesized forecast function, with approximating functions that take the form , where the size of the parameter vector depends on the sample size, T, through the number of terms, Typically the functional form for g is chosen to have good approximating properties for a wide set of potential forecast functions. We examine these methods in more detail in chapter 11. Chen (2007, Theorem 3.1) gives results for consistent estimation of for under general conditions. The conditions are similar to those stated for M-estimation above— identification of the model at a single point , continuity of the loss functions in the parameters, and uniform convergence— although since we now deal with sequences of models, results are somewhat more difficult to establish. In addition, there are restrictions on the parameters of the sieve models, i.e., the While consistency of the model is of course useful to establish, the relationship between the forecast and is not obvious in any practical forecast situation.
4.4 CONCLUSION
This chapter relates notions of forecast optimality under the classical approach, when the conditional distribution , including the parameters is known, to what we obtain in practice when the model and its parameters are unknown. Model and parameter uncertainty fundamentally affect our understanding of forecast optimality in several ways.
First, estimation methods that use objective functions different from the forecaster’s loss function should generally be avoided since there is no obvious reason why the parameter estimates should be good at minimizing the forecaster’s risk. This is the reason why Granger (1993) warns against basing estimation and evaluation on different criteria, a point that has been made a number of times in the forecasting literature (e.g., Weiss (1996).) The only exception to this arises when the forecast model is based on MLE estimates from a correctly specified model. However, we view correct specification of the joint density of the data to be highly unlikely for most forecasting problems.
Second, even when a model is correctly specified up to an unknown set of parameters and the parameters are estimated with the relevant loss function, at best only a family of minimum risk forecast methods, rather than a single optimal forecast method, can be established. This point was established in the previous chapter, and applies more generally to all of the results here.
Third, even when a forecast is constructed from a parametric model, the best-case scenario is that we obtain the optimal forecast from a restricted set of forecasting models even when we use the correct loss function.
Fourth, even when the forecast is based on the correct loss function, heterogeneity in the data means that it is likely that the estimator for (the best parameter for the current forecasting problem) actually converges to (the average best parameter for the forecasting problem) with averages taken over the random variables that generated the data. In practice this difference could be small, but this is not necessarily the case. Of course if the data heterogeneity can be modeled, then improvements can be made in this direction. Chapter 19 examines examples of this nature.
Fifth, among forecast models estimated with the correct loss function, different estimators trade off risk across different regions of the parameter space in order to reduce the effect of parameter estimation error. This is the reason for considering restricted loss functions. Much of the effort in constructing good forecasts with limited data samples—such as the use of Bayesian vector autoregressions—comes from examining these second-order effects.
练习题
What is the result of restricting the set of forecast models to a specific parametric form ?
What is a model selection problem in the context of parametric models?
What are the conditions for consistent estimation of for under general conditions?
Under model selection, the risk function is the same as that obtained if the model were truly known.
When the joint distributions of the outcome and predictors are unknown, a popular method to obtain a forecast model is ___.
Explain the relationship between the forecast and in practical forecast situations.
What is the form of the approximating functions used in sieve M-estimation?
Which of the following are true about the size of the parameter vector in sieve M-estimation?
Sieve M-estimation methods generalize the estimation procedure by replacing a known or hypothesized forecast function with approximating functions.
Chen (2007, Theorem 3.1) provides results for consistent estimation of for under conditions similar to those stated for ___.
What is the main challenge in establishing the relationship between the forecast and ?
What is the primary reason for the weaker notion of forecast optimality when restricting the set of forecast models to a specific parametric form?
When using sieve M-estimation for unknown joint distributions, which of the following is a key assumption for consistent estimation of for ?
Which of the following statements are true regarding the relationship between model selection and forecast optimality?
In the context of sieve M-estimation, the functional form for is typically chosen to have good approximating properties for a wide set of potential forecast functions, and the size of the parameter vector depends on the sample size, , through the number of terms, . This approach is used when the joint distributions of the outcome and predictors are ___.
Explain why the relationship between the forecast and is not obvious in practical forecast situations when using sieve M-estimation.
登录后解锁笔记、知识点解析、AI 问答
立即登录