正在学习
11.4 CONCLUSION
11.4 CONCLUSION
Sieve methods such as artificial neural networks have powerful approximating capabilities and are able to approximate a very wide set of nonlinear models. This property relies on including an infinite number of terms. However, it is worth realizing that any estimated sieve model is itself an approximation to this infinite-order approximation. Finite sample applications approximate the infinite series with a finite series and must rely on estimated parameters that are then subject to estimation error. As a result, there is a real risk that sampling error overwhelms any gains from being able to approximate the underlying model in a more flexible manner. Model selection methods can be used to decrease the effects of overparameterization, but the actual usefulness of nonparametric forecasting methods over other less flexible methods needs to be evaluated rather than taken for granted.
The main difficulty with nonparametric methods is the so-called curse of dimensionality. Even with a single predictor, these methods are not particularly parsimonious, and with multiple predictors a very large number of parameters may have to be estimated (in the series-type approaches) or be influenced by sparse data (in the kernel approaches). Methods such as projection pursuit regression attempt to avoid this problem, but there is limited experience with their performance in forecasting.
Another issue is that it can be difficult to interpret the results from nonparametric estimation. This need not be an issue to the extent that the methods are regarded as an automatic forecasting device. However, if interpretability of the model’s estimated relations is important, nonparametric methods are not as straightforward to use as parametric methods. In part, such issues can be addressed by using graphs to illustrate how input variables map into predictions, although this involves more effort than simply looking at estimates of individual regression coefficients.
Binary Forecasts
Many economic applications involve data with a restricted support. Examples include a consumer’s decision to buy an automobile chosen from a small set of possibilities or the Fed’s decision to change the Federal funds rate in multiples of 25 basis points. When the outcome variable has restricted support, this needs to be taken into account in constructing as well as evaluating a forecasting model. It also has implications for the choice of loss function. For illustration, consider the binary prediction problem in which the observed outcome, Y, takes one of two values. For example, we might forecast whether there is a recession, whether an asset price rises, whether a company or a household goes bankrupt, or whether a student is admitted to a college. In each of these cases a particular outcome either happens or does not happen. Clearly the forecast cannot sensibly take any value on the real number line for these cases.
As for all forecasting problems, we can consider either density forecasts or point forecasts of the outcome. In the binary model the density of Y given z is fully described by the conditional mean, so the literature on binary forecasts has focused on density forecasts.1 Moreover, in contrast with the majority of forecasting problems, the conditional mean is not (except in trivial cases) a possible outcome for binary variables and hence does not provide a point forecast for Y. Standard methods for obtaining the conditional mean therefore result in a density forecast rather than a point forecast. This unusual dichotomy extends to forecast evaluation; most methods for forecast evaluation of binary variables amount to density rather than point forecast evaluation.
Interesting issues arise for point forecasting of binary outcomes. First, there is a direct relationship between the density forecast and the point forecast. Since both the outcome variable and forecast can take only two possible values, the loss function is extremely simple and so it is much easier to match the loss function with the losses that could be incurred. For scoring rules—loss functions for distributional forecasting—direct relationships between utility functions and proper scoring rules have been developed. For point forecasting, the simplicity of the loss function gives rise to a fixed class of reasonable loss functions. Although the loss function is simple, it is not completely well behaved because it depends on the sign function, which is discontinuous. This makes estimation somewhat more challenging and has led to shortcuts in estimation which in many cases could result in poor forecasting models.
Much of the early work on binary predictions comes from weather forecasting and it is useful to bridge different areas of the forecasting literature by including a note on differences in nomenclature. In the weather forecasting literature the unconditional distribution of Y is called a climatological forecast and is often used as a baseline; a point forecast is referred to as a deterministic forecast, while a distributional forecast is called a probabilistic forecast.
Chapter 12 proceeds as follows. Section 12.1 covers point and probability forecasts in the case with binary outcomes, while section 12.2 discusses density forecasting for binary variables. The estimation and construction of point forecasts for binary outcome variables is further discussed in section 12.3 which also covers an empirical application. Section 12.4 presents an empirical application and section 12.5 concludes.
练习题
What property allows sieve methods like artificial neural networks to approximate a wide set of nonlinear models?
What is an estimated sieve model itself an approximation to?
What is the main risk associated with finite sample applications of sieve models?
What is the purpose of model selection methods in sieve models?
What are the main difficulties with nonparametric methods? (Select all that apply)
Nonparametric methods are as straightforward to use as parametric methods when interpretability of the model's estimated relations is important.
Using graphs can completely eliminate the issue of interpreting results from nonparametric estimation.
In the binary prediction problem, the observed outcome, , takes one of ___ values.
In the binary model, the density of given is fully described by the ___.
Explain the unusual dichotomy in binary forecast evaluation.
What is the relationship between the density forecast and the point forecast in binary outcomes?
Which of the following are true about point forecasting of binary outcomes? (Select all that apply)
Which of the following statements are true about nonparametric methods and the curse of dimensionality? (Select all that apply)
When using sieve methods like artificial neural networks for forecasting, which of the following statements is correct regarding the trade-off between flexibility and estimation error?
Which of the following statements are correct regarding the challenges of nonparametric forecasting methods?
Explain how the curse of dimensionality affects nonparametric methods and how projection pursuit regression attempts to address this issue.
登录后解锁笔记、知识点解析、AI 问答
立即登录