正在学习

11.4 CONCLUSION

11.4 CONCLUSION

Sieve methods such as artificial neural networks have powerful approximating capabilities and are able to approximate a very wide set of nonlinear models. This property relies on including an infinite number of terms. However, it is worth realizing that any estimated sieve model is itself an approximation to this infinite-order approximation. Finite sample applications approximate the infinite series with a finite series and must rely on estimated parameters that are then subject to estimation error. As a result, there is a real risk that sampling error overwhelms any gains from being able to approximate the underlying model in a more flexible manner. Model selection methods can be used to decrease the effects of overparameterization, but the actual usefulness of nonparametric forecasting methods over other less flexible methods needs to be evaluated rather than taken for granted.

The main difficulty with nonparametric methods is the so-called curse of dimensionality. Even with a single predictor, these methods are not particularly parsimonious, and with multiple predictors a very large number of parameters may have to be estimated (in the series-type approaches) or be influenced by sparse data (in the kernel approaches). Methods such as projection pursuit regression attempt to avoid this problem, but there is limited experience with their performance in forecasting.

Another issue is that it can be difficult to interpret the results from nonparametric estimation. This need not be an issue to the extent that the methods are regarded as an automatic forecasting device. However, if interpretability of the model’s estimated relations is important, nonparametric methods are not as straightforward to use as parametric methods. In part, such issues can be addressed by using graphs to illustrate how input variables map into predictions, although this involves more effort than simply looking at estimates of individual regression coefficients.

Binary Forecasts

Many economic applications involve data with a restricted support. Examples include a consumer’s decision to buy an automobile chosen from a small set of possibilities or the Fed’s decision to change the Federal funds rate in multiples of 25 basis points. When the outcome variable has restricted support, this needs to be taken into account in constructing as well as evaluating a forecasting model. It also has implications for the choice of loss function. For illustration, consider the binary prediction problem in which the observed outcome, Y, takes one of two values. For example, we might forecast whether there is a recession, whether an asset price rises, whether a company or a household goes bankrupt, or whether a student is admitted to a college. In each of these cases a particular outcome either happens or does not happen. Clearly the forecast cannot sensibly take any value on the real number line for these cases.

As for all forecasting problems, we can consider either density forecasts or point forecasts of the outcome. In the binary model the density of Y given z is fully described by the conditional mean, so the literature on binary forecasts has focused on density forecasts.1 Moreover, in contrast with the majority of forecasting problems, the conditional mean is not (except in trivial cases) a possible outcome for binary variables and hence does not provide a point forecast for Y. Standard methods for obtaining the conditional mean therefore result in a density forecast rather than a point forecast. This unusual dichotomy extends to forecast evaluation; most methods for forecast evaluation of binary variables amount to density rather than point forecast evaluation.

Interesting issues arise for point forecasting of binary outcomes. First, there is a direct relationship between the density forecast and the point forecast. Since both the outcome variable and forecast can take only two possible values, the loss function is extremely simple and so it is much easier to match the loss function with the losses that could be incurred. For scoring rules—loss functions for distributional forecasting—direct relationships between utility functions and proper scoring rules have been developed. For point forecasting, the simplicity of the loss function gives rise to a fixed class of reasonable loss functions. Although the loss function is simple, it is not completely well behaved because it depends on the sign function, which is discontinuous. This makes estimation somewhat more challenging and has led to shortcuts in estimation which in many cases could result in poor forecasting models.

Much of the early work on binary predictions comes from weather forecasting and it is useful to bridge different areas of the forecasting literature by including a note on differences in nomenclature. In the weather forecasting literature the unconditional distribution of Y is called a climatological forecast and is often used as a baseline; a point forecast is referred to as a deterministic forecast, while a distributional forecast is called a probabilistic forecast.

Chapter 12 proceeds as follows. Section 12.1 covers point and probability forecasts in the case with binary outcomes, while section 12.2 discusses density forecasting for binary variables. The estimation and construction of point forecasts for binary outcome variables is further discussed in section 12.3 which also covers an empirical application. Section 12.4 presents an empirical application and section 12.5 concludes.

练习题

What property allows sieve methods like artificial neural networks to approximate a wide set of nonlinear models?

A. Including a finite number of terms
B. Including an infinite number of terms
C. Using a single predictor
D. Having a limited number of parameters

What is an estimated sieve model itself an approximation to?

A. A finite-order approximation
B. A linear model
C. An infinite-order approximation
D. A parametric model

What is the main risk associated with finite sample applications of sieve models?

A. Overfitting the data
B. Underfitting the data
C. Sampling error overwhelming gains from flexible approximation
D. Inability to approximate any model

What is the purpose of model selection methods in sieve models?

A. To increase the number of parameters
B. To decrease the effects of overparameterization
C. To make the model more complex
D. To eliminate the need for estimation

What are the main difficulties with nonparametric methods? (Select all that apply)

A. They are too simple
B. The curse of dimensionality
C. They are always parsimonious
D. Difficulty in interpreting results
E. They require a small number of parameters

Nonparametric methods are as straightforward to use as parametric methods when interpretability of the model's estimated relations is important.

Using graphs can completely eliminate the issue of interpreting results from nonparametric estimation.

In the binary prediction problem, the observed outcome, , takes one of ___ values.

In the binary model, the density of given is fully described by the ___.

Explain the unusual dichotomy in binary forecast evaluation.

What is the relationship between the density forecast and the point forecast in binary outcomes?

Which of the following are true about point forecasting of binary outcomes? (Select all that apply)

A. The loss function is extremely simple
B. The loss function depends on a continuous function
C. It is easy to match the loss function with potential losses
D. The loss function is well - behaved and continuous
E. Estimation can be challenging due to the nature of the loss function

Which of the following statements are true about nonparametric methods and the curse of dimensionality? (Select all that apply)

A. Even with a single predictor, nonparametric methods are not very parsimonious
B. With multiple predictors, a small number of parameters need to be estimated in series - type approaches
C. With multiple predictors, kernel approaches are not influenced by sparse data
D. Projection pursuit regression attempts to address the curse of dimensionality
E. The curse of dimensionality only occurs with a large number of predictors

When using sieve methods like artificial neural networks for forecasting, which of the following statements is correct regarding the trade-off between flexibility and estimation error?

A. Increasing the number of terms in the sieve model always improves forecast accuracy without introducing any drawbacks.
B. Finite sample applications of sieve models approximate infinite series with finite series, introducing estimation error that may overwhelm gains from flexibility.
C. Sieve models with fewer parameters are always preferable because they eliminate the curse of dimensionality entirely.
D. The curse of dimensionality only affects kernel-based methods, not artificial neural networks.

Which of the following statements are correct regarding the challenges of nonparametric forecasting methods?

A. Nonparametric methods are immune to the curse of dimensionality.
B. Projection pursuit regression attempts to address the curse of dimensionality but has limited forecasting experience.
C. Interpretability of nonparametric estimation results is straightforward, similar to parametric methods.
D. Kernel-based methods can be influenced by sparse data when multiple predictors are used.
E. Boosted regression trees are inflexible and cannot capture local features of the data.

Explain how the curse of dimensionality affects nonparametric methods and how projection pursuit regression attempts to address this issue.

登录后解锁笔记、知识点解析、AI 问答

立即登录