正在学习
11.2 ESTIMATION OF SIEVE MODELS
11.2 ESTIMATION OF SIEVE MODELS
Sieve estimation, due to Grenander (1981), seeks to approximate an unknown function by combinations of functions of the data. Applied to nonparametric forecasting, the idea is to construct a forecasting model of the form
where is the number of scalar nonlinear terms, each denoted , and are weights on those terms. Variations in choice of the functions capture different approaches to the nonlinear approximation. Note that we could set , although often it is better in practice to let this coefficient be estimated. The inclusion of a linear lead term reflects that linear models often perform well empirically and many of the choices for the functions do not approximate linear models well unless is very large. The image of a sieve is useful here—think of a larger number of terms as a more finely meshed sieve; the number of functions the sieve will “catch” is larger for large values of because it is able to better approximate a wider set of possible functions. A sieve with fewer terms will catch fewer functional forms.
Once the functions in (11.4) are determined, estimation involves solving an extremum problem based on the loss function. Specifically, the free parameters , and perhaps the fineness of the mesh, , are chosen by minimizing the average sample loss,
Most often the loss criterion is MSE, and so sieve estimators will approximate the conditional mean of given . Nonlinear regression can be used in this case to construct estimates of the forecasting model.
Theoretical results suggest that for the right choice of functions, can become arbitrarily small as gets large for any member of a given set of families of models. This suggests that as , the forecasting model becomes a better and better approximation to any model in some set of nonlinear models, . This feature of being a good approximator for a very large set of models is often used to justify the use of nonparametric methods. However, any empirical application must use a finite number of terms in the approximation, thus introducing both approximation and estimation errors.
The number of terms in (11.4), , is subscripted by the sample size, because the approximating properties are established asymptotically by letting the number of terms increase with the sample size. From a practical perspective, needs to be small relative to the sample size to ensure that estimation of the parameters is not so imprecise that it overwhelms any gains from obtaining a better approximation to the optimal forecast model. If is chosen to be too small, the model can be a poor approximation to the true conditional mean even when the model has attractive properties for larger values of
Complications might arise when estimating which, depending on the selected functional forms, can enter the specification nonlinearly and hence requires using search methods. Notice that for a given value of , the model is linear in and hence these parameters can be estimated by OLS or by any of the shrinkage methods examined in chapter 6.
Sieve methods, of which there are many, differ in their specification of the functional form for . These functions have the property that as gets large the models are dense in some set M. Not all choices of functions have useful properties in this regard. The desire for good approximating properties leads directly to the use of basis functions as specifications for . Basis functions can be mutually orthogonal or correlated. We next provide a partial list of functions that can be employed.
11.2.1 Polynomials
When the form of the forecasting model, , is unknown, a seemingly reasonable approach is to take a Taylor-series approximation to the function. For a univariate predictor, , this suggests using the forecasting model,
Although this estimator can have useful approximation properties— particularly if the true function is very smooth—in most cases it is really a local approximation. When the parameters are estimated over a range of data, the estimated weights are not the local parameters at any particular point but instead an average of them. In practice, the method can therefore prove to be a poor predictor. To remedy that the parameter estimates are a data-weighted average obtained over many points, estimation can focus on a single point by extending the local linear regression in (11.3) to include polynomial terms. This approach is the most common form of application in empirical work.
练习题
What is the primary purpose of sieve estimation in nonparametric forecasting?
In the sieve estimation forecasting model, what does represent?
Which of the following are true about the role of the linear lead term in sieve estimation?
In sieve estimation, a larger number of terms results in a more finely meshed sieve, which can approximate a wider set of possible functions.
The free parameters in sieve estimation are chosen by minimizing the average sample loss, denoted as . This is known as solving an __________ problem.
Explain why sieve estimators approximate the conditional mean of given .
According to theoretical results, what happens to as gets large?
Any empirical application of sieve estimation must use a finite number of terms, which introduces both approximation and estimation errors.
Which of the following are considerations when choosing in sieve estimation?
What complications might arise when estimating in sieve estimation, and how can they be addressed?
What is a common approach to improve polynomial approximation in sieve estimation?
Polynomial approximation in sieve estimation is always a good predictor because it captures the true function very smoothly.
When the form of the forecasting model is unknown, a Taylor-series approximation suggests using the model for a univariate predictor . This approach is known as using __________.
Which of the following statements about basis functions in sieve estimation are true?
Explain the relationship between parametric and nonparametric forecasting models in the context of sieve estimation.
In sieve estimation, what is the role of the linear lead term in the forecasting model ?
What is the primary reason for choosing to be small relative to the sample size in sieve estimation?
Sieve estimators approximate the conditional mean of given when the loss criterion is MSE.
In sieve estimation, the number of terms is subscripted by the sample size because the approximating properties are established ___.
登录后解锁笔记、知识点解析、AI 问答
立即登录