正在学习

7.2.2 Choice of Lag Orders

7.2.2 Choice of Lag Orders

So far we have assumed that the lag orders, p and were known. In some financial forecasting problems, appeals to market efficiency may lead one to choose zero lags or a low-order moving average process, to reflect market microstructure effects on the dynamics of security prices. For quarterly data one might want at least four lags to account for a seasonal component. In most situations, forecasters do not have a great deal of knowledge about these parameters, however, other than the broad notion under stationarity that shocks in the distant past have less of an impact on today’s value than more recent shocks. This, combined with the finding that estimating a large number of parameters relative to the sample size is likely to result in imprecise estimates, suggests working with parsimonious ARMA models that include relatively few lags.

Different methods can be used to empirically determine the values for and q or, if noncontiguous lags are considered, which specific lags to include. Box and Jenkins (1970) originally suggested a judgemental approach based on examining the autocorrelations and partial autocorrelations of the data. However, automated methods for lag selection in general specifications are now more commonly used.

A widely used approach is to employ model selection criteria such as those discussed in chapter 6. Each of these defines a family of methods distinguished by its own trade-off between goodness of fit, which improves as more lags get included, versus a penalty term that grows as an increasing number of parameters are used. Varying the choice of ARMA order results in different models, and we can define the set of such models as , where represents model k and the search is conducted over K different combinations of p and q. Define the squared standard deviation of the residuals from model k as . For linear ARMA specifications, information criteria take the form

Here counts the number of estimated parameters for model so if there is no constant term, while if a constant is included. Here is a penalty term that is a function of the sample size, T . The objective is to choose a model that minimizes (7.16). As discussed in section 6.3, popular information criteria


Figure 7.4: Recursive lag length selections for the Akaike and Bayes information criteria (AIC, BIC) applied to models with up to 12 autoregressive lags.

routinely reported in regression packages use the following penalty terms:

Criteriong(T)
AIC (Akaike, 1974)2T-1
BIC (Schwarz, 1978)ln(T)/T
Hannan Quinn (1979)2ln(ln(T))/T.

Since the penalty terms differ but the measure of fit, , is the same across the various methods, different information criteria often favor different models. For example, the penalty term for the BIC is greater than that for the AIC provided that which holds whenever the sample size is greater than seven observations. The BIC therefore tends to choose more parsimonious models with fewer parameters than the AIC.

Figure 7.4 shows selection of the lag order for autoregressive models fitted to the four series introduced earlier. To select the lag order we use an expanding estimation window starting in 1970 and ending in 2014. For simplicity we compare only lag orders selected by the AIC and BIC. A maximum of 12 lags is considered, and we assume contiguous lags, i.e., we cannot include lag k without also including lags . For the inflation rate series, the AIC selects 12 lags most of the time, while the BIC selects 3 lags most of the time, at least after 1980. For the less persistent stock return series, BIC selects no lags, while the AIC selects 1 or 2 lags. The difference between the lag length selection of the AIC and BIC is even more pronounced for the unemployment rate series for which the AIC initially selects 2–3 lags, followed by 10 lags after the mideighties. In contrast, the BIC selects only 2 lags most of the time. Finally, for the interest rate series, the lag length selection is quite unstable during the first decade, albeit with the AIC choosing more lags than the BIC, only for the two criteria to settle on 12 (AIC) and 8 (BIC) lags towards the end of the sample.


Figure 7.5: Recursive out-of-sample forecasts generated by autoregressive models with lag length selected by the AIC or BIC.

How important are such differences in lag length selection to forecasting performance? To address this issue, we plot the forecasts from the models selected by the AIC and BIC in figure 7.5. While there are some minor differences between the forecasts, the series are clearly very similar, with correlations ranging from 0.27 between the two forecasts of stock returns to 0.95 (inflation rate) and 0.99 (unemployment and interest rate forecasts).

The reason the differences in predicted values are so small, despite large differences in model specification, is that many of the series are highly persistent. For such variables, neighboring lags and will be close substitutes. The one variable considered here for which this is not true—stock returns—is also the variable for which we observe the greatest difference between the forecasts generated by the models selected by AIC versus BIC.

练习题

In financial forecasting, what might lead one to choose zero lags or a low-order moving average process?

A. High sample size
B. Market efficiency appeals
C. Large number of parameters
D. High autocorrelation

What is the recommended approach when forecasters do not have much knowledge about lag orders and ?

A. Use high-order models
B. Use parsimonious ARMA models
C. Ignore stationarity conditions
D. Use noncontiguous lags

Which method was originally suggested by Box and Jenkins for lag selection?

A. Automated methods
B. Model selection criteria
C. Judgemental approach based on autocorrelations
D. Penalty term analysis

What is the form of information criteria for linear ARMA specifications?

A.
B.
C.
D.

Which of the following are penalty terms used by popular information criteria?

A.
B.
C.
D.
E.

The BIC tends to choose more parsimonious models with fewer parameters than the AIC.

For quarterly data, at least eight lags are needed to account for a seasonal component.

The objective is to choose a model that minimizes , where counts the number of estimated parameters for model as if there is no ___.

Explain why the BIC generally selects more parsimonious models compared to the AIC.

What is the main advantage of using parsimonious ARMA models in financial forecasting?

When selecting the optimal lag order for an ARMA model using information criteria, which of the following statements is correct regarding the trade-off between goodness of fit and model complexity?

A. The AIC always selects more parsimonious models than the BIC because its penalty term is larger.
B. The BIC tends to choose models with more parameters than the AIC when the sample size .
C. The AIC and BIC differ only in their penalty terms, with the BIC's penalty term being .
D. The Hannan-Quinn criterion uses a penalty term of , making it identical to the AIC.

Which of the following factors influence the choice of lag orders and in an ARMA model? Select all that apply.

A. The desire to reflect market microstructure effects, which may lead to choosing zero lags or a low-order moving average process.
B. The need to account for seasonal components, which may require at least four lags for quarterly data.
C. The preference for models with a large number of parameters to improve goodness of fit, regardless of sample size.
D. The principle of parsimony, which suggests working with models that include relatively few lags to avoid imprecise estimates.
E. The use of information criteria, which balance goodness of fit with a penalty for model complexity.

The BIC tends to select more parsimonious models with fewer parameters than the AIC because its penalty term is greater than that of the AIC when .

In the context of ARMA model selection, the objective is to choose a model that minimizes the information criterion , where is the squared standard deviation of the residuals, is the number of estimated parameters, and is a penalty term that is a function of the sample size . For the BIC, ___$.

登录后解锁笔记、知识点解析、AI 问答

立即登录