正在学习

13.2.2 GARCH Models

13.2.2 GARCH Models

A large class of models for volatility forecasting make use of some variant of the autoregressive conditional heteroskedasticity (ARCH) model introduced by Engle (1982). Consistent with (13.3), suppose that the volatility of shocks to the predicted variable is time varying and takes the form

Here is the conditional variance of given information at time . The ARCH(q) model proposed by Engle (1982) takes the form

Since the -values are positive, a large shock, , will induce higher future values of the conditional variances, . ARCH models can therefore capture clustering in the volatility of forecast errors: large squared errors tend to be followed by large squared errors and small errors tend to be followed by small errors. The sign of future errors is not predictable, however, as the model does not specify whether positive or negative shocks are more likely to occur.

The ARCH(q) model may require a large number of lag coefficients, to fit series with highly persistent volatility. A more parsimonious model that can capture persistent volatility dynamics is the Generalized ARCH or GARCH specification introduced by Bollerslev (1986). The GARCH( p, q) model with p autoregressive lags of the conditional variance and q lags of squared innovations takes the form

In empirical work, by far the most popular specification is the GARCH(1,1) model:

The term in the final bracket has zero mean since . It follows from (13.8) that is a measure of the persistence of the conditional variance process, while is a measure of the “news” impact of shocks, i.e., the effect of current (normalized) shocks on the forecast of next period’s conditional variance.

Many studies have examined the empirical performance of the GARCH(1,1) model and found it difficult to outperform. Hansen and Lunde (2005) compare 330 ARCH-type specifications, including very sophisticated volatility models. They find no significant evidence that any of these models outperforms the GARCH(1,1) model for daily exchange rate predictions. Conversely, for predictions of daily returns on IBM, models that account for leverage effects (discussed below) are found to be better than the GARCH(1,1) model. Given the prominence of this class of models, we next cover details of how such models can be used to generate density forecasts.

Note that recursive backward substitution in the first line in (13.8) allows us to express the conditional variance forecast as a sum of lagged squared innovations:

Hence the GARCH(1,1) model implies an ARCH model of infinite order with a particular decay structure in the coefficients on past squared innovations. This is very similar to the geometric decay seen for the AR(1) model analyzed in chapter 7.

As long as , it is clear from the last line in (13.8) that the volatility process converges. Taking unconditional expectations on both sides, the steady state—or unconditional—variance (if it exists) must satisfy

Thus, assuming that the process is covariance stationary, and so

which is the unconditional variance. Using this, we can rewrite (13.8) as

This allows us to write the future one-step-ahead conditional variance as follows (for :

Taking expectations conditional on period-t information, the expected one-period variance h-periods ahead, is

Hence, if the one-step-ahead conditional variance forecast exceeds the average variance forecast, , the multi-step-ahead variance forecasts will exceed the average forecast by an amount that is declining in the forecast horizon. This mean-reverting tendency is of course a defining property of stationary processes and so the importance of the assumption that becomes clear from (13.11).

Equation (13.11) gives the expected value of the future conditional variance given current information. The conditional density of the future variance is more complicated to characterize. Moving one step ahead and ignoring that the parameters are unknown, conditional on period-t information the only unknown variable in the last line in (13.9) is and so next period’s one-step-ahead conditional variance forecast follows a chi-squared distribution with one degree of freedom:

Moving one period further ahead to , from (13.10) we have

Notice the presence of cross-product terms such as involving future shocks that are unknown at time t. Such terms mean that the conditional multiperiod variance is no longer distributed as a chi-squared variable. In fact, the GARCH process can be viewed as a mixture of normals with time-varying weights. This leads to a distribution that can be considerably more fat tailed than the normal distribution. For example, as shown by Bollerslev (1986), the coefficient of excess kurtosis for the GARCH(1,1) process, measured relative to the Gaussian distribution,

This expression is always positive and so shows that even a GARCH process with a conditionally Gaussian one-step-ahead forecast can generate fat tails for multiperiod or unconditional forecasts.

Example 13.2.2 (Estimation of GARCH parameters based on a loss function). Skouras (2007) proposes to the parameters of a GARCH(1,1) model so as to minimize the expected loss of the forecast as used in an investor’s portfolio decision. The investor can invest in a risky asset paying a random excess return over the risk-free rate, . The investor has initial wealth with wealth evolving according to the equation . Following Skouras (2007), consider the GARCH(1,1) model for the return on the risky asset,

Suppose this approach is used by an investor with constant absolute risk aversion (CARA) utility

where A is the investor’s coefficient of absolute risk aversion. As pointed out by Skouras, under these assumptions, the investor’s Bayes’ rule (optimal holding of the risky asset)

takes the form

where are the unknown parameters. If the investor estimates the parameters that minimize the sample average loss, we get what Skouras calls the maximum utility estimator,

For the GARCH(1,1) model this estimator takes the form

One issue with this estimator is that it identifies only the ratio of the conditional mean to the conditional variance. In empirical results based on a GARCH(1, 1) model, Skouras finds that the volatility dynamics implied by the empirical maximum utility estimates tend to be less persistent than their QMLE counterparts. When used in an out-of-sample investment rule, the maximum utility estimates often lead to higher average utility than under the QMLE approach, although this does not hold during the last 10 years of the sample analyzed by Skouras.

练习题

Which of the following correctly represents the ARCH(q) model equation?

A.
B.
C.
D.

In the ARCH model, what is the effect of a large shock on future conditional variances?

A. It decreases future conditional variances.
B. It has no effect on future conditional variances.
C. It increases future conditional variances.
D. It makes future conditional variances negative.

Which of the following is the most parsimonious model for capturing persistent volatility dynamics compared to the ARCH(q) model?

A. ARCH(1) model
B. GARCH(1,1) model
C. ARCH(2) model
D. GARCH(2,2) model

Which of the following statements are true about the GARCH(1,1) model?

A. It has the form .
B. measures the persistence of the conditional variance process.
C. measures the effect of past shocks on current conditional variance.
D. It cannot capture clustering in volatility.
E. It assumes that the sign of future errors is predictable.

The GARCH(1,1) model has been found to outperform many sophisticated volatility models for daily exchange rate predictions according to Hansen and Lunde (2005).

In the GARCH(1,1) model, if , the volatility process will converge.

The term in the rewritten GARCH(1,1) equation represents the sum of lagged squared innovations with a particular decay structure in the coefficients, implying an ARCH model of ___ order.

Explain the significance of in the GARCH(1,1) model.

Which of the following are reasons for the popularity of density forecasts over point forecasts? (Select all that apply)

A. Density forecasts provide a full distribution of possible outcomes.
B. Point forecasts are sufficient for all users with different loss functions.
C. Density forecasts convey the precision of the forecast.
D. Policy makers need full distribution forecasts to consider the range of possible outcomes.
E. Risk-averse investors prefer point forecasts as they are simpler.

What is the unconditional variance in the GARCH(1,1) model, and how is it derived?

Which of the following statements correctly describes the relationship between the and parameters in the GARCH(1,1) model? Assume the model is covariance stationary.

A. measures the persistence of the conditional variance process, while measures the impact of current shocks
B. measures the persistence of the conditional variance process, while measures the impact of current shocks
C. measures the unconditional variance, while measures the conditional variance
D. must be greater than 1 for the volatility process to converge

Which of the following are true about the GARCH(1,1) model's conditional variance forecast? Select all that apply.

A. It can be expressed as a sum of lagged squared innovations
B. It requires an infinite number of lag coefficients to fit persistent volatility
C. It implies an ARCH model of infinite order with geometric decay in coefficients
D. It assumes the conditional mean is constant over time
E. It converges if

In the GARCH(1,1) model, the term in equation (13.8) has a zero mean because .

The unconditional variance of the GARCH(1,1) model is given by , provided that ___.

Explain how the GARCH(1,1) model captures volatility clustering and why it is more parsimonious than the ARCH(q) model for highly persistent volatility.

登录后解锁笔记、知识点解析、AI 问答

立即登录