正在学习
18.1.3 Empirical Application to Volatility Forecasting
18.1.3 Empirical Application to Volatility Forecasting
Applications of density forecasting in finance have often focused on volatility modeling. As discussed in chapter 2, Patton (2011) establishes that the MSE and QLIKE loss functions can be used to consistently rank volatility forecasts even when the outcome is measured with noise but can be proxied through the realized variance. Patton (2011) suggests using Diebold–Mariano tests applied to these loss functions.
Table 18.1 provides an illustration of this approach applied to daily stock market returns measured by the S&P500 index. As benchmark we use the risk metrics model which downweights past squared returns exponentially by a factor along with a 300-day rolling window volatility estimator. These forecasts are compared to a GARCH(1,1) and an AR model with lag length selected by AIC estimated on the realized variance series, using MSE loss. The initial estimation window uses 10 years of data from 2000 to the end of 2009, while the out-of-sample evaluation period runs from 2010 through 03/31/2015 and so covers more than five years. Forecast precision is measured using the realized variance as a proxy for the outcome. Positive values of the Diebold–Mariano test show that the squared errors of the model listed first are higher than those of the model listed last. The results show that the AR model performs best and in fact beats the GARCH and rolling window models with values close to 5%. Conversely, the rolling window method performs quite poorly in this application.
18.2 EVALUATING FEATURES OF DISTRIBUTIONAL FORECASTS
An alternative to examining the density forecast by means of some loss function is to review certain features of the density forecast. The idea is to check whether certain features of the distribution of the outcomes line up well with features implied by the distributional forecasts. This can be done either by directly relating the predicted and “actual” distributions and examining whether they are identical, or by considering certain features of the probability distribution.
Feature-based evaluation has been most fully developed for the case where Y is either 0 or 1 and so is a binary random variable with a probability forecast . Murphy (1973) showed that the standard quadratic probability score, , can be decomposed as follows (see equation (15.3)):
The first term in (18.9) is the unconditional variance of the outcome variable and so is independent of the forecast; the second term is the average squared calibration error; the last term is known as the resolution term and is the squared average difference between the conditional mean given the forecast and the unconditional mean. Calibration and resolution are examined below; the hope is that calibration is small (0) while resolution is large, leading to a smaller MSE.
To compute the decomposition in (18.9), for the case where follows a discrete distribution we divide the line from 0 to 1 into m bins with centers Suppose again that we have (out-of-sample) observations to evaluate forecasting performance and observe observations in the i th bin so that . For each bin we predict of the outcomes to be equal to 1. For each i, let the outcomes be and let be the associated sample mean of the outcomes. For the special case with , Murphy (1973) suggests the following decomposition:
A good forecast, i.e., one with a small MSE, has calibration as close to 0 as possible and a resolution as high as possible. Calibration and resolution are correlated and it can be unclear how to adjust these measures to minimize the mean squared loss since any attempts to minimize the calibration term might also decrease the resolution.
To derive the decomposition in (18.10), note that
Here we used that . The first term is the calibration (ignoring the scaling by , while the second term is the resolution.
Other decompositions of the Brier score exist (see, e.g., Sanders, 1963), one of which results in a property known as sharpness which refers to how concentrated the distributional forecast is. Sharpness is a feature of the forecasts themselves rather than how they relate to the outcomes. It is commonly evaluated based on a histogram of the probability forecasts on [0, 1]. A sharp forecast has a lot of probability mass near 1 or 0, the idea being that the forecast is giving decisive signals on which outcome will occur.
Extensions from the binary problem to general distributions have only recently been tackled. Our discussion reflects the small formal literature on such extensions and ignores some less used properties of density forecasts. We first cover calibration, then resolution, and finally sharpness.
练习题
According to Patton (2011), which loss functions can be used to consistently rank volatility forecasts even when the outcome is measured with noise?
What test does Patton (2011) suggest applying to the MSE and QLIKE loss functions for volatility forecasts?
Which models were compared to the risk metrics model in the volatility forecasting example?
The initial estimation window for the volatility forecasting example used data from 2000 to the end of 2009.
The rolling window method performed better than the GARCH model in the volatility forecasting example.
The AR model performed best and beat the GARCH and rolling window models with -values close to ___.
Explain the purpose of using Diebold–Mariano tests in volatility forecasting.
What is the primary focus of feature-based evaluation of distributional forecasts?
Which term represents the squared average difference between the conditional mean given the forecast and the unconditional mean in the decomposition of the standard quadratic probability score?
What are the components of the decomposition of the standard quadratic probability score?
A good forecast has a calibration term as close to 1 as possible.
Calibration and resolution are independent measures in feature-based evaluation.
The standard quadratic probability score can be decomposed into three terms: unconditional variance, average squared calibration error, and the ___.
Explain the relationship between calibration and resolution in feature-based evaluation.
What is the purpose of decomposing the standard quadratic probability score?
When comparing the volatility forecasts of an AR model and a GARCH(1,1) model using Diebold–Mariano tests with MSE loss, which of the following statements is correct if the Diebold–Mariano test statistic shows a positive value with a -value close to 5%?
Which of the following statements are true regarding the evaluation of volatility forecasts using MSE and QLIKE loss functions?
The decomposition of the standard quadratic probability score shows that a good forecast, with a small MSE, should have calibration as close to 0 as possible and resolution as high as possible.
登录后解锁笔记、知识点解析、AI 问答
立即登录