正在学习

17.2.2 The Diebold–Mariano Test

17.2.2 The Diebold–Mariano Test

The approach of Diebold and Mariano (1995) explicitly takes into account the underlying loss function as well as sampling variation in the average losses. Suppose that two forecasts are available so we have two sets of losses, and . Consider the difference between the two losses:

Samples of time-series observations on the two losses and the associated loss differential, , form the basis of a test. Specifically, we can test the null hypothesis,

against either a one-sided alternative—if we want to see whether a particular method is better than the other—or against a two-sided alternative if we want to see whether the performances are different on average.

Diebold (2015) discuss conditions under which a set of simple tests of equal predictive performance can be conducted. The key assumptions are that, for all t,

Under these assumptions, a variety of methods are available to conduct the hypothesis test. The most obvious idea is to use a t-statistic which requires an estimate of the variance of the individual losses for scaling purposes. When the suitably scaled vector of out-of-sample losses defined in (17.2) satisfies a central limit theorem, , then

where and the covariance matrix  is partitioned after the first row and first column. Diebold and Mariano (1995) suggest estimating the scale nuisance parameter directly from the sequence of data , using the method of Newey and West (1987), although they note that any consistent estimator of the long-run variance would be applicable. For h-step-ahead forecasts with overlapping data, an structure is induced in if the underlying forecasts are rational so a window of at least observations should be used to construct the standard errors in (17.18).

Diebold and Mariano (1995) take forecasts as primitives. This means that models whose forecasts are strongly affected by parameter estimation error tend to be deemed inferior relative to models that are less strongly affected by estimation error, even if the latter model could provide the best fit to the data in large samples. Thus, we cannot conclude from the outcome of the Diebold–Mariano test which model is “best”—we can only make statements about which method generates the best forecast in a particular sample. Diebold and Mariano (1995) use the usual variance estimator as a consistent estimator for the scale of the difference, in (17.18).

When are the assumptions in (17.15)–(17.17) likely to hold? Suppose the parameters of a small and a large forecasting model have been estimated using a very short initial sample (e.g., 20 observations) and we are interested in studying the subsequent sequence of forecasts generated by the two models. In this situation, the initial forecasts from the large model are likely to be particularly strongly affected by estimation error. This effect will be reduced as the length of the estimation sample size increases, provided an expanding estimation window is used. In this situation, the sequence of loss differentials will not be stationary and (17.15)–(17.17) is unlikely to provide a good characterization of the behavior of the loss differentials. Specifically, the variance of the differential loss will decline over time, violating the second assumption in (17.17).

In other, less extreme, situations where either the initial estimation window is large or remains fixed so that the distribution of the sequence of loss differentials does not change over time (a case covered by Giacomini and White (2006) as we shall see below), assumption (17.17) is far more likely to be a good approximation to the data. Comparison of the accuracy of survey forecasts is another example in which the DM test can be useful. The methods that account for estimation error in the evaluation of forecast performance require that forecasts were generated according to a tight set of conditions that can be very hard to verify for such data.

If the focus is on finding out which method provides the best fit to the data and forecast errors are constructed using estimated model parameters, standard errors may have to be adjusted as suggested by West (1996) or West and McCracken (1998). In some cases one can use the corrected version of  from West (1996), replacing the usual estimator of  with that suggested in West (1996). Once again, nested models cannot be compared; if the models are nested,  will not be of full rank, and the long-run variance of converges to 0. The results of West and McCracken (1998) would be relevant for the Morgan–Granger–Newbold regression test, while those of West (1996) are appropriate for the Diebold–Mariano test.

Harvey, Leybourne, and Newbold (1997) modify the Diebold and Mariano t-test in two ways; first, by altering the divisor on the variance term and, second, by suggesting that the t-distribution with degrees of freedom should be used to construct the critical values. The modification follows from noting that the Diebold– Mariano test does not use degrees-of-freedom-adjusted variances. The suggestion to use the t-distribution is not based on theoretical concerns but serves to increase the critical values for tests that appear oversized in Monte Carlo experiments. For a general h-period forecast horizon, their modified t-statistic, labeled , is

where is the original Diebold–Mariano t-statistic. Notice that the correction adds second- and third-order terms which will have little effect except in very small sample sizes or for long horizons. for all h, the first-order effect is to reduce the size of the test statistic.7

Monte Carlo evidence in Harvey, Leybourne, and Newbold (1998) shows that their corrections have the desired directional effect on size—the corrections reduce the rejection rate under the null hypothesis by both increasing the critical value and decreasing the value of the test statistic. Size is not controlled for all values of h and all sample sizes, but the distortions are smaller than for the uncorrected t-statistic.

Most of the simulation evidence has been conducted under MSE loss, so is the difference between two squared terms. Squared errors tend to be quite skewed since they are bounded below at 0, and hence it is unsurprising that small sample tests based on the asymptotic normal distribution are oversized. This effect is exacerbated when the underlying forecast errors are fat tailed. Diebold and Mariano (1995) present simulation evidence that illustrates these effects when the forecast errors are drawn from stable distributions so the Monte Carlo design does not allow estimation error in the construction of the forecasts. Busetti and Marcucci (2013) conduct a Monte Carlo study of the size and power properties of a variety of tests for equal predictive accuracy under squared error loss for nested regression models. They find that the ranking of different tests is quite robust across settings with misspecified models and across different forecast horizons. They also find that highly persistent regressors give rise to a loss in power but do not affect the size of the test.

练习题

In the Diebold–Mariano test, what does the null hypothesis represent?

A. The two forecasts have the same expected loss
B. One forecast is better than the other
C. The variances of the two forecasts are equal
D. The forecasts are unbiased

Which of the following is a key assumption for conducting the Diebold–Mariano test according to the given text?

A. is a non - constant function of
B. for all
C.
D. The forecasts are always perfectly accurate

When using the Diebold–Mariano test for h - step - ahead forecasts with overlapping data, what should be used to construct the standard errors in (17.18)?

A. A window of 1 observation
B. A window of at least observations
C. A window of observations
D. No window is needed

Which of the following statements about the Diebold–Mariano test are correct?

A. It takes into account the underlying loss function
B. It only considers one - sided alternatives
C. It can be used to test the equal predictive performance of two forecasts
D. It assumes that models with less parameter estimation error are always better

What are the possible alternatives to the null hypothesis in the Diebold–Mariano test?

A. A one - sided alternative to see if a particular method is better
B. A two - sided alternative to see if the performances are different on average
C. The variances of the two forecasts are equal
D. The forecasts are uncorrelated

The Diebold–Mariano test can conclude which model is the “best” overall.

The Diebold–Mariano test assumes that the forecasts are primitives and ignores the impact of parameter estimation error on the forecasts.

In the Diebold–Mariano test, the difference between the two losses is defined as . The null hypothesis is ___.

For h - step - ahead forecasts with overlapping data in the Diebold–Mariano test, an structure is induced in if the underlying forecasts are rational, and a window of at least ___ observations should be used to construct the standard errors.

Explain the significance of the assumption in the Diebold–Mariano test.

How does the Diebold - Mariano test relate to the concept of forecast encompassing under MSE loss (kp_17_1_006)?

A. Both are used to determine the best - fitting model in large samples
B. The Diebold - Mariano test compares expected losses, while forecast encompassing under MSE loss determines if one forecast can be written as a function of others to minimize loss
C. They are completely unrelated concepts
D. The Diebold - Mariano test is only for one - step - ahead forecasts, while forecast encompassing is for multi - step

The Diebold - Mariano test and the regression - based test for loss equivalence (kp_17_2_005) both use the concept of comparing some form of difference in forecast errors.

In the regression - based test for loss equivalence (kp_17_2_005), the regression equation is . The null hypothesis is that ___.

Explain how the Diebold - Mariano test and the ENC - T test by Clark and McCracken (kp_17_1_1_5) are different in their focus and application.

登录后解锁笔记、知识点解析、AI 问答

立即登录